Research records provisional
Independent research

AI evaluation & assurance

Studies & methods

Can the evaluation
support the decision?

Node & Norm examines whether reviewers can apply a method consistently and whether scoring rules work as described. Explore the findings, source materials, and limits of each study.

One preserved reviewer comparison

7 / 7

matching findings
2 reviewers · 1 frozen packet

HIT · Documentary assessment

Can independent reviewers reach the same assessment?

Human Influence Telemetry examines documented human influence. In this historical exercise, two scorers agreed on all seven items under contract 0.1.0.

Agreement concerns this packet. It does not establish broad reliability or effective oversight.

Research in progress

An evaluation method under review

Scoring contract
unresolved.

Written formula · weights · calculator

CDFI · Evaluation-method review

Do the scoring rules match the calculation?

The Catholic Doctrinal Fidelity Index provides a domain-specific example of AI evaluation. Its source review exposes disagreements between the written formula, weights, and calculator.

The scoring contract remains unresolved. CDFI continues to evolve alongside the external SAICRED project.

Evolving

Start with the claim. Make the test inspectable.

Before using an assessment to approve a system or change a workflow, a team needs to know what the result supports. Could another reviewer reproduce it? Did the calculation follow the stated rules? What evidence would change the conclusion?

Each note identifies the source record, the method and version used, the finding, and the questions still open. Current work statuses are recorded separately from earlier results in the canonical Registry.

Inspect the work behind the result.

Use the preserved materials to examine the rules, trace a finding, or plan a replication. The notes identify missing data and changes between the reviewed version and later work.