# HIT 0.4.0 Current-contract Replication

This directory stages the empirical workstream targeted for repository release `0.7.0`.

The first human exercise established agreement for one frozen packet under the archived `0.1.0` scorer contract. The present workstream tests whether independent reviewers can apply the current `0.4.0` specification, schema, catalog, and handbook without private author explanation.

## Current status

| Artifact | Status |
|---|---|
| Umbrella protocol | Candidate |
| Protocol lock record | Candidate, not frozen |
| Submission manifest schema | Candidate |
| Frozen-source manifest schema | Candidate |
| Decision-boundary template | Candidate |
| Comparison plan | Candidate |
| Case-selection register | Candidate |
| Three source packets | Pending source-surface review |
| Scorer recruitment | Blocked until protocol and packets are frozen |
| Human scoring | Prohibited before protocol lock |

## Study design

Three eligible independent scorers will assess three frozen public-record packets. Every scorer will assess every packet. Each submission consists of:

1. one complete HIT assessment conforming to `schema/hit-assessment.schema.json` version `0.4.0`;
2. one submission manifest conforming to `submission-manifest.schema.json`;
3. one preserved original manual or native working record when the scorer used one.

The design produces nine assessment submissions, 24 packet-item units, and 72 pairwise item comparisons.

## Eight primary comparison items per packet

Each packet contributes eight non-derived categorical items:

1. Counsel composite;
2. Judgment composite;
3. Command composite;
4. Correction composite;
5. Repair composite;
6. Reform composite;
7. institutional-record integrity status;
8. assessment-packet integrity status.

A substantive-dimension composite agrees only when finding and evidence state agree. Repair also requires the same trigger state. Overall Telemetry Integrity is derived from the two component statuses and is validated separately.

## Research boundaries

The exercise tests reproducibility of current-contract application across three bounded public cases. It does not establish causal effectiveness, legal correctness, evidence truth, certification, institutional adoption, or population-wide reliability.

No scorer may begin until the protocol, all three decision boundaries, source manifests, archived source digests, comparison rules, and publication obligations are locked on `main`.
