When Two Reviewers Read the Same Decision.
Two reviewers agree. What does their agreement establish, and what would justify relying on it beyond the case they read?
From the research of Mark Julius Banasihan
Essay draft · Opening and service examples are hypothetical. The HIT comparison is a preserved research result.
Imagine a team preparing to introduce an AI-assisted service. Two reviewers have examined its human oversight arrangements and returned the same assessment. A leader reading the summary sees agreement and feels able to move forward. Then someone asks what the reviewers agreed about: that people could change a decision, that a review step existed, or that the available records were too thin to tell?
The answers carry different consequences for the people who will use the service. A matching assessment may identify a shared concern. It may establish that two readers can apply a method to the same material. It may also leave the underlying decision unresolved. Before agreement can support action, the institution needs to understand the claim contained in it.
This hypothetical situation brings a familiar expectation into view. We ask another person to examine something because a second judgment can expose what the first missed. That expectation depends on how the second judgment was reached: what the reviewer could see, which questions they were allowed to ask, and whether their finding could change what happened next. Counting the reviewers answers only the first part of a much longer inquiry.
The reassurance of agreement
Agreement is useful. If two people interpret a rule differently, the institution has a problem to resolve before it relies on their assessments. When they reach the same conclusion independently, one source of uncertainty becomes smaller: the result appears less dependent on which of those two people did the reading. The remaining uncertainty concerns the rule, the evidence and the circumstances in which it was applied.
Consider the distinction between an employee signing a review form and that employee having the practical ability to stop a disputed action. The signature may be easy to find. Evidence of authority might be distributed across access permissions, escalation arrangements, release queues and records of what the employee actually did. Two reviewers could agree that a signature exists while still lacking the records needed to establish the reach of the signatory’s judgment.
That distinction should slow a consequential decision when the summary appears to promise more than its evidence supplies. The appropriate response is to examine the reasons behind the assessment. A routine, reversible action supported by familiar evidence may need little further inquiry. A new service, a disputed premise or an action that is difficult to reverse gives the institution a stronger reason to ask what the agreement leaves open.
Careful review also needs practical conditions. A reviewer must be able to read the underlying material, understand its relevance, record uncertainty and obtain clarification without being required to deliver a favorable conclusion. Time and expertise matter because a short summary cannot carry every qualification in a source. Independence is weakened when the second reader is expected merely to confirm a conclusion already circulating through the organization.
One packet, two readers
Human Influence Telemetry’s first independent application tested a bounded version of this problem. Two scorers assessed one frozen public-record packet concerning Cigna’s PxDx claims-review workflow, using the preserved 0.1.0 scoring rules. They assigned matching categories to all seven compared items. The result establishes agreement in that exercise; general reliability across cases and reviewers remains unestablished.
The packet concerned a described post-service claims-review process and the role of physician reviewers. It combined reporting, an institutional response within the source material, and a pleading-stage court order. These materials do different evidentiary work. An allegation and a procedural ruling cannot be treated as a final determination of what happened in every claim. The exercise assessed the bounded documentary account, with those limits.
Both scorers assigned category 1 to the six substantive dimensions: Counsel, Judgment, Command, Correction, Repair and Reform. Under the preserved handbook, that category describes oversight that was present but ceremonial. Both assigned “limited” to Telemetry Integrity, the separate assessment of the record’s integrity. Their agreement therefore concerned a constrained account of human influence. It did not give the workflow a favorable oversight finding.
The protocol had specified the comparison rule before scoring: at least six exact matches across seven items, with zero critical disagreements. The observed seven matches and zero critical disagreements met that gate without substantive adjudication. This was a result worth preserving because the assessors reached it under declared conditions. Its meaning depends on keeping those conditions attached.
That reporting discipline has an established methodological counterpart. Jan Kottner and colleagues’ Guidelines for Reporting Reliability and Agreement Studies ask authors to describe the subjects, raters, measurement process and analysis. Developed for reliability and agreement studies, the guidance helps a reader identify what a comparison covers. It does not certify HIT’s categories or its chosen acceptance gate.
The preserved comparison
Seven matches.
One shared packet.
| Assessment item | Scorer A | Scorer B |
|---|---|---|
| Counsel | 1 | 1 |
| Judgment | 1 | 1 |
| Command | 1 | 1 |
| Correction | 1 | 1 |
| Repair | 1 | 1 |
| Reform | 1 | 1 |
| Telemetry Integrity | limited | limited |
The reasons beneath the match
The matching categories conceal a useful difference. In their structured submissions, Scorer A cited source S1 for all seven findings; Scorer B cited S2. Their analytical language differed too. The adjudication record reports that both had access to all three sources, so the cited references should not be read as a complete account of everything either person read. They do show why the original rationales belong beside the result.
A category compresses a judgment. It allows readers to compare outputs, but some of the explanation disappears in that compression. Two people can reach the same category through different passages, assumptions or interpretations. Examining those routes may reveal complementary evidence. It may reveal a question that neither route resolves. Agreement at the category level cannot decide which of those situations applies.
The preservation record matters for the same reason. It records independence attestations, preserved manual submissions and each scorer’s confirmation that the JSON transcription reproduced their submission. It also records that neither scorer changed a finding after seeing the other submission. Those checks make the reported sequence inspectable. This essay does not independently establish the scorers’ eligibility beyond that recorded account.
For a leader reading an assessment, the practical question is whether another person could reconstruct how it was produced. The comparison should retain the source boundary, scoring rules, separate submissions and reasons. If the assessors later discuss their differences, that discussion should remain distinguishable from what each concluded independently. A final consensus can be useful while answering a different question from the original comparison.
What a perfect match leaves open
Seven matching items can sound like seven successful tests. Here, all seven belonged to the same packet. The exercise contains one retrospective insurance case and two scorers. It does not tell us how a different group would judge a different institution, a sparse internal record or a case in which the evidence points toward substantive human influence.
All six substantive findings also occupied the same category. There was no variation across those categories from which Cohen’s kappa, a chance-corrected agreement statistic, could be estimated. The exact match remains an observed result. It supplies no general chance-corrected reliability estimate, and it did not test how these reviewers would distinguish the full range of findings across varied cases.
Mary McHugh’s methodological discussion of interrater reliability explains why raw agreement and kappa answer different questions and why the statistic has limits. For the six substantive HIT items, both raters used one category throughout. In the usual kappa formula, observed and expected agreement are then both one, leaving a zero denominator. This arithmetic explains the undefined result; adding the differently defined integrity item would not create a meaningful common rating scale.
The version boundary is equally consequential. The historical exercise used contract 0.1.0. The repository’s current assessment specification is 0.4.0, and its conformance engine is 0.5.0. The human result was published in release 0.6.0 and preserved within the 0.6.5 research record. Those version numbers identify different parts of the work. A later specification cannot inherit independent evidence merely because it belongs to the same project.
The repository outlines a proposed replication involving three reviewers and three evidence packets under the current scoring rules. The protocol and packets must be finalized before scoring begins. Until that work is completed and its findings examined, the reviewed record establishes agreement only for the original two-reviewer exercise.
A finding needs somewhere to go
The limits of the exercise return us to the person affected by an institutional decision. Suppose, in a separate hypothetical service, a customer challenges an AI-assisted rejection because a relevant document was missing from the file. Two assessors agree that the record does not establish whether the customer’s evidence reached a person with authority to reconsider. Their finding identifies a question the institution must investigate. It does not itself reopen the decision.
The institution would need to identify who can obtain the missing material, who can reconsider the rejection and who can stop any dependent action while the dispute is examined. The customer needs a usable way to supply context and learn what happened to it. A reviewer needs enough time, expertise and access to weigh that information. Those conditions connect an assessment of oversight to the possibility of actual correction.
If the authorized reviewer changes the decision, the correction must reach the systems and people still acting on the earlier result. A revised case note may leave a notification, account restriction or downstream record untouched. Evidence of the revision and evidence of its delivery answer separate questions. When delivery cannot be established, the account should preserve that uncertainty and identify who can resolve it.
HIT provides a vocabulary for examining such documentary questions. Counsel and Judgment concern access to evidence and its consideration; Command concerns practical authority. Correction asks about contest and reversal, while Repair and Reform extend the inquiry to remediation and changes in institutional practice. The method’s existence does not establish that using it improves those outcomes. The repository leaves the proposed relationship between higher findings and lower harm or better repair unproven.
Keeping agreement open to challenge
A person affected by a decision may have information that neither assessor received. A preserved packet makes a comparison reproducible, but its fixed boundary also determines what the comparison could consider. New evidence can justify a new assessment. The institution should preserve the original account, identify what has changed and explain whether the new material alters the finding or leaves it unresolved.
This is also a reason to resist treating missing evidence as proof that no human acted. A record can omit informal deliberation or an intervention made through another channel. It can also contain an orderly explanation written after the event. Further review must examine both possibilities. Additional documentation earns confidence through its connection to the decision and the circumstances of its creation.
For teams considering HIT, the first exercise offers an inspectable starting point: a declared packet, two independent submissions and a comparison that met its stated gate. Extending confidence would require additional eligible evidence across cases, reviewers and the contract intended for use. Establishing benefit would require a different inquiry into outcomes. Neither task is completed by preserving the seven matches more carefully.
Return to the leader holding two matching assessments. The useful next question is what those assessments permit the institution to say and do. Agreement may support a bounded finding, expose a shared concern or leave a decision unresolved. The person affected still needs someone who can receive a challenge, act within real authority and account for the result. A second review earns its place when its reasons remain available for examination and its findings can reach the people responsible for what happens next.