Research records provisional
Independent research

AI evaluation & assurance

Research note

Can independent reviewers assess human influence consistently?

AUTHOR

Mark Julius Banasihan

HUMAN RESULT

18 July 2026

Research status: Research in progress. The earlier bounded application remains available; broader research and validation continue. Status record · 2026-09-15

Human Influence Telemetry.

HIT is a documentary assurance method for assessing whether human authority had practical force in an institutional decision process. This note explains its first bounded independent application.

What the study found.

Two independent scorers assigned the same categories to all seven items in one frozen Cigna PxDx packet. The predeclared advancement gate was met without substantive adjudication.

What that means.

The exercise establishes agreement for this packet under scorer contract 0.1.0. It does not establish broad reliability, evidence truth, effective oversight, or lower harm. Replication under the current contract remains open.

Source materials.

Result summary · Repository citation · Registry record

Seven matching findings on one case

Both scorers assigned category 1 to Counsel, Judgment, Command, Correction, Repair, and Reform, and limited to Telemetry Integrity.

Under the preserved handbook, category 1 means the form of oversight was present but the records did not show practical capacity to alter the decision path. Agreement therefore does not mean the process received a favorable oversight assessment.

For an institutional reader, the result answers a narrower question: could two reviewers apply these categories to this same record and reach the same output? The answer in this exercise was yes.

Exact agreements
7 / 7
Reviewers
2
Frozen packet
1
Assessment itemScorer AScorer B
Counsel11
Judgment11
Command11
Correction11
Repair11
Reform11
Telemetry Integritylimitedlimited
Figure 1. The original categorical findings, displayed side by side. Category 1 means present but ceremonial; Telemetry Integrity uses its own vocabulary.
0
Absent
1
Present but ceremonial
2
Substantively exercised
IE
Insufficient evidence

Category labels are not percentages. Telemetry Integrity is assessed separately. No aggregate influence score is calculated.

What did human review contribute?

A signature or review step records human involvement. HIT asks what the available evidence supports about the person’s actual contribution and the institution’s capacity to correct, repair, and change its process.

For a technology or operations leader, the practical starting point is a bounded workflow: what records would let a second reviewer reconstruct the human role? The dimensions alongside translate that question into evidence requests.

These are proposed uses of the method. This study does not measure the benefit of introducing HIT into an institution.

Preserved application handbook.

Counsel
Information and evidence available to the reviewer
Judgment
Reasons, alternatives, dissent, and deliberation
Command
Authority to approve, reject, modify, or stop
Correction
Appeal, escalation, interruption, and reversal
Repair
Remediation, corrected decisions, and responsibility for repair
Reform
Changes to rules, systems, or institutional practice
Telemetry Integrity
Provenance, completeness, edit control, and retention

What exactly was assessed?

The exercise examined the described post-service PxDx claims-review workflow and the role of medical directors or physician reviewers. It used a fixed public record, rather than a representative sample of claim files.

The packet combines reporting, an institutional response, and a pleading-stage court order. Those sources have different evidentiary roles. An allegation or procedural ruling is not treated as a final merits finding.

For an institution designing a similar review, fixing the period, process, and source packet lets readers understand what the findings cover and what additional records could change them.

Full decision boundary.

Packet
HIT-IR-CIGNA-PXDX-001
Period
Reported 2022 activity through 31 March 2025
Sources
Investigative report; follow-up report; federal pleading-stage order
Scorers
Two eligible independent reviewers
Scorer contract
0.1.0
Protocol
HIT-IRP-CIGNA-001 · 1.0.0
Comparison
Six substantive findings plus Telemetry Integrity

How was independence preserved?

The protocol fixed the packet, submission requirements, comparison rule, and advancement threshold before scoring. Each scorer submitted findings separately.

The adjudication record reports independence attestations, access to all three sources, preserved manual submissions, and scorer confirmation of the JSON transcriptions. It also records that neither scorer changed a finding after seeing the other submission.

These are recorded attestations and preservation checks. This page has not independently investigated scorer eligibility. For an evaluation team, their value is an inspectable account of how the result was produced.

Adjudication and preservation record.

  1. FreezeDefine packet, rules, and advancement threshold.
  2. AssessObtain independent findings and rationales.
  3. PreserveVerify transcriptions and lock the submitted records.
  4. CompareCalculate agreement before adjudication.
  5. InterpretKeep the result attached to its original contract.

What does seven out of seven establish?

The predeclared gate required at least six exact agreements across seven items and zero critical disagreements. The observed result was seven agreements and zero critical disagreements.

The simple agreement proportion is the number of matching items divided by the total compared. Seven dimensions are not seven independent cases; the sample still contains one packet and two scorers.

Cohen’s kappa is not estimable here because all six substantive ratings occupied the same category. Perfect observed agreement therefore supplies no general chance-corrected reliability estimate.

This result supported the repository’s Level 2 / Applicable maturity decision. That designation remains bounded to the exercise; it is not a deployment certification.

Predeclared advancement rule

Required

6 / 7 or more exact matches

0 critical disagreements

Observed

7 / 7 exact matches

0 critical disagreements

The predeclared gate was met.

Exact agreement = 7 / 7 = 1.0

Substantive Cohen’s kappa: not estimable. All six substantive ratings occupy one category.

Original comparison and threshold

Matching categories can have different reasoning

Scorer A cited S1 for all seven findings; Scorer B cited S2 for all seven. The adjudication record notes differences in their analytical vocabulary as well. Those differences did not produce a categorical disagreement.

For an assurance reviewer, this is a reason to preserve rationales alongside categories. A matching output alone does not reveal whether reviewers relied on the same evidence or made the same interpretive steps.

No substantive adjudication was required. The original findings remain the primary result.

The categories match. The source references and reasoning remain separate records.

A traceable claim can still lack adequate evidence

The repository’s v0.6.5 audit records five checks: traceability, integrity, human support review, evidence fitness, and dependency closure. The figure groups the evidence-fitness findings and preserves the results for the other four checks.

For example, H3 concerns the completed frozen-packet exercise. H5 proposes that higher HIT findings predict lower harm or better repair; the claim map marks it blocked because an outcome study is absent.

For a technology leader, this distinction prevents a working implementation or well-documented result from being used as evidence for a stronger claim it never tested. The figure reports the repository’s own audit state, not an independent validation of the method.

Evidence fitness · 13 mapped claims

Pass
6 claims
Fail
5 claims
Indeterminate
2 claims
Traceability
13 / 13 pass
Integrity
13 / 13 pass
Human support review
13 / 13 pass
Dependency closure
13 / 13 pass

6 of 13 claims are eligible for a conclusion in the source audit.

Figure 2. Node & Norm presentation of the repository’s v0.6.5 gate data. These are claim checks, not participants or deployment outcomes.

What needs to be tested next?

The supplied repository stages a separate current-contract replication. It proposes three reviewers each assessing three frozen packets using the 0.4.0 specification.

The plan compares eight non-derived items per packet, including separate institutional-record and assessment-packet integrity statuses. It is not a rerun of the seven-item historical contract.

Protocol and packet freezing remain prerequisites. Candidate materials do not establish a new completed result or advance maturity. For teams considering HIT, current-contract reproducibility and transfer to their own evidence conditions remain open questions.

Candidate replication design.

Preserved historical exercise

Scorer contract · 0.1.0
The rules used for the seven-item comparison.
Human result · 0.6.0
Released 18 July 2026.
Repository citation · 0.6.5
Released 9 August 2026.

Current specification and implementation

Normative specification · 0.4.0
A separate contract from the historical result.
Conformance engine · 0.5.0
The implementation layer.

Proposed replication

0.7.0 candidate
Three reviewers × three frozen packets proposed. Not locked; scoring prohibited before lock.

Claim-by-claim evidence record

Each entry preserves the source claim, its status, evidence-fitness result, and eligibility to enter a conclusion. All thirteen claims are visible here; the source JSON retains the longer rationales and evidence locators.

The scorer ratings appear in Figure 1 above. These records describe a documentary method and one application. They do not establish legal correctness, evidence truth, causal effectiveness, institutional adoption, or a validated aggregate influence score.

  1. H1Supported

    HIT represents six substantive dimensions and split Telemetry Integrity in a machine-readable assessment.

    Conclusion eligible: Yes · Evidence fitness: pass

  2. H2Supported

    The evidence-state and finding rules distinguish affirmative absence, ceremonial presence, substantive exercise, and insufficient evidence in the executable boundary suite.

    Conclusion eligible: Yes · Evidence fitness: pass

  3. H3Supported

    Two eligible independent reviewers reached the predeclared threshold on one frozen Cigna packet under the preserved 0.1.0 scorer contract.

    Conclusion eligible: Yes · Evidence fitness: pass

  4. H4Provisional

    HIT can classify ceremonial oversight where a binary human-presence label would record human involvement.

    Conclusion eligible: No · Evidence fitness: indeterminate

  5. H5Blocked

    Higher HIT findings predict lower harm or better repair.

    Conclusion eligible: No · Evidence fitness: fail

  6. H6Unsupported

    A high HIT finding proves meaningful human control in a causal or legal sense.

    Conclusion eligible: No · Evidence fitness: fail

  7. H7Provisional

    HIT evidence may inform external oversight, accountability, and contestability analysis within a separately reviewed applicability boundary.

    Conclusion eligible: No · Evidence fitness: indeterminate

  8. H8Blocked

    HIT has been independently adopted by an institution outside the author's control.

    Conclusion eligible: No · Evidence fitness: fail

  9. H9Pending

    The public HIT package can be installed, operated, and explained without private author interpretation.

    Conclusion eligible: No · Evidence fitness: fail

  10. PAPER-C01Supported

    HIT is a documentary assurance method with a stable 0.4.0 normative contract and a separate 0.5.0 conformance engine.

    Conclusion eligible: Yes · Evidence fitness: pass

  11. PAPER-C02Supported

    The published human result is bounded to one frozen packet and does not resolve current-contract replication.

    Conclusion eligible: Yes · Evidence fitness: pass

  12. PAPER-C03Supported

    The v0.6.5 audit detects its prespecified corruptions and blocks mapped claims whose required gates do not pass.

    Conclusion eligible: Yes · Evidence fitness: pass

  13. PAPER-C04Blocked

    HIT has established population-wide inter-rater reliability under the current 0.4.0 contract.

    Conclusion eligible: No · Evidence fitness: fail

Sources and version notes

This page was checked against the author-supplied human-influence-telemetry-main.zip on 11 September 2026. The archive’s citation identifies version 0.6.5. Its commit identity has not been inferred from the archive name.

The ZIP includes historical and unreleased entries; older references to 0.6.4 do not supersede the 0.6.5 citation. The human result remains 0.6.0. Candidate 0.7.0, 0.9.0, and 1.0.0 documents describe future work.

Archive SHA-256: e555c53784dc8966147330a1acebf63d6155aae874b4916bcbab5445e70814d2.

Citation metadata ↓ · Release history · Source license · Complete claim-evidence map

Source downloads remain unchanged copies. Figures 1 and 2 are Node & Norm presentations of the source findings and gate data; the original figure presentation remains linked. Archive inspection verifies transcription and provenance; research admission remains provisional.