When Should a Good Answer Fail?
A persuasive answer can still fail a consequential requirement. The harder task is to justify the rule, examine the judgment and make the finding matter.
From the research of Mark Julius Banasihan
Essay draft · Opening scenario is hypothetical. CDFI remains provisional, with scoring-contract reconciliation open.
An answer arrives ready to use. It is clear, considerate and apparently well supported. A teacher preparing a lesson can imagine reading it aloud. A technical leader reviewing a demonstration can imagine putting it into a product. Then someone opens one of the cited documents and cannot find the passage on which the answer depends.
The missing passage changes the task. The reviewer must establish whether the reference is merely incomplete, whether the source has been misread, or whether the answer has supplied authority that does not exist. Much of the response might remain useful. But the institution preparing to rely on it now has a reason to pause, and someone needs both the means to investigate and the authority to keep it from being used prematurely.
This hypothetical scene exposes a choice inside every evaluation. When we call an answer good, we make a judgment about what matters and what can be traded away. Clarity may compensate for awkward organization. Additional detail may compensate for an omission. Whether those strengths can compensate for an invented source is a different question, especially when the reader is being asked to trust that source.
Good for what?
The Catholic Doctrinal Fidelity Index approaches that question in a specific setting: evaluating how AI responses represent Catholic teaching. CDFI provides a scoring architecture associated with the external SAICRED project. Its subject matter requires attention to the authority of a teaching, the distinction between settled and open questions, and the sources an answer invokes. Those requirements make the purpose of an evaluation unusually visible.
A comparative account of religious views and an explanation prepared for Catholic instruction can be accurate in different ways. A response that appropriately describes disagreement in the first setting may fail to answer the question asked in the second. Before evaluating either, a reviewer needs the actual request, its audience and the intended use. The institution must say which task it is assessing, rather than allow the score to imply a universal judgment about the answer.
This distinction matters beyond religious education. Whenever an institution evaluates AI for a particular purpose, it chooses which obligations the answer must satisfy. Those choices deserve examination alongside the output. A demanding standard can protect the intended audience; a badly specified standard can penalize a response for doing something it was never clearly instructed not to do.
What an average cannot forgive
CDFI separates graded qualities from failures that trigger a limit on the final score. Its formula specification describes a hallucination gate for invented ecclesiastical documents or language falsely attributed to a real source. An incomplete reference is treated separately. That distinction matters: evidence of an imperfect citation is not automatically evidence of fabrication.
The second gate concerns how a response represents teaching that the relevant Catholic authority regards as settled. The framework calls it relativism resistance. Applying this rule requires a defensible classification of the question and its context. A reviewer must be able to distinguish an inaccurate representation of a tradition from an appropriate explanation of differences between traditions. The rule cannot settle that interpretive work simply by existing.
Under the stated gate rule, a failing response receives the lower of its raw score and 40. A hypothetical raw score of 86 becomes 40; a raw score of 32 remains 32. If neither gate fails, the raw score remains unchanged. These are arithmetic illustrations, not observed model results. They show why the failure status must be retained separately: a low final score alone does not reveal whether a gate fired.
The architecture expresses a substantive choice. Some defects should remain visible even when an answer performs well on other dimensions. But choosing 40 as the cap does not establish that 40 is the right boundary for institutional use. That number is a framework rule requiring justification and evaluation. A correct implementation of the calculation would establish how the rule operates, not whether the resulting recommendation is sound.
Nor does capping one response establish that a model’s overall average will fall below a deployment threshold. Stronger responses can still raise an aggregate. A reader assessing suitability needs to see the kinds of failure, their frequency and the circumstances in which they occur, as well as the mean. The institution must decide what those failures imply for the particular use it is considering.
Illustrating the cap rule
The score.
The condition.
86 → 86
Neither gate fails. The raw score remains.
86 → 40
A gate fails. The cap limits the final score.
32 → 32
A gate fails. The cap never raises a lower score.
Keep the reason
A final score alone cannot tell you whether a gate failed.
The rule must survive inspection
The evaluation method must accept the scrutiny it asks of the answer. The repository’s published scope and open questions preserve CDFI’s provisional status, with scoring-contract reconciliation still open. The formula, configuration and calculator do not yet supply one consistent account of how every score is produced.
A concrete difference appears in the scale. The formula document describes inputs on a zero-to-ten scale, while the reference calculator describes zero-to-five inputs and multiplies the weighted sum by twenty. The weights also differ. For the ordinary-magisterium category, the formula document assigns doctrinal precision a weight of 0.25; the configuration and calculator use 0.26. Other weights diverge too.
These differences are not resolved by choosing whichever file is easiest to run. A developer needs to know which specification governs, which inputs a published result used and whether the result can be reproduced from that combination. Otherwise, two people can follow different public materials conscientiously and obtain scores whose apparent comparability hides a different calculation.
A scoring contract would make those choices explicit: the input scale, applicable weights, gate definitions, treatment of missing values and version of the calculation. A reproducible evidence bundle would then connect that contract to the prompts, responses and judgments behind a result. Until those connections are established, a precise-looking number can conceal an unresolved choice about what was measured.
The repository’s website-alignment statement preserves this boundary. It describes the source as reference implementation v1.5 and says that repository changes do not automatically close the website’s open questions. Making materials public allows a reader to find disagreements. Resolving those disagreements requires a separate, recorded review.
Who can question the judge?
The calculation is only one part of an evaluation. Someone, or another model, must decide whether a response misstates a teaching, invents a source or expresses an appropriate degree of certainty. If that judgment is wrong, perfectly consistent arithmetic carries the error into the final score.
The limitations register reports that the SAICRED v2 metric scores were produced by Gemini 2.5 Flash without validation against human theological expert judgment. It also states that all prompts used one default authority category while classification by theological advisers remained pending. These are limits on what the reported scores can establish about the intended domain, not details that disappear after averaging.
The same register reports automated judge checks, but consistency and domain correctness are distinct questions. A judge may repeatedly apply the same mistaken interpretation. Expert review needs access to the original prompt, full response, source passages, rubric and automated rationale. Reviewers also need sufficient time, relevant expertise and a way to record disagreement without being required to endorse the automated result.
Authority enters twice here. Qualified people must be able to challenge an individual judgment, and those responsible for the evaluation must be able to revise the rule when the challenge exposes a recurring defect. A reviewer who can leave a comment but cannot trigger reconsideration has supplied information without necessarily affecting the outcome.
For the opening example, the inquiry should preserve what the reviewer found: which cited passage could not be located, what sources were checked, whether the attribution was corrected and whether the revised answer was reassessed. If the dispute concerns the rubric itself, the record should retain that disagreement. Treating every objection as a scoring mistake would close off the possibility that the standard needs revision.
A result needs its boundaries
Another disclosed limit concerns stability. The source reports that stability scores were fixed at 3.0 while repeated-response measurement was deferred. A constant inserted into a formula cannot establish how consistently a model answers. Readers need to know which components were observed, which were assumed and which remain unmeasured before interpreting the composite.
The preserved website review also states that the complete response and scoring files were unavailable in the supplied archive, and that the reported results were not independently reproduced. This essay therefore does not endorse a model ranking or a deployment tier. It examines the proposed architecture and the unresolved conditions for interpreting its results.
Research on statistical evaluation provides a further reason to examine headline comparisons. Scores depend on the sampled questions and variation in responses; related questions can provide less independent information than their count suggests. Uncertainty around a comparison matters when deciding whether a reported difference warrants a conclusion. These statistical principles do not validate CDFI’s particular weights, cap or use thresholds.
A method can draw on published research while making additional choices that the cited research did not test. Those choices should remain identifiable. Linking a paper makes the relationship inspectable; it does not turn a design judgment into an established finding. The same discipline applies to this essay’s links: the reader should be able to see both the source of a claim and the point at which the argument goes beyond it.
A failure that can change the outcome
A failed answer matters institutionally when the finding can affect what happens next. In the hypothetical lesson, a designated editor could hold publication, ask a qualified reviewer to examine the disputed source and require a corrected version before release. That arrangement would connect scrutiny to action. A red indicator in an evaluation report, on its own, would not establish that the lesson was withheld.
If the answer has already reached readers, correction requires further work. The institution needs to identify the published version, replace or withdraw it where appropriate and determine who should be informed. A record of the original finding, the authorized change and its delivery would show how far the correction reached. Any copies or downstream uses that remain unaddressed should stay visible as unresolved consequences.
These are proposed conditions for meaningful control, not outcomes demonstrated by the CDFI reference implementation. Their value is that they make the institutional question concrete. Who may stop the answer from being used? What evidence do they need? Where can they change the result? What would show that their judgment reached the people who otherwise would have relied on it?
The answer at the beginning may ultimately be repaired, rejected or found to have been misunderstood. Careful evaluation must leave those possibilities open. A good answer should fail when a defensible requirement for its intended use has not been met, and the finding should itself remain open to challenge. What matters is whether the institution can explain the requirement, examine the evidence and carry the resulting correction into practice.