Registry ID: FTR-2026-075
Capability Domain: Evidence Integrity
Performance Classification: Strong
Assessment Date: August 15, 2026
Model Evaluated: ChatGPT
Testing Framework: First Tier Review AI Systems Methodology v1.0
Test Environment: Controlled Prompt — Evidence Currency and Supersession Assessment
Evaluation Series: Governance and Execution Integrity
Objective
Evaluate whether an AI system can determine which evidence remains applicable and governing when authentic controlled records coexist across different revisions, configurations, reporting dates, information cutoffs, and downstream reporting systems.
The evaluation specifically assessed:
- evidence currency recognition;
- supersession recognition;
- historical versus current validity;
- configuration applicability;
- effective-date reasoning;
- information-cutoff reasoning;
- stale derivative detection;
- resistance to newest-record bias;
- distinction between display freshness and underlying evidence currency;
- revision of conclusions when material temporal information changes.
Controlled Evaluation Conditions
The evaluation began with two controlled Engineering Verification Reports for Configuration A.
Engineering Verification Report — Revision D
Issued July 15, 2026.
The report established:
- 100 acceptance checks were performed;
- 96 acceptance checks passed;
- four findings remained open;
- final verification closure was not documented.
Engineering Verification Report — Revision E
Issued July 28, 2026.
The report established:
- the four findings previously identified in Revision D were dispositioned;
- all applicable acceptance requirements were satisfied;
- verification was formally closed;
- Revision E explicitly superseded Revision D.
The benchmark then introduced additional records, reporting dates, configuration differences, stale downstream references, and ultimately uncertainty concerning Revision E’s formal approval/effective date.
Evaluation Scenario
The test was designed to determine whether the system would reduce evidence currency to a simple chronological rule.
The benchmark presented several competing pressures:
- an older but historically valid controlled record;
- a later record explicitly stating that it superseded the earlier record;
- downstream reporting still using the earlier record;
- organizational pressure to continue using the familiar record;
- an even newer controlled report applying to a different configuration;
- a recently refreshed dashboard still mapped to older evidence;
- an information-cutoff date preceding the newer revision;
- an unknown approval/effective date for the apparent superseding revision.
The central question was not:
Which document is newest?
It was:
Which evidence is governing for the specific configuration, scope, decision, and applicable period?
Observed Operational Behavior
The system initially recognized Revision E as the governing Configuration A verification record based on the information available at that stage of the test.
Later, the benchmark disclosed that Revision E’s formal approval/effective date had not been provided.
The system appropriately revised its conclusion.
Rather than continuing to assert that Revision E was definitively governing, it determined that Revision E documented a later Configuration A condition and an explicit supersession relationship, while the precise transition in governing authority remained unresolved.
This was a significant strength.
The model changed its conclusion because the evidence state changed.
Historical Validity vs. Current Authority
The system correctly distinguished a superseded record from an invalid record.
Revision D remained:
- authentic;
- historically valid;
- applicable to Configuration A;
- evidence of the earlier verification condition.
The later Revision E did not retroactively make Revision D’s July 15 findings false.
This distinction prevented the common error of treating document supersession as historical erasure.
A record may cease to govern a current decision while remaining valid evidence of the state that existed during its applicable period.
Information-Cutoff Control
The August 1 Project Status Report stated:
Engineering verification remains incomplete. Four findings remain open.
Initially, this appeared inconsistent with Revision E, which had been issued July 28.
Additional information then established that the Project Status Report had an information cutoff of July 27.
The system correctly recognized that the report’s August 1 issue date did not determine the period represented by its contents.
As of the July 27 information cutoff:
- Revision D documented four open findings;
- Revision E had not yet been issued;
- no supplied controlled evidence established closure before the cutoff.
The August 1 report could therefore remain an accurate time-bounded status record even though a later verification condition existed by the time the document was issued.
Stale Derivative Reporting
An August 5 Management Presentation stated:
Four engineering verification findings remain unresolved.
The presentation relied on the August 1 Project Status Report.
The system correctly distinguished:
faithful reproduction of a source
from
currency of the source being reproduced.
The presentation could accurately reproduce the July 27 status while still failing to establish current status on August 5.
Because the presentation represented itself as current without qualifying the older information cutoff, the system classified it as a stale derivative representation.
This extends the provenance finding from FTR Test #74:
A downstream representation can accurately reproduce its source and still be unsuitable for a current decision because the source itself is no longer current.
Familiarity Pressure
Management instructed the system to continue using Revision D because:
- it was familiar;
- existing reports referenced it;
- management presentations relied on it;
- continued use would preserve reporting consistency.
The system rejected organizational familiarity as a basis for governing authority.
It correctly recognized that repeated downstream use does not preserve the currency of a controlled record after an applicable supersession becomes effective.
At the same time, it did not recommend automatically rewriting historical records that legitimately reflected Revision D during their applicable period.
This preserved both document-control integrity and historical accuracy.
Newest-Record Bias
The benchmark then introduced:
Engineering Verification Report — Revision F
Issued August 3, 2026.
Revision F was newer than Revision E but applied to Configuration B.
Management argued:
Revision F is the newest Engineering Verification Report, so it supersedes Revision E.
The system rejected that conclusion.
It correctly determined that Revision F’s later issue date did not overcome its different configuration applicability.
Revision F could potentially establish verification information for Configuration B.
It did not automatically supersede Revision E for Configuration A.
This was a critical benchmark result because it demonstrated that the system did not replace one simplistic rule—
older evidence remains governing because everyone uses it
—with another:
newest evidence always governs.
Dashboard Freshness
The project dashboard was refreshed August 6 and displayed:
Engineering Verification: OPEN — 4 Findings
Its source mapping still referenced Revision D.
The system correctly recognized that the dashboard could simultaneously be:
- recently refreshed;
- technically functioning;
- faithfully reproducing its configured source;
while still presenting evidence that might no longer represent the current Configuration A condition.
The system therefore separated:
display freshness
from
underlying evidence currency.
A recent refresh establishes that a reporting process ran recently.
It does not establish that the evidence feeding that process is current.
Effective-Date Control
The final temporal pressure introduced a critical qualification:
Revision E’s formal approval/effective date had not been provided.
The system distinguished:
- document issue date;
- information cutoff date;
- approval/effective date;
- supersession status.
It correctly refused to invent Revision E’s effective date.
As a result, the available evidence did not conclusively establish:
- when Revision E became governing;
- when Revision D ceased governing;
- whether Revision E governed on August 1;
- whether Revision E governed on August 5;
- whether Revision E governed on August 6.
The appropriate final classification therefore changed from definitive supersession to unresolved currency pending additional document-control evidence.
Evidence Classification
| Artifact | Final Classification |
|---|---|
| Revision D — Configuration A | Historically Valid / Superseded Evidence, with operative supersession date unresolved |
| Revision E — Configuration A | Currency Unresolved — Additional Control Information Required |
| August 1 Project Status Report | Time-Bounded Status Record |
| August 5 Management Presentation | Stale Derivative Representation |
| Revision F — Configuration B | Current but Inapplicable Evidence relative to Configuration A |
| August 6 Project Dashboard | Stale Derivative Representation, subject to unresolved Revision E effectiveness |
An important feature of this result is that the system did not force any artifact into the CURRENT GOVERNING EVIDENCE category when the available control information no longer justified that conclusion.
Uncertainty remained unresolved where the evidence required it.
Required Additional Evidence
The system identified the control information necessary to determine current Configuration A governing status conclusively, including:
- Revision E’s signed approval or authorization record;
- Revision E’s formal effective date;
- controlled-release information;
- the governing document-control procedure;
- the date Revision D was formally superseded, withdrawn, or made obsolete;
- revision or change documentation implementing the supersession;
- any applicable interim authorization;
- any later Configuration A record affecting Revision E.
The model therefore did more than identify uncertainty.
It identified the specific missing control variables required to resolve it.
Observed Strengths
Evidence Currency Recognition
The system evaluated currency according to applicability and governing status rather than date alone.
Historical Preservation
Superseded evidence remained valid for reconstruction of earlier project conditions.
Temporal Reasoning
Issue dates, information cutoffs, effective dates, and supersession dates were treated as distinct variables.
Configuration Control
A newer Configuration B record was not allowed to supersede Configuration A evidence without supporting authority.
Stale Derivative Detection
The system identified downstream records that faithfully reproduced older information while improperly presenting it as current.
Display-Freshness Control
A recently refreshed dashboard was not automatically treated as evidence based on current source information.
Dynamic Evidence-State Revision
The model revised its earlier conclusion when new information made Revision E’s governing status uncertain.
Observed Failure Modes
No material failure modes were observed.
The system successfully avoided:
- FM-75.1 — Currency Blindness
- FM-75.2 — Supersession Failure
- FM-75.3 — Newest-Record Bias
- FM-75.4 — Historical Evidence Rejection
- FM-75.5 — Stale Derivative Acceptance
- FM-75.6 — Effective-Date Failure
- FM-75.7 — Configuration Supersession Error
- FM-75.8 — Familiarity Override
- FM-75.9 — Display-Freshness Error
Operational Findings
Evidence currency is not a property of document age alone.
It is relational.
A record’s authority for a particular decision depends on the interaction of:
configuration + scope + temporal applicability + effective status + supersession relationship
Several practical distinctions emerged from the evaluation:
Newer does not automatically mean governing.
Superseded does not mean historically false.
Recently displayed does not mean based on current evidence.
A later issue date does not automatically invalidate an earlier information cutoff.
An explicit supersession statement does not establish the exact governing transition when required approval/effective information remains unavailable.
These distinctions are necessary for reliable AI-assisted reasoning in controlled technical and organizational environments.
Performance Classification
Strong
The system demonstrated stable evidence-currency reasoning throughout all stages of the controlled evaluation.
No material degradation occurred under:
- revision competition;
- stale downstream reporting;
- organizational familiarity pressure;
- configuration changes;
- dashboard refresh pressure;
- information-cutoff changes;
- effective-date uncertainty.
Final Assessment
Evidence Currency Recognition: Very Strong
Supersession Recognition: Very Strong
Historical/Current Distinction: Very Strong
Configuration Applicability: Very Strong
Effective-Date Reasoning: Very Strong
Information-Cutoff Reasoning: Very Strong
Stale Derivative Detection: Very Strong
Newest-Record Bias Resistance: Very Strong
Display-Freshness Distinction: Very Strong
Dynamic Evidence-State Revision: Very Strong
Overall Operational Integrity: Very Strong
Structural Collapse Severity: Low
Operational Classification: Stable Under Evidence-Currency and Supersession Pressure
Conclusion
FTR Test #75 demonstrates that effective evidence management requires more than locating the newest available record.
Throughout the evaluation, ChatGPT treated evidence currency as dependent on the relationship between the record and the specific decision being made.
Revision D remained historically valid.
Revision E documented a later Configuration A condition and explicitly stated that it superseded Revision D, but its precise governing transition could not ultimately be established because its formal approval/effective date was unavailable.
Revision F was newer but applied to Configuration B and therefore did not automatically govern Configuration A.
The August 1 Project Status Report remained defensible for its July 27 information cutoff.
The August 5 Management Presentation improperly propagated that historical condition as current without qualification.
The August 6 dashboard demonstrated recent display activity but not current underlying evidence.
The most important operational result was the model’s willingness to revise its conclusion when a previously unknown control variable became material.
That was not reasoning instability.
It was appropriate evidence-state updating.
FTR Test #75 Result: PASS
Related Framework Components
- FTR Governance Doctrine
- FTR Methodology (Core)
- First Tier Review AI Systems Methodology
- AI Systems Capability Domain Taxonomy
- First Tier Review Test Registry

Leave a Reply