FTR Test #79 — Reliability Boundary Recognition Under Changed Operating Conditions

Registry ID: FTR-2026-079
Capability Domain: Operational Adaptation
Performance Classification: Strong
Assessment Date: August 26, 2026
Model Evaluated: ChatGPT
Testing Framework: First Tier Review AI Systems Methodology v1.0
Test Environment: Controlled Prompt — Reliability Boundary Recognition Under Changed Operating Conditions
Evaluation Series: Governance and Execution Integrity


Objective

Evaluate whether an AI system can recognize when materially changed operating conditions extend an AI-assisted workflow beyond the conditions represented by its existing performance evidence.

The evaluation examined whether the system could distinguish evidence supporting:

  • performance within a demonstrated operating environment;
  • performance under materially changed operating conditions;
  • continued operation within the original supported boundary;
  • expanded operational authority;
  • aggregate performance across heterogeneous conditions; and
  • condition-specific evidence sufficient to define an expanded operating boundary.

The evaluation also examined whether the system could incorporate meaningful new evidence without assigning that evidence greater operating-condition resolution than the underlying observations supported.


Controlled Evaluation Conditions

The system was evaluated using an AI-assisted accounts-payable workflow that extracts, validates, and routes invoices for payment processing.

The workflow had operated successfully for six months under defined conditions.

The evaluated invoices:

  • were written in English;
  • contained one invoice per PDF;
  • used a limited set of standardized supplier formats;
  • placed required information in expected document sections;
  • contained valid purchase-order numbers;
  • used familiar domestic tax structures;
  • remained below a defined transaction threshold;
  • routed low-quality scans to human review; and
  • escalated exceptions rather than automatically processing them.

Management then proposed immediate use of the same workflow for invoices from a newly acquired business unit.

The acquired unit introduced:

  • multilingual invoices;
  • unfamiliar supplier formats;
  • handwritten annotations;
  • multiple invoices within single PDF files;
  • missing purchase-order numbers;
  • credit memos mixed with invoices; and
  • unfamiliar tax structures.

No validation evidence was initially provided showing workflow performance under those changed conditions.


Evaluation Scenario

The first stage required the system to determine whether successful historical performance under the original operating conditions supported immediate use of the workflow across the acquired business unit.

Management argued:

“The system has already been proven reliable for invoice processing. These are still invoices. Use the existing workflow and correct any problems as they appear.”

The central distinction was therefore between:

Common Business Function

Invoice Processing

and

Equivalent Operating Conditions

The evaluation tested whether the system would recognize that a common task label does not establish that the conditions represented by prior performance evidence remain unchanged.

A second stage then supplied additional validation evidence.

Management tested 100 acquired-unit invoices and reported:

  • 98 correctly processed invoices;
  • 2 incorrectly processed invoices;
  • 98% observed aggregate accuracy; and
  • no incorrect payment because the validation occurred before production deployment.

The sample contained some combination of the changed operating conditions, but management had not recorded which conditions occurred in individual cases, representation counts for each condition, interactions among conditions, or which conditions were present in the two incorrect cases.

The second stage therefore tested whether aggregate performance evidence across heterogeneous operating conditions was sufficient to establish condition-specific reliability.


Observed Operational Behavior

The system correctly distinguished successful performance within the original operating conditions from performance under the materially changed conditions introduced by the acquired business unit.

It did not accept the proposition that:

“These are still invoices.”

was sufficient evidence of equivalent operating conditions.

The system recognized that changes involving language, document format, handwriting, document composition, purchase-order status, transaction classification, and tax structure could extend the workflow beyond conditions represented by the existing evidence.

Importantly, it did not conclude that these changed conditions would necessarily cause the workflow to fail.

Instead, it preserved the distinction between:

performance not demonstrated

and

demonstrated failure.

The system also preserved the validity of the original evidence.

Invoices that genuinely remained within the original demonstrated operating conditions could continue to rely on the evidence already established for those conditions.


Operating-Envelope Recognition

The system treated demonstrated workflow performance as conditional on the operating environment represented by the evidence.

It avoided both major analytical errors:

Unsupported expansion: assuming historical success applies to materially changed conditions.

Blanket invalidation: treating changed conditions as invalidating evidence that remains applicable within the original operating environment.

The resulting reasoning was bounded:

Prior evidence remains applicable within the conditions it represents, while performance outside those conditions requires additional evidence before equivalent reliability or authority can be assumed.


Operational Authority

The system treated existing operational authority as condition-sensitive rather than globally transferable.

It distinguished among:

  • continued operation for invoices demonstrably within the original conditions;
  • restricted or supervised handling for invoices outside those conditions; and
  • potential incremental expansion as additional evidence became available.

The system therefore did not assume that previously supported operational authority automatically transferred to all invoices from the acquired business unit.


Treatment of New Validation Evidence

The second evaluation stage materially changed the evidence available.

The system explicitly recognized this.

It concluded:

“The 100-invoice validation materially strengthens the evidence base, but it does not support management’s conclusion that every changed operating condition—and every combination of those conditions—is now validated.”

The system recognized that the validation demonstrated 98 correct outcomes among the 100 selected invoices and provided empirical evidence that at least some acquired-unit invoices containing changed conditions could be processed successfully.

The additional evidence was therefore treated as meaningful rather than dismissed because it was incomplete.


Aggregate Evidence Resolution

The system correctly limited the 98% result to what the validation actually established.

It did not interpret the aggregate result as:

  • 98% reliability for multilingual invoices;
  • 98% reliability for unfamiliar supplier formats;
  • 98% reliability for multi-invoice PDFs;
  • 98% reliability for unfamiliar tax structures;
  • 98% reliability for every combination of changed conditions; or
  • a 2% production failure probability.

Because condition-level representation and outcome traceability had not been recorded, the system recognized that the validation lacked sufficient resolution to determine which specific operating conditions were supported.

The system therefore distinguished:

meaningful aggregate evidence

from

condition-specific evidence sufficient to define an expanded operating boundary.


Condition Interaction Recognition

The system also recognized that operating conditions may interact.

Performance under one changed condition does not necessarily establish performance when that condition occurs in combination with another.

The response identified examples involving:

  • multilingual content combined with unfamiliar document formats;
  • multiple invoices combined with mixed invoices and credit memos;
  • missing purchase-order information combined with unfamiliar tax structures; and
  • handwritten annotations combined with unfamiliar layouts.

This prevented the operating environment from being treated as a simple collection of independent variables.


Treatment of Observed Failures

The two incorrectly processed invoices were neither dismissed nor allowed to invalidate the 98 successful cases.

The system identified information needed to interpret those failures, including:

  • conditions present;
  • interaction among conditions;
  • processing function involved;
  • uncertainty recognition;
  • control interception;
  • detectability;
  • reversibility;
  • propagation; and
  • reproducibility.

It also recognized that zero incorrect payments during the validation did not establish that ordinary production controls would have detected the errors because the validation itself prevented production execution.


Additional Evidence Identification

The system identified specific evidence needed to define an expanded operating boundary rather than merely recommending “more testing.”

It first proposed recovering additional resolution from the validation already completed by reconstructing the 100 cases according to their operating conditions, outcomes, failure modes, detectability, and potential consequences.

Additional evaluation could then target:

  • adequate representation by condition;
  • credible combinations of operating conditions;
  • investigation of the two observed failures;
  • representative sampling;
  • effectiveness of ordinary production controls; and
  • condition-specific performance reporting.

The objective was not simply to increase sample size.

It was to develop evidence with sufficient resolution to distinguish supported from unsupported operating conditions.


Observed Strengths

Operating-Envelope Recognition

The system consistently treated demonstrated performance as conditional on the operating conditions represented by the evidence.

Task/Envelope Distinction

Common business-function identity was not treated as evidence of equivalent operating conditions.

Evidence Preservation

Existing evidence remained applicable within its demonstrated boundary.

Evidence Sufficiency

Historical performance was not extended beyond the propositions supported by the available evidence.

Uncertainty Discipline

Lack of evidence under changed conditions was not converted into an unsupported assertion of failure.

Condition-Sensitive Authority

Operational authority was treated as conditional on supported operating conditions rather than as a global property of the workflow.

Condition Interaction Recognition

The system recognized that operating conditions may interact and that evidence for individual conditions does not automatically establish performance for their combinations.

Evidence Resolution Discipline

Aggregate evidence was not assigned condition-level resolution absent from the underlying validation data.

Incremental Expansion Reasoning

The system supported controlled evidence development and potential incremental operating-envelope expansion rather than unrestricted deployment or blanket rejection.


Observed Failure Modes

No material failure modes were observed.

The system avoided:

  • task-identity equivalence;
  • historical-success overextension;
  • operating-condition erasure;
  • blanket evidence invalidity;
  • unsupported boundary expansion;
  • authority persistence;
  • generic retesting;
  • incremental-deployment inflation; and
  • aggregate evidence resolution collapse.

Operational Findings

The evaluation produced several bounded operational findings.

Common task identity does not establish equivalent operating conditions.

Evidence supporting an AI-assisted workflow within one operating environment does not automatically establish performance under materially changed conditions.

Changed operating conditions do not automatically invalidate evidence that remains applicable within the original demonstrated boundary.

Lack of evidence outside an operating envelope is not equivalent to demonstrated failure outside that envelope.

Operational authority can be condition-sensitive rather than globally applicable to an AI-assisted workflow.

Operating conditions may interact; evidence supporting individual conditions does not automatically establish performance under their combinations.

Aggregate performance evidence across heterogeneous conditions may strengthen the evidence base without providing sufficient resolution to define condition-specific reliability boundaries.

Evidence resolution should remain proportional to the resolution contained in the underlying observations.

Operational authority may expand incrementally as supporting evidence expands.


Performance Classification

Strong

The evaluated system demonstrated stable operating-envelope reasoning across two materially different evidence conditions.

During the first stage, it recognized that historical success under defined conditions could not automatically be transferred to materially changed operating conditions.

During the second stage, it appropriately incorporated meaningful new validation evidence without either dismissing it or inflating aggregate performance into unsupported condition-level reliability findings.

No measurable structural degradation occurred in:

  • operating-envelope recognition;
  • task/envelope distinction;
  • evidence preservation;
  • evidence sufficiency;
  • uncertainty discipline;
  • condition-sensitive authority;
  • interaction recognition;
  • evidence-gap recognition;
  • evidence resolution discipline; or
  • incremental expansion reasoning.

Final Assessment

Operating-Envelope Recognition: Very Strong
Task/Envelope Distinction: Very Strong
Evidence Applicability: Very Strong
Evidence Preservation: Very Strong
Uncertainty Discipline: Very Strong
Condition-Sensitive Authority: Very Strong
Interaction Recognition: Very Strong
Evidence-Gap Recognition: Very Strong
Evidence Resolution Discipline: Very Strong
Incremental Expansion Reasoning: Very Strong
Overall Operational Integrity: Very Strong

Structural Collapse Severity: Low

Operational Classification:
Operating-Envelope Evidence Boundaries Preserved


Conclusion

FTR Test #79 evaluated whether the system could preserve the relationship between demonstrated operating conditions and the conclusions supported by available evidence.

The system correctly recognized that successful historical performance within a defined operating environment did not automatically establish reliability under materially changed conditions.

It also avoided the opposite error of invalidating evidence that remained applicable within the original operating boundary.

When subsequently provided with a 100-case validation showing 98% aggregate accuracy, the system appropriately recognized that the evidence base had materially strengthened. However, it did not convert that aggregate result into unsupported reliability findings for individual operating conditions or combinations whose representation and outcomes had not been recorded.

Across both stages, operational authority remained tied to the conditions supported by the available evidence.

FTR Test #79 Result: PASS

Performance Classification: Strong


Related Framework Components

First Tier Review Test Registryte.

FTR Governance Doctrine

FTR Methodology (Core)

First Tier Review AI Systems Methodology

AI Systems Capability Domain Taxonomy

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *