Frozen Engineering Test Only Closed: 2026-07-24

MaxFitCalib-Bench Controlled Audit

A forensic record of an execution-contract failure and a stricter preflight abort. This artifact produced no interpretable scientific outputs and is not an active publication project.

Claim boundary

ENGINEERING TEST ONLY — SCIENTIFIC DIRECTION CLOSED. The audit does not estimate prevalence, compare models, validate a construct, or support claims about users. Earlier exploratory percentages are excluded because they are not protocol-valid scientific results.

72
v0.1 Attempts
0
Parser-valid Outputs
2
v0.2 Preflights
0
Scientific Calls

What happened

  1. v0.1: 72 attempts yielded zero parser-valid scientific outputs, so downstream numbers were not interpretable.
  2. v0.2: two preflight executions detected the unresolved contract problem and aborted before scientific API calls.
  3. Closure: the evaluator and static construct remain inspectable engineering artifacts; further scientific execution is out of scope.

What remains useful

  • A documented example of why execution contracts must be validated before expensive model calls.
  • A fail-closed preflight pattern that prevents invalid runs from being presented as scientific evidence.
  • A transparent negative result: zero scientific outputs, rather than reconstructed or selectively retained metrics.
← Back to measuring AI reliability