Research artifact · Measuring AI reliability

Evaluation Runtime

When an AI-agent evaluation crashes mid-execution, what happened? Most harnesses answer by overwriting: the retry replaces the failed attempt, a missing success record becomes permission to repeat an external action, and one blended rate hides what was actually measured. This artifact is a deliberately small runtime built so that logical work, execution attempts, ambiguity, and report denominators stay separately inspectable — under a machine-readable reliability contract.

Active · prototype · public repository · updated 2026-09-17

Research question

Can an evaluation runtime record failure without inventing certainty — and refuse to repeat an external action it cannot account for?

The question is operational, not scientific: given one append-only execution history, can the runtime distinguish a logical task from its attempts, preserve ambiguity instead of resolving it by guessing, bind judgments to the specification that was actually executed against, and keep every report denominator explicit?

Artifact · four scenarios

What the runtime makes inspectable

Each scenario is a static, deterministic projection of documented, test-pinned semantics — Published public repository (fresh-root history). Every projection below is illustrative: it transcribes documented, test-pinned semantics of the runtime into a static view. No new experimental results were generated.. Select a scenario.

F1 · Task vs attempt identity

Failure mode. Conflating a logical task with its executions. If a task_id were reused as the identity of an attempt, expressing "ran it again" would require overwriting history — the first attempt's failure record would be destroyed.

What happens

  1. one logical task_id
  2. attempt 1 → new run_id, appended
  3. attempt 2 → new run_id, appended
  4. attempt 3 → new run_id, appended
  5. attempt ordinals are derived from durable run_started events, never from an in-memory counter
Illustrative projection of retry semantics: one task_id, three attempts, three distinct run_id values
AttemptRun identityOutcome
1run-1a2b (run_id A)failed — stored as a fact, never erased
2run-9f3c (run_id B)failed — a separate appended record
3run-77d0 (run_id C)success — committed as a third record

What the runtime preserves

All three records remain in the append-only event history. The eventual success does not overwrite the failures; per-attempt status and task-level outcome stay separable.

Claim boundary

This does not demonstrate exactly-once execution, deduplication, or any real-provider behavior. The runtime deliberately rejects duplicate attempt identities instead of deduplicating them.

F2 · Ambiguous external execution

Failure mode. A crash between dispatching external work and committing its outcome. If recovery treats a missing success record as permission to repeat the action, an already-executed (possibly billed) invocation is silently duplicated.

What happens

  1. run_started persisted
  2. provider_invocation_started persisted immediately before the call
  3. process interrupted before outcome/terminal could be committed
  4. fold projects: ambiguous_external_execution
  5. resume returns AMBIGUOUS — no provider invocation
Durable execution-state projection (documented contract; fold and resume decisions)
Durable event prefixFolded execution_stateAutomatic restart action
No run_startedno Run existsFRESH
Start, no invocation markerknown_not_executedRETRY if attempt budget remains
Invocation marker, no outcomeambiguous_external_executionAMBIGUOUS — no provider call
Typed retryable failure outcomeknown_failedRETRY if attempt budget remains
Untyped failure outcome, no terminalambiguous_external_executionAMBIGUOUS — no provider call
Success outcome, no terminal commitambiguous_external_executionAMBIGUOUS — no provider call
Committed success terminalknown_executed_and_committedSKIP

What the runtime preserves

Uncertainty itself. A durable invocation without a committed successful terminal is never automatically re-invoked unless durable evidence proves a retryable failure. Ambiguity is preserved across restarts rather than resolved by guessing.

Claim boundary

This demonstrates preservation and fail-closed recovery, not resolution. No reconciliation mechanism, no external-repair API, and no proof that the ambiguous invocation did not execute externally exists (RC-14: NOT_SUPPORTED).

F3 · Evaluator version binding

Failure mode. Re-judging a stored output against a silently changed specification. If the mutable current task definition were the scoring reference, a later spec edit would rewrite the meaning of every historical judgment.

What happens

  1. run executes with expected = "Paris"; the run persists expected_snapshot = "Paris"
  2. output returned: " paris " (edge whitespace, different case)
  3. current task fixture is later edited to expected = "London" (drift)
  4. v1 (exact) judges the run: FAIL — against the snapshot, never "London"
  5. v2 (normalized) judges the same run: PASS — same snapshot reference
Illustrative projection: one stored run, two evaluator versions, one immutable reference
Scoring referencev1 exactv2 normalized-text
expected_snapshot = "Paris" (captured at execution time)FAIL — " paris " ≠ "Paris"PASS — "paris" = "paris"
current tasks.jsonl = "London" (drifted)never used as scoring referencenever used as scoring reference

What the runtime preserves

Judgment identity is the pair (run_id, evaluator_version). Each version's verdict is a separate durable record; execution metrics are invariant across evaluator versions; a run without a captured snapshot is not eligible for new evaluation at all (no current-task fallback).

Claim boundary

Version binding establishes which semantics produced a judgment. It does not establish that any evaluator version measures output quality validly, nor anything about model quality.

F4 · Denominator discipline

Failure mode. Collapsing different questions into one rate. Coverage, per-attempt reliability, recovery, and output quality have different denominators; merging them turns "nothing was evaluated" into a fake 0.0 and hides quarantined records inside an average.

What happens

  1. four report signals, four explicit denominators
  2. raw vs valid vs quarantined counts are all exposed, never hidden
  3. an undefined ratio is emitted as null, not 0.0
  4. structural contradictions abort the report entirely (no numbers at all)
  5. referential violations and ambiguous duplicate judgments are quarantined and listed
The four report signals (Signal Semantics v0) — each answers a different question
SignalNumeratorDenominatorUndefined when
Task execution coverageunique attempted task_idsdefined tasksno tasks defined → null
Attempt execution reliabilitysuccessful valid runsall valid runs (incl. interrupted)no valid runs → null
Recovery-aware task executionattempted tasks with ≥1 successunique attempted tasksnothing attempted → null
Conditional output qualitypassed evaluationscompleted evaluations onlyno completed evaluations → null

What the runtime preserves

Quarantined observations stay visible: an evaluation referencing a nonexistent or failed run, or an ambiguous duplicate (run_id, evaluator_version) group, is excluded from every denominator and listed in the report's integrity block. Failed evaluator runs appear as raw counts, never as silent zeros.

Claim boundary

No end-to-end task pass rate exists (deliberately deferred), and no metric here says anything about model quality or about the scientific validity of any evaluation.

Evidence

Reliability contract classification

Every guarantee the runtime makes is a claim record with an explicit status and named evidence tests. The totals below are contract classifications, not a quality score: 7 of 15 is not a rating, and the NOT_SUPPORTED rows are as much a part of the artifact as the SUPPORTED ones.

7
Supported
RC-01 · RC-02 · RC-03 · RC-04 · RC-05 · RC-13 · RC-15
4
Partially supported
RC-06 · RC-07 · RC-08 · RC-09
3
Not supported
RC-10 · RC-11 · RC-14
1
Not applicable
RC-12

Source: src/reliability_contract.py in the public repository · statuses unchanged from the prepared candidate (7 / 4 / 3 / 1).

Scientific boundary

What this artifact demonstrates — and what it does not

Demonstrated here

  • Logical task identity and execution-attempt identity are distinct, durable, and test-pinned.
  • Retries append new history; committed results are immutable across resume and replay.
  • Ambiguous external execution is preserved and blocks automatic re-invocation (fail closed).
  • Evaluations bind (run_id, evaluator_version) and judge only the execution-time snapshot.
  • Report denominators are explicit; undefined ratios are null; invalid records are quarantined, not averaged in.
  • A machine-readable reliability contract classifies every claim with named evidence tests.

Not established

  • No real LLM/provider integration — the agent is a deterministic mock plus injected faults.
  • Single process, single writer; no database, queue, or web service.
  • No exactly-once provider execution and no event-idempotent append (RC-10/RC-14: NOT_SUPPORTED).
  • Not production-grade infrastructure: append+flush is not fsync, transactional, or concurrent-writer safe.
  • Not a model-quality benchmark and not an evaluation-methodology validation.
  • Malformed returned output can be committed as execution success (no output schema validation).
Non-guarantees

Stated limits, by claim id

These are contract records in the snapshot, not disclaimers added for this page.

RC-10 · Event replay idempotence

Status NOT_SUPPORTED

Meaning Duplicate events fail loudly; they are not deduplicated. No idempotency key or exactly-once append.

RC-11 · Failure replay determinism

Status NOT_SUPPORTED

Meaning Replay accepts successful snapshot-bearing runs only; failed runs are not replayable.

RC-14 · Exactly-once provider execution

Status NOT_SUPPORTED

Meaning Uncertainty is preserved and automatic replay blocked, but external execution cannot be proven either way without reconciliation.

RC-12 · Cost / token budget safety

Status NOT_APPLICABLE

Meaning No cost or token budget subsystem exists; the only budget is the request-attempt budget (RC-02).

Source & reproducibility

Where the evidence lives

The runtime is a standard-library-only Python package whose semantics are pinned by an executable test suite; the prepared snapshot was verified green before this page was built. It is intentionally small enough to read in one sitting.

The runtime is published as a public repository with fresh-root history, so the source and test references below link into it directly.

Reading rule: an interactive artifact, an executable test suite, and a validated scientific result are different evidence classes. This page contains only the first two.