Evaluation Runtime
When an AI-agent evaluation crashes mid-execution, what happened? Most harnesses answer by overwriting: the retry replaces the failed attempt, a missing success record becomes permission to repeat an external action, and one blended rate hides what was actually measured. This artifact is a deliberately small runtime built so that logical work, execution attempts, ambiguity, and report denominators stay separately inspectable — under a machine-readable reliability contract.
Active · prototype · public repository · updated 2026-09-17
Can an evaluation runtime record failure without inventing certainty — and refuse to repeat an external action it cannot account for?
The question is operational, not scientific: given one append-only execution history, can the runtime distinguish a logical task from its attempts, preserve ambiguity instead of resolving it by guessing, bind judgments to the specification that was actually executed against, and keep every report denominator explicit?
What the runtime makes inspectable
Each scenario is a static, deterministic projection of documented, test-pinned semantics — Published public repository (fresh-root history). Every projection below is illustrative: it transcribes documented, test-pinned semantics of the runtime into a static view. No new experimental results were generated.. Select a scenario.
F1 · Task vs attempt identity
Failure mode. Conflating a logical task with its executions. If a task_id were reused as the identity of an attempt, expressing "ran it again" would require overwriting history — the first attempt's failure record would be destroyed.
What happens
- one logical task_id
- attempt 1 → new run_id, appended
- attempt 2 → new run_id, appended
- attempt 3 → new run_id, appended
- attempt ordinals are derived from durable run_started events, never from an in-memory counter
| Attempt | Run identity | Outcome |
|---|---|---|
| 1 | run-1a2b (run_id A) | failed — stored as a fact, never erased |
| 2 | run-9f3c (run_id B) | failed — a separate appended record |
| 3 | run-77d0 (run_id C) | success — committed as a third record |
What the runtime preserves
All three records remain in the append-only event history. The eventual success does not overwrite the failures; per-attempt status and task-level outcome stay separable.
Executable evidence
Claim boundary
This does not demonstrate exactly-once execution, deduplication, or any real-provider behavior. The runtime deliberately rejects duplicate attempt identities instead of deduplicating them.
F2 · Ambiguous external execution
Failure mode. A crash between dispatching external work and committing its outcome. If recovery treats a missing success record as permission to repeat the action, an already-executed (possibly billed) invocation is silently duplicated.
What happens
- run_started persisted
- provider_invocation_started persisted immediately before the call
- process interrupted before outcome/terminal could be committed
- fold projects: ambiguous_external_execution
- resume returns AMBIGUOUS — no provider invocation
| Durable event prefix | Folded execution_state | Automatic restart action |
|---|---|---|
| No run_started | no Run exists | FRESH |
| Start, no invocation marker | known_not_executed | RETRY if attempt budget remains |
| Invocation marker, no outcome | ambiguous_external_execution | AMBIGUOUS — no provider call |
| Typed retryable failure outcome | known_failed | RETRY if attempt budget remains |
| Untyped failure outcome, no terminal | ambiguous_external_execution | AMBIGUOUS — no provider call |
| Success outcome, no terminal commit | ambiguous_external_execution | AMBIGUOUS — no provider call |
| Committed success terminal | known_executed_and_committed | SKIP |
What the runtime preserves
Uncertainty itself. A durable invocation without a committed successful terminal is never automatically re-invoked unless durable evidence proves a retryable failure. Ambiguity is preserved across restarts rather than resolved by guessing.
Executable evidence
Claim boundary
This demonstrates preservation and fail-closed recovery, not resolution. No reconciliation mechanism, no external-repair API, and no proof that the ambiguous invocation did not execute externally exists (RC-14: NOT_SUPPORTED).
F3 · Evaluator version binding
Failure mode. Re-judging a stored output against a silently changed specification. If the mutable current task definition were the scoring reference, a later spec edit would rewrite the meaning of every historical judgment.
What happens
- run executes with expected = "Paris"; the run persists expected_snapshot = "Paris"
- output returned: " paris " (edge whitespace, different case)
- current task fixture is later edited to expected = "London" (drift)
- v1 (exact) judges the run: FAIL — against the snapshot, never "London"
- v2 (normalized) judges the same run: PASS — same snapshot reference
| Scoring reference | v1 exact | v2 normalized-text |
|---|---|---|
| expected_snapshot = "Paris" (captured at execution time) | FAIL — " paris " ≠ "Paris" | PASS — "paris" = "paris" |
| current tasks.jsonl = "London" (drifted) | never used as scoring reference | never used as scoring reference |
What the runtime preserves
Judgment identity is the pair (run_id, evaluator_version). Each version's verdict is a separate durable record; execution metrics are invariant across evaluator versions; a run without a captured snapshot is not eligible for new evaluation at all (no current-task fallback).
Executable evidence
Claim boundary
Version binding establishes which semantics produced a judgment. It does not establish that any evaluator version measures output quality validly, nor anything about model quality.
F4 · Denominator discipline
Failure mode. Collapsing different questions into one rate. Coverage, per-attempt reliability, recovery, and output quality have different denominators; merging them turns "nothing was evaluated" into a fake 0.0 and hides quarantined records inside an average.
What happens
- four report signals, four explicit denominators
- raw vs valid vs quarantined counts are all exposed, never hidden
- an undefined ratio is emitted as null, not 0.0
- structural contradictions abort the report entirely (no numbers at all)
- referential violations and ambiguous duplicate judgments are quarantined and listed
| Signal | Numerator | Denominator | Undefined when |
|---|---|---|---|
| Task execution coverage | unique attempted task_ids | defined tasks | no tasks defined → null |
| Attempt execution reliability | successful valid runs | all valid runs (incl. interrupted) | no valid runs → null |
| Recovery-aware task execution | attempted tasks with ≥1 success | unique attempted tasks | nothing attempted → null |
| Conditional output quality | passed evaluations | completed evaluations only | no completed evaluations → null |
What the runtime preserves
Quarantined observations stay visible: an evaluation referencing a nonexistent or failed run, or an ambiguous duplicate (run_id, evaluator_version) group, is excluded from every denominator and listed in the report's integrity block. Failed evaluator runs appear as raw counts, never as silent zeros.
Executable evidence
Claim boundary
No end-to-end task pass rate exists (deliberately deferred), and no metric here says anything about model quality or about the scientific validity of any evaluation.
Reliability contract classification
Every guarantee the runtime makes is a claim record with an explicit status and named evidence tests. The totals below are contract classifications, not a quality score: 7 of 15 is not a rating, and the NOT_SUPPORTED rows are as much a part of the artifact as the SUPPORTED ones.
Source: src/reliability_contract.py in the public repository · statuses unchanged from the prepared candidate (7 / 4 / 3 / 1).
What this artifact demonstrates — and what it does not
Demonstrated here
- Logical task identity and execution-attempt identity are distinct, durable, and test-pinned.
- Retries append new history; committed results are immutable across resume and replay.
- Ambiguous external execution is preserved and blocks automatic re-invocation (fail closed).
- Evaluations bind (run_id, evaluator_version) and judge only the execution-time snapshot.
- Report denominators are explicit; undefined ratios are null; invalid records are quarantined, not averaged in.
- A machine-readable reliability contract classifies every claim with named evidence tests.
Not established
- No real LLM/provider integration — the agent is a deterministic mock plus injected faults.
- Single process, single writer; no database, queue, or web service.
- No exactly-once provider execution and no event-idempotent append (RC-10/RC-14: NOT_SUPPORTED).
- Not production-grade infrastructure: append+flush is not fsync, transactional, or concurrent-writer safe.
- Not a model-quality benchmark and not an evaluation-methodology validation.
- Malformed returned output can be committed as execution success (no output schema validation).
Stated limits, by claim id
These are contract records in the snapshot, not disclaimers added for this page.
| Claim | Status | What this means |
|---|---|---|
| RC-10 · Event replay idempotence | NOT_SUPPORTED | Duplicate events fail loudly; they are not deduplicated. No idempotency key or exactly-once append. |
| RC-11 · Failure replay determinism | NOT_SUPPORTED | Replay accepts successful snapshot-bearing runs only; failed runs are not replayable. |
| RC-14 · Exactly-once provider execution | NOT_SUPPORTED | Uncertainty is preserved and automatic replay blocked, but external execution cannot be proven either way without reconciliation. |
| RC-12 · Cost / token budget safety | NOT_APPLICABLE | No cost or token budget subsystem exists; the only budget is the request-attempt budget (RC-02). |
RC-10 · Event replay idempotence
Status NOT_SUPPORTED
Meaning Duplicate events fail loudly; they are not deduplicated. No idempotency key or exactly-once append.
RC-11 · Failure replay determinism
Status NOT_SUPPORTED
Meaning Replay accepts successful snapshot-bearing runs only; failed runs are not replayable.
RC-14 · Exactly-once provider execution
Status NOT_SUPPORTED
Meaning Uncertainty is preserved and automatic replay blocked, but external execution cannot be proven either way without reconciliation.
RC-12 · Cost / token budget safety
Status NOT_APPLICABLE
Meaning No cost or token budget subsystem exists; the only budget is the request-attempt budget (RC-02).
Where the evidence lives
The runtime is a standard-library-only Python package whose semantics are pinned by an executable test suite; the prepared snapshot was verified green before this page was built. It is intentionally small enough to read in one sitting.
The runtime is published as a public repository with fresh-root history, so the source and test references below link into it directly.
Reading rule: an interactive artifact, an executable test suite, and a validated scientific result are different evidence classes. This page contains only the first two.