← All research

Measuring AI Reliability

Scientific Question:

How reliable are AI systems in high-stakes domains?

This archive documents a benchmark attempt that did not produce protocol-valid scientific outputs. It is retained to make the execution failure and fail-closed response inspectable.

Projects

P Frozen Engineering Audit

MaxFitCalib-Bench

Execution-contract audit; zero interpretable scientific outputs; direction closed