← All research
Research Program
Measuring AI Reliability
Scientific Question:
How reliable are AI systems in high-stakes domains?
This archive documents a benchmark attempt that did not produce protocol-valid scientific outputs. It is retained to make the execution failure and fail-closed response inspectable.
Projects
P Frozen Engineering Audit
MaxFitCalib-Bench
Execution-contract audit; zero interpretable scientific outputs; direction closed