CheckMyCoach
A bounded intervention prototype that tests whether evidential-limit failures in AI advice can be operationally detected, routed, and repaired without overstating certainty.
Active · offline-evaluated · updated 2026-07-31
Can a system detect when AI advice exceeds its evidential support — and repair the response without overstating certainty?
Fitness advice from a generative AI can be confident and specific while resting on thin evidence. The question here is operational: can a bounded system detect when an answer exceeds its support, route that case for repair, and produce a revision that keeps useful content while dropping unsupported certainty?
Testing detection, routing, and repair of evidential-limit failures
CheckMyCoach is a bounded intervention prototype. It tests whether evidential-limit failures in AI advice — the kind identified in earlier measurement and audit work — can be detected, routed, and repaired in a deterministic, inspectable way before any human study is run.
The four-module pipeline
CheckMyCoach routes advice cases, assigns rule-based tags, generates candidate revisions, and verifies them with deterministic surface checks.
Routing
Selects cases for candidate generation under the historical heuristic.
Rule-based tagging
Assigns rule-based tags to routed cases.
Candidate generation
Returns a candidate revision for every routed case.
Surface checks
Runs deterministic target-and-retention checks on delivered outputs.
What the current evaluation establishes
One fixed 40-case constructed corpus, one execution. The table separates documented execution facts from what they do not establish.
| Fact | What it is | What it is not |
|---|---|---|
| 40 constructed cases | A fixed, constructed fitness-advice corpus. | Not a real-world dataset. |
| 15 / 40 routed | Cases routed to candidate generation under the heuristic (M1). | Routing coverage is not detection accuracy. |
| 12 / 40 removed target feature | Delivered outputs that removed the predefined target feature (M4). | Removal is not validated detection or repair effectiveness. |
| 0 / 40 passed every check | No delivered output passed every target-and-retention check. | No validated effectiveness is claimed. |
| Generic = Generic+Oracle | Identical per-case decisions in this single execution. | Not an equivalence result; oracle value remains unassessed. |
Fixed-corpus execution — descriptive facts, not validated outcomes
40 constructed cases
What it is A fixed, constructed fitness-advice corpus.
What it is not Not a real-world dataset.
15 / 40 routed
What it is Cases routed to candidate generation under the heuristic (M1).
What it is not Routing coverage is not detection accuracy.
12 / 40 removed target feature
What it is Delivered outputs that removed the predefined target feature (M4).
What it is not Removal is not validated detection or repair effectiveness.
0 / 40 passed every check
What it is No delivered output passed every target-and-retention check.
What it is not No validated effectiveness is claimed.
Generic = Generic+Oracle
What it is Identical per-case decisions in this single execution.
What it is not Not an equivalence result; oracle value remains unassessed.
What this supports — and what it does not yet establish
Supports
- Four-module decomposition executes end-to-end on a fixed constructed corpus.
- Routing, removal, acceptance, and end-to-end denominators are documented for one run.
- A controlled way to separate routing coverage, revision execution, and surface-check acceptance before testing human effects.
Does not yet establish
- Validated detection, superiority, equivalence, or user-effect claims.
- Human benefit or cross-domain generalization.
- Any human-participant finding.
Human validation: the current empirical step
The active empirical step is human validation of CheckMyCoach's target-removal and retention judgments.
Human validation is not completed, and no human-participant findings are claimed. The standalone CheckMyCoach system note is archived and remains a reusable writing asset; it is not an active manuscript submission or peer-reviewed publication.