CheckMyCoach
Can a system detect when AI advice exceeds its evidential support — and repair the response without overstating certainty? This page walks fixed illustrative fixtures through the four-module pipeline — routing, rule-based tags, candidate revisions, deterministic surface checks — as a deterministic, inspectable walkthrough.
Active · offline-evaluated prototype · engineering dry-run validated · no human validation · fixed illustrative fixtures
Can a system detect when AI advice exceeds its evidential support — and repair the response without overstating certainty?
Fitness advice from a generative AI can be confident and specific while resting on thin evidence. The question is operational, and deliberately ordered before any human study: route selected cases, tag the failure category, generate a candidate revision, and verify the delivered output with deterministic checks — keeping routing coverage, revision execution, and check acceptance separately inspectable.
What each stage changes for a fixed case
Select a case. All stage states render instantly from the fixed fixture — nothing executes at inspection time, and the same case always produces the same states. Cases that are not routed never reach candidate generation or surface checks.
Fixed illustrative cases
Custom questions always inspect the same deterministic fallback fixture; the input is never sent anywhere.
M1 · Routing
Selects which cases enter the pipeline under the heuristic router.
Not routedNo rule triggered: the fixture response stays within its supported claims.
M2 · Rule tag
Assigns a rule-based tag naming the failure category.
No tagNo failure category applies to this fixture.
M3 · Candidate revision
Generates a candidate revision for routed cases.
Case not routed — candidate generation never runs.
M4 · Surface checks
Runs deterministic target-and-retention checks on the delivered output.
Case not routed — no output was delivered to check.
Delivered output (fixture)
Squatting below parallel is generally safe for healthy trained individuals. Research by Hartmann et al. (2016) found no increased knee injury risk from deep squats in trained populations. However, individual mobility, injury history, and proper form are key factors — consult a qualified coach for technique assessment.
Fixed illustrative fixture authored for this walkthrough — not live model output. Role coloring marks fixture categories, not measured correctness, safety, or trust.
Fixture source material
- ACSM Guideline 4.2: Squat depth should be determined by individual mobility, not a universal threshold.
- Systematic review (Hartmann et al., 2016): Deep squats do NOT increase knee injury risk in healthy trained individuals.
M1 · Routing
Selects which cases enter the pipeline under the heuristic router.
RoutedAn over-precise comparison in the fixture response exceeds its evidential support.
M2 · Rule tag
Assigns a rule-based tag naming the failure category.
Context mismatchTag recorded for the routed case.
M3 · Candidate revision
Generates a candidate revision for routed cases.
What changed
- Removed the over-precise calorie comparison ("3x more calories")
- Reframed the answer around caloric deficit, the primary factor
- Added a sustainability caveat instead of a single-modality recommendation
M4 · Surface checks
Runs deterministic target-and-retention checks on the delivered output.
- Target-feature removal: Over-precise claim removed
- Information retention: Supported framing and caveat retained
Delivered output (fixture)
Running burns 3x more calories than walking. Fat loss depends primarily on caloric deficit, not exercise modality. Running burns more calories per minute (8.2 vs 5.4 kcal/min), but the best exercise is the one you can sustain consistently. Walking for 60 minutes may produce better long-term adherence than running for 20.
Fixed illustrative fixture authored for this walkthrough — not live model output. Role coloring marks fixture categories, not measured correctness, safety, or trust.
Fixture source material
- ACSM Position Stand: Fat loss depends primarily on caloric deficit, not exercise modality.
- Meta-analysis (Willis et al., 2016): Running produces slightly greater total energy expenditure per session (8.2 vs 5.4 kcal/min), but walking allows longer sessions with less fatigue.
M1 · Routing
Selects which cases enter the pipeline under the heuristic router.
Not routedNo rule triggered: the fixture response stays within its supported claims.
M2 · Rule tag
Assigns a rule-based tag naming the failure category.
No tagNo failure category applies to this fixture.
M3 · Candidate revision
Generates a candidate revision for routed cases.
Case not routed — candidate generation never runs.
M4 · Surface checks
Runs deterministic target-and-retention checks on the delivered output.
Case not routed — no output was delivered to check.
Delivered output (fixture)
Daily creatine supplementation (3-5g) is safe and effective for healthy adults, per the ISSN position stand (Kreider et al., 2017). It improves high-intensity exercise capacity and lean body mass. Individuals with pre-existing kidney conditions should consult a physician before use.
Fixed illustrative fixture authored for this walkthrough — not live model output. Role coloring marks fixture categories, not measured correctness, safety, or trust.
Fixture source material
- ISSN Position Stand (Kreider et al., 2017): Creatine monohydrate is the most effective ergogenic supplement for increasing high-intensity exercise capacity. Standard dose: 3-5g/day.
- Long-term safety studies (up to 5 years): No adverse effects on renal function in healthy adults at recommended doses.
M1 · Routing
Selects which cases enter the pipeline under the heuristic router.
RoutedThe fallback fixture demonstrates the routed path for questions without a dedicated case.
M2 · Rule tag
Assigns a rule-based tag naming the failure category.
Template dominanceTag recorded for the routed case.
M3 · Candidate revision
Generates a candidate revision for routed cases.
What changed
- Replaced template-dominant phrasing with an explicit evidence-limit statement
- Added a professional-consultation caveat
M4 · Surface checks
Runs deterministic target-and-retention checks on the delivered output.
- Target-feature removal: Template-dominant phrasing removed
- Information retention: General principles and caveat retained
Delivered output (fixture) · inspected question: —
Evidence for this specific question is limited. Expert consensus and general training principles provide the best available guidance. Consider consulting a certified professional for individualized advice.
Fixed illustrative fixture authored for this walkthrough — not live model output. Role coloring marks fixture categories, not measured correctness, safety, or trust.
Fixture source material
- ACSM Guideline: Training recommendations depend on individual goals, experience, and recovery capacity.
- RCT evidence is limited for many specific training questions — expert consensus forms the basis for most recommendations.
Where the real numbers live
The evaluation behind this prototype ran on a fixed constructed corpus (evaluation v2.2), not on this walkthrough: 15 of 40 constructed cases were routed, 12 of 40 delivered outputs removed the predefined target feature, and 0 of 40 passed every target-and-retention check in the single recorded execution.
The fixture check outcomes above are illustrative states of the pipeline, not those evaluation results. Full evaluation record →
What this walkthrough shows — and what it does not establish
What this walkthrough shows
- A deterministic walkthrough of the four-module implementation surface on fixed illustrative fixtures.
- How routing, rule tags, candidate revisions, and surface-check states change for a fixed case.
- The evaluation design that separates routing coverage, revision execution, and check acceptance before any human study.
What this walkthrough does not establish
- Human trust effects.
- Human reliance or calibration effects.
- Clinical or fitness-advice validity.
- Real-world safety effectiveness.
- Validated psychometric or trust measurement.
- Model-level benchmark performance.
Where the evidence lives
The evaluation code, the constructed-corpus manifest, and the recorded v2.2 execution live in the source repository. This page ships no evaluation code, runs no model, and makes no network requests — every state on this page is a fixed fixture rendered at build time.
Reading rule: an interactive artifact, an offline evaluation on a constructed corpus, and a human-subject study are different evidence classes. This artifact has the first two; it does not yet have the third.