Research artifact · Bounded advice intervention

CheckMyCoach

Can a system detect when AI advice exceeds its evidential support — and repair the response without overstating certainty? This page walks fixed illustrative fixtures through the four-module pipeline — routing, rule-based tags, candidate revisions, deterministic surface checks — as a deterministic, inspectable walkthrough.

Active · offline-evaluated prototype · engineering dry-run validated · no human validation · fixed illustrative fixtures

Research question

Can a system detect when AI advice exceeds its evidential support — and repair the response without overstating certainty?

Fitness advice from a generative AI can be confident and specific while resting on thin evidence. The question is operational, and deliberately ordered before any human study: route selected cases, tag the failure category, generate a candidate revision, and verify the delivered output with deterministic checks — keeping routing coverage, revision execution, and check acceptance separately inspectable.

Artifact · pipeline walkthrough

What each stage changes for a fixed case

Select a case. All stage states render instantly from the fixed fixture — nothing executes at inspection time, and the same case always produces the same states. Cases that are not routed never reach candidate generation or surface checks.

Fixed illustrative cases

Custom questions always inspect the same deterministic fallback fixture; the input is never sent anywhere.

M1 · Routing

Selects which cases enter the pipeline under the heuristic router.

Routed

An over-precise comparison in the fixture response exceeds its evidential support.

M2 · Rule tag

Assigns a rule-based tag naming the failure category.

Context mismatch

Tag recorded for the routed case.

M3 · Candidate revision

Generates a candidate revision for routed cases.

Candidate generated

What changed

  • Removed the over-precise calorie comparison ("3x more calories")
  • Reframed the answer around caloric deficit, the primary factor
  • Added a sustainability caveat instead of a single-modality recommendation

M4 · Surface checks

Runs deterministic target-and-retention checks on the delivered output.

Checks rendered (illustrative)
  • Target-feature removal: Over-precise claim removed
  • Information retention: Supported framing and caveat retained

Delivered output (fixture)

Running burns 3x more calories than walking. Fat loss depends primarily on caloric deficit, not exercise modality. Running burns more calories per minute (8.2 vs 5.4 kcal/min), but the best exercise is the one you can sustain consistently. Walking for 60 minutes may produce better long-term adherence than running for 20.

supported claim (fixture role) uncertainty / caveat retained over-precise claim removed

Fixed illustrative fixture authored for this walkthrough — not live model output. Role coloring marks fixture categories, not measured correctness, safety, or trust.

Fixture source material

  • ACSM Position Stand: Fat loss depends primarily on caloric deficit, not exercise modality.
  • Meta-analysis (Willis et al., 2016): Running produces slightly greater total energy expenditure per session (8.2 vs 5.4 kcal/min), but walking allows longer sessions with less fatigue.
Evaluation context

Where the real numbers live

The evaluation behind this prototype ran on a fixed constructed corpus (evaluation v2.2), not on this walkthrough: 15 of 40 constructed cases were routed, 12 of 40 delivered outputs removed the predefined target feature, and 0 of 40 passed every target-and-retention check in the single recorded execution.

The fixture check outcomes above are illustrative states of the pipeline, not those evaluation results. Full evaluation record →

Scientific boundary

What this walkthrough shows — and what it does not establish

What this walkthrough shows

  • A deterministic walkthrough of the four-module implementation surface on fixed illustrative fixtures.
  • How routing, rule tags, candidate revisions, and surface-check states change for a fixed case.
  • The evaluation design that separates routing coverage, revision execution, and check acceptance before any human study.

What this walkthrough does not establish

  • Human trust effects.
  • Human reliance or calibration effects.
  • Clinical or fitness-advice validity.
  • Real-world safety effectiveness.
  • Validated psychometric or trust measurement.
  • Model-level benchmark performance.
Source & reproducibility

Where the evidence lives

The evaluation code, the constructed-corpus manifest, and the recorded v2.2 execution live in the source repository. This page ships no evaluation code, runs no model, and makes no network requests — every state on this page is a fixed fixture rendered at build time.

Reading rule: an interactive artifact, an offline evaluation on a constructed corpus, and a human-subject study are different evidence classes. This artifact has the first two; it does not yet have the third.