Primary research system

CheckMyCoach

A bounded intervention prototype that tests whether evidential-limit failures in AI advice can be operationally detected, routed, and repaired without overstating certainty.

Active · offline-evaluated · updated 2026-07-31

Research question

Can a system detect when AI advice exceeds its evidential support — and repair the response without overstating certainty?

Fitness advice from a generative AI can be confident and specific while resting on thin evidence. The question here is operational: can a bounded system detect when an answer exceeds its support, route that case for repair, and produce a revision that keeps useful content while dropping unsupported certainty?

Why this exists

Testing detection, routing, and repair of evidential-limit failures

CheckMyCoach is a bounded intervention prototype. It tests whether evidential-limit failures in AI advice — the kind identified in earlier measurement and audit work — can be detected, routed, and repaired in a deterministic, inspectable way before any human study is run.

System

The four-module pipeline

CheckMyCoach routes advice cases, assigns rule-based tags, generates candidate revisions, and verifies them with deterministic surface checks.

M1

Routing

Selects cases for candidate generation under the historical heuristic.

M2

Rule-based tagging

Assigns rule-based tags to routed cases.

M3

Candidate generation

Returns a candidate revision for every routed case.

M4

Surface checks

Runs deterministic target-and-retention checks on delivered outputs.

Evidence

What the current evaluation establishes

One fixed 40-case constructed corpus, one execution. The table separates documented execution facts from what they do not establish.

Fixed-corpus execution — descriptive facts, not validated outcomes

40 constructed cases

What it is A fixed, constructed fitness-advice corpus.

What it is not Not a real-world dataset.

15 / 40 routed

What it is Cases routed to candidate generation under the heuristic (M1).

What it is not Routing coverage is not detection accuracy.

12 / 40 removed target feature

What it is Delivered outputs that removed the predefined target feature (M4).

What it is not Removal is not validated detection or repair effectiveness.

0 / 40 passed every check

What it is No delivered output passed every target-and-retention check.

What it is not No validated effectiveness is claimed.

Generic = Generic+Oracle

What it is Identical per-case decisions in this single execution.

What it is not Not an equivalence result; oracle value remains unassessed.

Scientific boundary

What this supports — and what it does not yet establish

Supports

  • Four-module decomposition executes end-to-end on a fixed constructed corpus.
  • Routing, removal, acceptance, and end-to-end denominators are documented for one run.
  • A controlled way to separate routing coverage, revision execution, and surface-check acceptance before testing human effects.

Does not yet establish

  • Validated detection, superiority, equivalence, or user-effect claims.
  • Human benefit or cross-domain generalization.
  • Any human-participant finding.
Next step

Human validation: the current empirical step

The active empirical step is human validation of CheckMyCoach's target-removal and retention judgments.

Human validation is not completed, and no human-participant findings are claimed. The standalone CheckMyCoach system note is archived and remains a reusable writing asset; it is not an active manuscript submission or peer-reviewed publication.

Place in the program

Research lineage

Artifacts

Resources