← InteractionKit
Frozen methodological asset Study 2 · design frozen, materials not built

Can corrective interfaces be matched to how AI communication fails?

A designed human-AI decision study would test whether a corrective card improves appropriate reliance specifically when its information structure matches the AI answer's communication failure. The study is deferred and not an active priority.

Single research question

Does matching an evidence-based corrective interface to an AI answer's communication-failure family improve appropriate reliance relative to an equally formatted, truthful but failure-mismatched correction?

Why this is identifiable

Two failure families

Unsupported numerical precision and omitted decision boundaries are operationalized in parallel answer variants.

Two corrective cards

A numerical-warrant card and a boundary-condition card are crossed with both failure families.

Matched counterfactual

Every participant receives matched and mismatched corrections; mismatch is irrelevant to the failure, not false or lower quality.

Stimulus generalization

The planned design uses 24 independently grounded scenarios and models participants and scenarios as crossed sources of variation.

Frozen design

DesignWithin-participant 2 × 2 × 2 × 2, with evidence support varied across scenarios
Planned exposure16 trials per participant; 8 matched and 8 mismatched
Scenario pool24 exercise-and-health decisions; 12 per evidence-support level
Primary outcomeChange in probability assigned to the correct final decision
Secondary outcomeFinal decision accuracy
Provisional N240 analyzable participants; binding only after exact-schedule post-pilot simulation passes

What exists now

3
typed primitives
7
contract tests
0
human participants claimed

The publicly tagged v1.0.0 release implements typed interaction specifications, composition checks, schema generation, and structured logging. The Study 2 design, confounding audit, and simulation code exist locally. Study 2 scenarios, answer variants, corrective-card text, independent ratings, ethics approval, registration, and participant data do not yet exist.

Decision gates before recruitment

  1. Independently adjudicate 24 scenarios and their evidence support.
  2. Blind-code failure-family purity and remove mixed-failure answer variants.
  3. Pretest card relevance, leakage, length, credibility, and comprehension.
  4. Verify the exact allocation schedule, full-rank design matrix, logging, and recovery behavior.
  5. Pass the post-pilot simulation gate: at least 80% power with acceptable Type I error, coverage, and wrong-sign probability.
  6. Obtain institutional ethics approval and freeze the public preregistration before confirmatory recruitment.

Contribution boundary

The intended contribution is a failure-contingent interaction principle: uncertainty support should expose the information missing from a specific communication failure. It is not a claim that more explanation, provenance, or interface complexity generally increases trust.

Current evidence supports software conformance and design auditability only. Causal, construct-validity, cross-domain, and cross-lab claims remain untested.