Human-centered AI · Decision support · Evaluation

Baixin Guo

Max · independent researcher

Designing inspectable interventions; testing what their evidence supports.

I design and evaluate human-centered AI systems for evidence-bounded decisions. My work focuses on system specifications, inspectable evaluation structures, and evidence review. CheckMyCoach is my primary research system, with human validation in preparation. My academic background is in Applied Psychology; current system evidence comes from implementation and bounded machine evaluation.

I am preparing human validation of CheckMyCoach. Human evaluation has not yet been conducted.

Explore the research program ↓
Research program

Evidence-bounded decision support

Earlier work includes a frozen measurement asset, engineering audit records, a frozen simulation prototype, and a retired behavioral research thread. These materials provide historical context for my current focus on CheckMyCoach.

01 · Background

Earlier work

M0 is a frozen measurement asset, not a validated instrument. MaxFitCalib-Bench is retained as engineering audit history, not a source of prevalence estimates or model rankings. The Precision Illusion behavioral thread is retired and provides historical motivation rather than an active study. Historical context

02 · Current priority

CheckMyCoach

CheckMyCoach is my primary research system. InteractionKit is released, frozen methodological software, and Knowledge Compiler is supporting evidence infrastructure with partial provenance recovery. Primary research system

03 · Historical context

CalTrust

CalTrust is frozen and simulation-only. Frozen · Simulation only

Selected work

Selected work

CheckMyCoach

A bounded prototype that routes selected AI outputs, generates candidate revisions, and checks target removal and information retention. I co-designed its evaluation decomposition, specified system behavior, and reviewed outputs against these requirements. Human validation is in preparation; human evaluation has not yet been conducted.

Primary research system

CheckMyCoach →

InteractionKit

Released, frozen software for typed AI interaction experiments and contract checks. I originated and co-designed the typed experiment architecture, including specifications and contract requirements. The release makes experimental structure inspectable; it does not establish construct validity or equivalence across implementations.

Released methodological software · Frozen

InteractionKit →

Knowledge Compiler

Supporting infrastructure for typed evidence objects, structural checks, and partial provenance recovery. I designed recovery and quarantine rules and reviewed their implementation against structural and evidence-boundary requirements. These checks expose unresolved evidence links without establishing universal source fidelity or downstream validity.

Supporting evidence infrastructure · Partial provenance recovery

Knowledge Compiler →
Latest

Activity

More →
Sep 17

Evaluation Runtime evidence live — bounded reliability scenarios; source repository public

Artifact
Jul 31

InteractionKit v1.0.0 released — 7 contract tests, type checking, production build, and clean-install reproduction

Release
Jul 11

CheckMyCoach Interactive Demo live — M1–M4 pipeline visualization

Release
Jul 3

Expanded discussion with a dual-process theory framing

Writing