Research program Evidence-bounded decision support
I design and evaluate human-centered AI systems for evidence-bounded decision support through system specification, inspectable evaluation structures, and evidence review.
The research question is when an AI output communicates more certainty than its evidence supports, when a system should intervene, and how to evaluate the consequences for human reliance. CheckMyCoach is the primary system through which I pursue this question. InteractionKit provides released methodological software, while Knowledge Compiler supports evidence handling. These projects have distinct roles and evidence limits; they are not a single validated platform or three equally active research studies.
Primary research system CheckMyCoach
A bounded four-module prototype that routes selected outputs, assigns rule-based tags, generates candidate revisions, and applies deterministic target-and-retention checks. My role centers on co-designing the evaluation decomposition, specifying system behavior, and reviewing outputs against target-removal and information-retention requirements. In a fixed 40-case run on a constructed development corpus, the pipeline routed 15 cases; 12 delivered outputs removed the predefined target feature, while none passed every target-and-retention check. This is a development result, not evidence of generalization or improved human decisions.
Historical context
Earlier work provides historical motivation rather than active research. M0 is a frozen measurement asset, not a validated instrument. MaxFitCalib-Bench is retained as engineering audit history, not a source of prevalence estimates or model rankings. CalTrust is frozen and simulation-only. The Precision Illusion behavioral thread is retired.
Current focus
I am preparing human validation of CheckMyCoach. Human evaluation has not yet been conducted.
CheckMyCoach project → Reading the evidence
Software implementation, bounded machine evaluation, and human validation answer different questions. The first two are represented in these HCAI artifacts; completed human validation is not.