← All projects
Frozen · Simulation Only v1.0 Updated: 2026-07-03

CalTrust

Simulation-tested adaptive-intervention research prototype.

CalTrust simulation prototype with an XGBoost acceptance proxy

The research question: Could different user contexts benefit from different calibration interventions? CalTrust tests this design hypothesis in simulation; it does not yet establish differential human effects.

The simulation prototype: CalTrust uses a LinUCB contextual bandit to select among four interventions using synthetic user contexts and output characteristics. An XGBoost proxy is trained on simulated trial data to predict simulated acceptance and confidence outcomes; the training-run metrics were not archived as a report artifact, and human validation has not yet been conducted.

Why this matters for HCI

Tests in simulation whether different synthetic user contexts may benefit from different uncertainty-presentation strategies, motivating future human-subject evaluation.

Intervention Selection Logic

Contextual bandit selects among four interventions using synthetic user features and a historical UCS heuristic User Profile naive / intermediate / expert LLM Output historical UCS heuristic LinUCB Contextual Bandit No Intervention output unchanged Verbal Uncertainty "This may vary…" Numerical Range "10-15% reduction" Evidence Citation "per ACSM guidelines"

Simulation Result Boundary

The current archived report supports a simulation prototype description, but not a reproducible profile-level policy comparison. No human outcome or calibrated-trust result is claimed.

Key Insight

Within the simulation, the bandit frequently selected "no intervention" for the naive synthetic profile. This is simulated policy behavior, not evidence about real users or cognitive load.

Tech Stack

Python 3.12 XGBoost LinUCB scikit-learn Streamlit

Key Takeaways

  1. LinUCB simulation prototype selects among four intervention types using synthetic user and output contexts.
  2. XGBoost acceptance/confidence proxies are trained on simulated trial data to predict simulated acceptance and confidence outcomes; the training-run metrics were not archived as a report artifact.
  3. The simulated policy often selected "no intervention" for the naive synthetic profile; human validation has not yet been conducted.

Changelog

v1.02026-07-03 — Simulation pipeline complete.
v0.22026-06-15 — LinUCB + Streamlit integrated.
v0.12026-02-01 — Prototype.