Baixin (Max) Guo
Human-Centered AI · AI Evaluation & Learning · Reliable AI Agents
Canonical research-outreach CV · Updated September 2026
Research Profile
Applied Psychology graduate working at the intersection of human-centered AI and reliable AI systems. I build reproducible evaluation and agent infrastructure, and study how uncertainty, imperfect feedback, and evidence quality affect AI decisions and oversight. I am increasingly interested in how these signals can support learning and adaptation.
Research interests: human/evaluator feedback; preference learning and post-training; reliable AI agents; reproducible AI evaluation.
Selected artifacts: CheckMyCoach · Evaluation Runtime
Selected Research & Technical Experience
- Built an offline AI decision-support pipeline with evidence retrieval, rule-based routing/tagging, candidate revision generation, CLI/MCP interfaces, and structured JSONL audit logs.
- Defined bounded evaluation criteria for routing, conditional revision, preservation, and end-to-end behavior; human-participant validation is in preparation.
- Specified and co-designed a serial runtime that keeps execution attempts, evaluation judgments, recovery decisions, and report denominators separately inspectable; directed implementation of the reproducible CLI.
- Defined ambiguity-aware retry, event-consistency, and report-construction semantics; candidate validation reproduced 279 tests, while replay of unchanged stored inputs produced no duplicate runs/evaluations and byte-identical reports.
Undergraduate Thesis — Evaluating Chinese LLMs on CBT-Bench
- Designed a fixed zero-shot evaluation across five Chinese LLM services on 442 structured CBT-Bench examples; implemented batch API execution with checkpointing and retry logic.
- Reported exact match, partial-overlap, and sample-average F1 metrics while bounding claims to benchmark evidence rather than clinical or human-outcome validity.
Trojan Sports Research Lab, USC Marshall School of Business
Research collaboration · Supervisor: Prof. Lorena Martin
- Collected, cleaned, visualized, and analyzed multi-source baseball data for research on pitcher workload, pitch characteristics, and injury risk.
Education
Changchun Humanities and Sciences College
Bachelor’s degree in Applied Psychology
Technical & Research Skills
Programming / Engineering
Python; JavaScript/TypeScript; SQL/SQLite; Git/GitHub; pytest; CLI tooling; structured data pipelines
ML / AI
LLM/API pipelines; LLM evaluation; agent prototyping; supervised ML workflows; contextual bandits (LinUCB)
Research / Measurement
Experimental design; psychometrics; inter-rater reliability; behavioral study design; statistical analysis