The work moves from measured decisions to model structure, agent behavior, constraint checks, and consequences in a coupled environment.
Selected work
What I build and investigate
Each project shows a different part of the same practice: evaluating decisions, tracing their consequences, and governing when they may alter a system.
01In preparation
Human-Grounded LLM Evaluation
Compare LLM-generated decisions with measured human pathways across social groups.
I designed the human-versus-model comparison, subgroup lenses, and repeated-run stability checks.
Research infrastructure with inspectable failure behavior
Selected repositories support research operations, agent orchestration, and multilingual learning. Counts are refreshed through a reviewed build-time data workflow.
awesome-agentic-ai-zh
Maintained trilingual roadmap for agentic AI, with automated cross-language checks.
Build evaluations that connect model behavior to real evidence
I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.