Selected work

Evidence-grounded agent systems

The work moves from measured decisions to model structure, agent behavior, constraint checks, and consequences in a coupled environment.

Selected work

What I build and investigate

Each project shows a different part of the same practice: evaluating decisions, tracing their consequences, and governing when they may alter a system.

01In preparation

Human-Grounded LLM Evaluation

Compare LLM-generated decisions with measured human pathways across social groups.

I designed the human-versus-model comparison, subgroup lenses, and repeated-run stability checks.

  • Evaluation design
  • Psychometrics and SEM
  • Stability analysis
View case study: Human-Grounded LLM Evaluation
02Research prototype · archived

FLOODABM

Connect household adaptation to flood damage, insurance, and financial exposure.

I built the household decision model and connected adaptation choices to catastrophe, insurance, and financial mechanics.

  • Agent-based modeling
  • Coupled simulation
  • Risk systems
View case study: FLOODABM
03In preparation for release

WAGF

Check and repair agent decisions before they update a coupled simulation.

I designed the validation and repair path that checks an agent proposal before the simulation state can change.

  • Structured outputs
  • Constraint validation
  • Auditable repair
View case study: WAGF

Open-source systems proof

Research infrastructure with inspectable failure behavior

Selected repositories support research operations, agent orchestration, and multilingual learning. Counts are refreshed through a reviewed build-time data workflow.

awesome-agentic-ai-zh

Maintained trilingual roadmap for agentic AI, with automated cross-language checks.

6.2K stars · 836 forks

awesome-agentic-ai-zh

ai-research-skills

Composable research workflow from gap discovery through study design and drafting.

220 stars · 17 forks

ai-research-skills

research-hub

Local-first Zotero, Obsidian, and NotebookLM workspace with provenance checks.

52 stars · 8 forks

research-hub

agent-collab-skills

Task splitting, reconciliation, debate, and acceptance gates for multi-agent work.

23 stars · 6 forks

agent-collab-skills

codex-delegate

Cross-platform delegation wrapper that verifies completion from repository state.

63 stars · 6 forks

codex-delegate

Updated: 2026-08-23 · Source: GitHub REST API snapshot

PDF

Documents

Choose the document for the context. The website keeps one identity; the PDFs change emphasis for the reader.

Contact

Build evaluations that connect model behavior to real evidence

I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.