awesome-agentic-ai-zh
Maintained trilingual roadmap for agentic AI, with automated cross-language checks.
7K stars · 948 forks
awesome-agentic-ai-zhLLM Evaluation & Agent Systems Research Engineer · Lehigh University
I evaluate whether LLM-agent decisions correspond to measured human behavior, and build governed systems that validate actions before they change real or simulated environments.
Human-grounded evaluation is my core specialty; agent governance and behavioral simulation let me inspect decisions before and after they affect a system.

Selected work
Each project shows a different part of the same practice: evaluating decisions, tracing their consequences, and governing when they may alter a system.
I designed the human-versus-model comparison, subgroup lenses, and repeated-run stability checks.
I built the household decision model and connected adaptation choices to catastrophe, insurance, and financial mechanics.
I designed the validation and repair path that checks an agent proposal before the simulation state can change.
Decision Provenance Explorer
Switch lenses to follow evidence, context, an LLM decision, its validation, and the state change that follows.
Interactive controls require JavaScript. The complete flow and sources remain below.
Measured constructs and reported choices establish the empirical reference.
Published evidenceA label-blind synthetic profile carries only approved attributes into the prompt.
Illustrative exampleRepeated model runs produce a choice and a concise rationale for comparison.
Illustrative exampleChecks compare direction, subgroup behavior, and run-to-run stability without exposing respondent records.
Public artifactThe result becomes an evaluation finding, not a claim that the model represents a person.
Illustrative exampleBehavioral and physical constraints define what a plausible action must respect.
Published evidenceA synthetic agent state limits the actions the model may propose.
Illustrative exampleThe LLM returns a structured proposal rather than mutating the simulation directly.
Illustrative exampleValidators reject invalid fields, request targeted repair, and retain an audit trace.
Public artifactOnly an accepted proposal updates the coupled model state.
Illustrative exampleSurvey evidence and public flood-risk sources ground household behavior.
Published evidenceHousehold tenure, risk perception, resources, and exposure form the decision context.
Published evidenceAgents choose adaptation, insurance, or no action under modeled constraints.
Public artifactModel checks keep behavioral, financial, and physical state transitions consistent.
Public artifactFlood losses and adaptation alter the household and environmental state for the next cycle.
Published evidenceOpen-source systems proof
Selected repositories support research operations, agent orchestration, and multilingual learning. Counts are refreshed through a reviewed build-time data workflow.
Maintained trilingual roadmap for agentic AI, with automated cross-language checks.
7K stars · 948 forks
awesome-agentic-ai-zhComposable research workflow from gap discovery through study design and drafting.
252 stars · 16 forks
ai-research-skillsTask splitting, reconciliation, debate, and acceptance gates for multi-agent work.
27 stars · 6 forks
agent-collab-skillsUpdated: 2026-09-15 · Source: GitHub REST API snapshot
Articles
Three practical notes on evaluating behavior, governing agent actions, and tracing decisions into system consequences. They explain methods without disclosing unpublished results.
A behavioral evaluation should compare decision structure, subgroup patterns, and stability, not just answer similarity.
Treat an LLM output as a proposal. Deterministic checks decide whether it may alter the system.
A decision model becomes consequential when actions alter the environment that shapes the next decision.
Contact
I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.
View recruiter brief