LLM Evaluation & Agent Systems Research Engineer · Lehigh University

Wenyu Chiou

I evaluate whether LLM-agent decisions correspond to measured human behavior, and build governed systems that validate actions before they change real or simulated environments.

Human-grounded evaluation is my core specialty; agent governance and behavioral simulation let me inspect decisions before and after they affect a system.

  • LLM behavior evaluation
  • Agent governance
  • Behavioral simulation
  • Reproducible research systems
Explore selected work

Open to Summer 2027 internships · Ph.D. expected Dec 2027 · Full-time opportunities after graduation

Wenyu Chiou beside his coupled flood-adaptation modeling poster at AGU 2025.
Presenting coupled human–flood modeling at AGU 2025, New Orleans.

AI recruiter fit explorer

Map a role to evidence, not a generic score.

Choose a hiring lens to see where my work is a direct fit, where experience transfers, and what still needs a conversation.

Selected work

What I build and investigate

Each project shows a different part of the same practice: evaluating decisions, tracing their consequences, and governing when they may alter a system.

01In preparation

Human-Grounded LLM Evaluation

Compare LLM-generated decisions with measured human pathways across social groups.
Evaluation design · Psychometrics and SEM · Stability analysis

I designed the human-versus-model comparison, subgroup lenses, and repeated-run stability checks.

  • Evaluation design
  • Psychometrics and SEM
  • Stability analysis
View case study: Human-Grounded LLM Evaluation
02Research prototype · archived

FLOODABM

Connect household adaptation to flood damage, insurance, and financial exposure.
Agent-based modeling · Coupled simulation · Risk systems

I built the household decision model and connected adaptation choices to catastrophe, insurance, and financial mechanics.

  • Agent-based modeling
  • Coupled simulation
  • Risk systems
View case study: FLOODABM
03In preparation for release

WAGF

Check and repair agent decisions before they update a coupled simulation.
Structured outputs · Constraint validation · Auditable repair

I designed the validation and repair path that checks an agent proposal before the simulation state can change.

  • Structured outputs
  • Constraint validation
  • Auditable repair
View case study: WAGF

Decision Provenance Explorer

Inspect what supports a decision before trusting its consequence.

Switch lenses to follow evidence, context, an LLM decision, its validation, and the state change that follows.

Wenyu Chiou: human evidence and context inform an LLM proposal, validation and repair, and human-environment feedback.Wenyu Chiou: human evidence and context inform an LLM proposal, validation and repair, and human-environment feedback.

Interactive controls require JavaScript. The complete flow and sources remain below.

Evaluation

  1. Human evidence

    Measured constructs and reported choices establish the empirical reference.

    Published evidence
  2. Context

    A label-blind synthetic profile carries only approved attributes into the prompt.

    Illustrative example
  3. LLM proposal

    Repeated model runs produce a choice and a concise rationale for comparison.

    Illustrative example
  4. Validation / repair

    Checks compare direction, subgroup behavior, and run-to-run stability without exposing respondent records.

    Public artifact
  5. Consequence

    The result becomes an evaluation finding, not a claim that the model represents a person.

    Illustrative example
Inspect the complete case

Governance

  1. Human evidence

    Behavioral and physical constraints define what a plausible action must respect.

    Published evidence
  2. Context

    A synthetic agent state limits the actions the model may propose.

    Illustrative example
  3. LLM proposal

    The LLM returns a structured proposal rather than mutating the simulation directly.

    Illustrative example
  4. Validation / repair

    Validators reject invalid fields, request targeted repair, and retain an audit trace.

    Public artifact
  5. Consequence

    Only an accepted proposal updates the coupled model state.

    Illustrative example
Inspect the complete case

Simulation

  1. Human evidence

    Survey evidence and public flood-risk sources ground household behavior.

    Published evidence
  2. Context

    Household tenure, risk perception, resources, and exposure form the decision context.

    Published evidence
  3. Agent decision

    Agents choose adaptation, insurance, or no action under modeled constraints.

    Public artifact
  4. Validation / repair

    Model checks keep behavioral, financial, and physical state transitions consistent.

    Public artifact
  5. Consequence

    Flood losses and adaptation alter the household and environmental state for the next cycle.

    Published evidence
Inspect the complete case

Open-source systems proof

Research infrastructure with inspectable failure behavior

Selected repositories support research operations, agent orchestration, and multilingual learning. Counts are refreshed through a reviewed build-time data workflow.

awesome-agentic-ai-zh

Maintained trilingual roadmap for agentic AI, with automated cross-language checks.

7K stars · 948 forks

awesome-agentic-ai-zh

ai-research-skills

Composable research workflow from gap discovery through study design and drafting.

252 stars · 16 forks

ai-research-skills

agent-collab-skills

Task splitting, reconciliation, debate, and acceptance gates for multi-agent work.

27 stars · 6 forks

agent-collab-skills

Updated: 2026-09-15 · Source: GitHub REST API snapshot

Articles

Methods, failure modes, and engineering judgment

Three practical notes on evaluating behavior, governing agent actions, and tracing decisions into system consequences. They explain methods without disclosing unpublished results.

Articles

Contact

Build evaluations that connect model behavior to real evidence

I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.

View recruiter brief