Recruiter brief

Wenyu Chiou — LLM Evaluation & AI Research Profile

A 60-second view of the research engineering work I can own, the evidence behind it, and my availability.

Wenyu Chiou beside his coupled flood-adaptation modeling poster at AGU 2025.
Presenting coupled human–flood modeling at AGU 2025, New Orleans.

01

Role fit in 60 seconds

I am an LLM Evaluation and AI Research Engineer focused on whether model decisions correspond to measured human behavior. I also build governed agent systems and behavioral simulations that make decision inputs, failures, repairs, and downstream consequences inspectable.

02

What I can own

  1. 01

    Human-grounded LLM evaluation

    Design subgroup-aware and stability-aware evaluations that compare generated decisions with measured behavioral evidence and state validity limits clearly.

  2. 02

    Governed agent workflows

    Define structured outputs, deterministic validators, targeted repair, and audit traces before model proposals can change system state.

  3. 03

    Behavioral simulation and AI for science

    Connect individual decisions to agent-based and coupled simulations so downstream environmental, financial, and institutional effects can be tested.

03

Evidence of practice

Three projects show the evaluation, governance, and simulation work behind this profile. Public status is stated directly.

01

In preparation

Human-Grounded LLM Evaluation

Problem
Compare LLM-generated decisions with measured human pathways across social groups.
My role
I designed the human-versus-model comparison, subgroup lenses, and repeated-run stability checks.
Verified capabilities
Evaluation design · Psychometrics and SEM · Stability analysis
Public status
In preparation. Public descriptions and synthetic examples only.
02

Research prototype · archived

FLOODABM

Problem
Connect household adaptation to flood damage, insurance, and financial exposure.
My role
I built the household decision model and connected adaptation choices to catastrophe, insurance, and financial mechanics.
Verified capabilities
Agent-based modeling · Coupled simulation · Risk systems
Public status
Public WRR article, AGU poster, and repository; the case-study analysis remains in revision.
03

In preparation for release

WAGF

Problem
Check and repair agent decisions before they update a coupled simulation.
My role
I designed the validation and repair path that checks an agent proposal before the simulation state can change.
Verified capabilities
Structured outputs · Constraint validation · Auditable repair
Public status
Framework in preparation. The trace below is synthetic and illustrative.

04

Verified technical fit

These tools and methods reflect my research, open-source work, and working development practice.

Programming and research engineering
Python · R · MATLAB · pytest/CI · Git
LLM and agent systems
LLM APIs · LangChain · Ollama · Model Context Protocol (MCP) · Retrieval-Augmented Generation (RAG) · agent memory · tool use · Agent Skills · plugin architectures · multi-agent workflows
AI engineering tools
OpenAI Codex · Claude Code · AI coding agents
Modeling and inference
Structural Equation Modeling (SEM) · Bayesian calibration · Agent-Based Modeling (ABM) · Hydrological Modeling · Sociohydrological Modeling

05

Availability and work authorization

  • Open to Summer 2027 internships
  • Ph.D. expected Dec 2027
  • F-1 student; CPT eligible for internships
  • Open to full-time opportunities after graduation

06

Ask the portfolio

Start with a recruiter question. The portfolio shows local evidence immediately and may use NVIDIA to route you to cited pages.

07

Resume and contact

For LLM evaluation, AI research, agent systems, or AI for science roles, use email or LinkedIn. The English industry resume is the shortest complete record.