Case study 01 · LLM behavior evaluation
In preparation
Human-Grounded LLM Evaluation
A subgroup-aware study of whether LLM-generated flood-adaptation responses correspond to decision pathways measured in 937 household profiles.
- 937 approved human profiles
- overall and social-group lenses
- repeated-run stability design
Problem
Why this problem matters
Fluent responses can look plausible while preserving the wrong relationships among risk perception, efficacy, constraints, and action. The study evaluates structure, not just surface similarity.
System artifact
Inspect the research logic
The public diagram is a study-design schematic. Numeric comparison results remain unpublished and are intentionally omitted.
Study-design schematic · no unpublished coefficients
What this lens emphasizesCompare the full pathway structure against the survey-grounded reference.
Role & method
What I built
I designed the human-versus-LLM evaluation, prepared label-blind personas, specified the psychometric and multi-group SEM comparison, and built repeated-run stability checks.
- Map approved survey attributes into personas without group labels.
- Collect structured generated responses under a fixed protocol.
- Estimate construct and pathway structure with the same approved measurement logic.
- Compare overall, subgroup, and repeated-run behavior without claiming human equivalence.
Validation and limitations
Validation
- Same construct definitions across human and generated responses
- Subgroup-aware comparison
- Repeated-run stability checks
- Explicit manuscript and result status
Limitations
- Survey responses are self-reported and context-specific.
- Persona prompts cannot encode every social or situational factor.
- Pathway correspondence is not evidence that an LLM is a human substitute.
- Unpublished effect sizes and subgroup results are not shown.
What changed
Behavioral evaluation becomes more useful when the unit of comparison is a decision pathway with declared validity limits, rather than one plausible answer.
Contact
Build evaluations that connect model behavior to real evidence
I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.