In preparation
LLM evaluation study — human pathways vs. generated responses
Problem
An LLM can produce a fluent answer without reproducing how real households make decisions. The difficult question is not whether a response sounds human; it is whether a synthetic decision-maker preserves empirically observed pathways and meaningful differences between social groups.
Why it matters
If LLM agents replace human behavior inside consequential simulations, apparent plausibility is not enough. Evaluation needs a human ground truth, a frozen comparison design, and visible limits before synthetic behavior is used as a behavioral claim.
Approach
My role
Designed the evaluation scope, redesigned the elicitation protocol, implemented the analysis pipeline, and maintained the regression tests that protect the study from circular persona labels and coding drift.
Method
A two-arm design: first establish the corrected human decision pathways with multi-group structural equation modeling, then ask LLMs to respond as label-blind socioeconomic personas drawn from the same survey population. The primary comparison asks whether the generated responses reproduce pathway structure and social-group differences.
What was built
The study is grounded in 937 household records and a redesigned elicitation. Full subgroup and model-pilot details remain private while the study is in preparation.
Key challenge
Keeping the comparison label-blind and non-circular while separating genuine behavioral signal from prompt, coding, and model artifacts. The study records redesign decisions and retracted artifacts instead of carrying an attractive but unsupported finding into the paper.
Validation & limitations
Current milestones are the implemented redesign and corrected human ground truth. No final agreement score or subgroup claim is reported here until the full run and scope validation are complete.
Results & status
In preparation — planned submission to Progress in Disaster Science.
Links
Transferable relevance
Academic A validity study for generative agents grounded in primary survey data and explicit social-group comparisons.
Industry Human-grounded LLM evaluation, regression-protected experiment design, and honest reporting of pending scope — directly legible to evaluation, agent, safety, and research-engineering teams.