In preparation

LLM evaluation study — human pathways vs. generated responses

Problem

An LLM can produce a fluent answer without reproducing how real households make decisions. The difficult question is not whether a response sounds human; it is whether a synthetic decision-maker preserves empirically observed pathways and meaningful differences between social groups.

Why it matters

If LLM agents replace human behavior inside consequential simulations, apparent plausibility is not enough. Evaluation needs a human ground truth, a frozen comparison design, and visible limits before synthetic behavior is used as a behavioral claim.

Approach

My role

Designed the evaluation scope, redesigned the elicitation protocol, implemented the analysis pipeline, and maintained the regression tests that protect the study from circular persona labels and coding drift.

Method

A two-arm design: first establish the corrected human decision pathways with multi-group structural equation modeling, then ask LLMs to respond as label-blind socioeconomic personas drawn from the same survey population. The primary comparison asks whether the generated responses reproduce pathway structure and social-group differences.

What was built

The study is grounded in 937 household records and a redesigned elicitation. Full subgroup and model-pilot details remain private while the study is in preparation.

Key challenge

Keeping the comparison label-blind and non-circular while separating genuine behavioral signal from prompt, coding, and model artifacts. The study records redesign decisions and retracted artifacts instead of carrying an attractive but unsupported finding into the paper.

Validation & limitations

Current milestones are the implemented redesign and corrected human ground truth. No final agreement score or subgroup claim is reported here until the full run and scope validation are complete.

Results & status

In preparation — planned submission to Progress in Disaster Science.

Links

Transferable relevance

Academic A validity study for generative agents grounded in primary survey data and explicit social-group comparisons.

Industry Human-grounded LLM evaluation, regression-protected experiment design, and honest reporting of pending scope — directly legible to evaluation, agent, safety, and research-engineering teams.