Case study 01 · LLM behavior evaluation

In preparation

Human-Grounded LLM Evaluation

A subgroup-aware study of whether LLM-generated flood-adaptation responses correspond to decision pathways measured in 937 household profiles.

  • 937 approved human profiles
  • overall and social-group lenses
  • repeated-run stability design

Problem

Why this problem matters

Fluent responses can look plausible while preserving the wrong relationships among risk perception, efficacy, constraints, and action. The study evaluates structure, not just surface similarity.

System artifact

Inspect the research logic

The public diagram is a study-design schematic. Numeric comparison results remain unpublished and are intentionally omitted.

Study-design schematic · no unpublished coefficients

What this lens emphasizesCompare the full pathway structure against the survey-grounded reference.

Role & method

What I built

I designed the human-versus-LLM evaluation, prepared label-blind personas, specified the psychometric and multi-group SEM comparison, and built repeated-run stability checks.

  1. Map approved survey attributes into personas without group labels.
  2. Collect structured generated responses under a fixed protocol.
  3. Estimate construct and pathway structure with the same approved measurement logic.
  4. Compare overall, subgroup, and repeated-run behavior without claiming human equivalence.

Validation and limitations

Validation

  • Same construct definitions across human and generated responses
  • Subgroup-aware comparison
  • Repeated-run stability checks
  • Explicit manuscript and result status

Limitations

  • Survey responses are self-reported and context-specific.
  • Persona prompts cannot encode every social or situational factor.
  • Pathway correspondence is not evidence that an LLM is a human substitute.
  • Unpublished effect sizes and subgroup results are not shown.

What changed

Behavioral evaluation becomes more useful when the unit of comparison is a decision pathway with declared validity limits, rather than one plausible answer.

Contact

Build evaluations that connect model behavior to real evidence

I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.