Start with a behavioral reference
An LLM can produce a fluent, plausible decision and still rely on a different structure from the people it is meant to simulate. The reference therefore cannot be a preferred answer written by the evaluator. It needs measured constructs, reported choices, and a model of how those variables relate.
The comparison is about structure. Does risk perception move the generated decision in the expected direction? Do constraints suppress an action when they should? Are differences between social groups preserved without placing the group label in the prompt?
Separate prompt construction from evaluation
Personas should contain only approved inputs and should be generated through a reproducible mapping. Labels that reveal the expected outcome must stay out. Repeated runs then expose whether a conclusion survives ordinary model variability.
Evaluation can combine choice agreement, pathway direction, subgroup contrasts, and stability. No single measure proves human equivalence. Together they locate where the model is useful and where it is only persuasive.
Failure modes to report
Common failures include collapsing distinct groups into one average response, producing the correct choice for the wrong stated reason, changing behavior across runs, and overclaiming causality from an associational human model. These are findings, not nuisances to hide.
Practical implication
For simulated-user systems, the deliverable is a bounded validity statement: which behaviors, populations, and conditions were examined; what aligned; what did not; and what remains unknown. That statement is more useful than a broad claim that the agent is realistic.