Research program
From human decision evidence to governed agent action
My research asks where model-generated behavior corresponds to observed pathways, where it diverges, and what controls are needed before those decisions affect an external system.
- 01
Measurement
Which constructs and subgroup differences are supported by human response data?
- 02
Behavioral correspondence
Do generated responses preserve relevant pathway structure across repeated runs?
- 03
Governance
Which physical, financial, and behavioral constraints should gate an agent action?
- 04
Coupling
How do decisions alter exposure, damage, insurance, and household financial burden over time?
Research toolkit
Methods and systems
- Survey design and psychometric measurement
- Confirmatory factor analysis and structural equation modeling
- Subgroup-aware and repeated-run LLM evaluation
- Agent-based and catastrophe-model coupling
- Structured outputs, validators, targeted repair, and audit logs
- Python research software, tests, CI, and versioned research artifacts
Validity boundaries
- Survey evidence is specific to its population, place, constructs, and collection process.
- Statistical model fit is evidence about a specified structure, not proof of causality.
- LLM response correspondence does not make an agent equivalent to a person.
- Governance rules expose and control declared constraints; they cannot cover unknown failure modes.
- Simulation results remain conditional on model structure, calibration, and scenario assumptions.
The Behavioral Systems Observatory
Follow a decision through the full research system
Each stage exposes a different validity question: what was measured, how pathways were estimated, what the model changes, and where agent behavior can fail.
Open a stage to inspect its evidence, risk, and relevance.
01Human evidence
What do people report doing under flood risk?
- Input
- Household profiles from the Passaic River Basin survey.
- Method
- Survey design, psychometrics, and subgroup-aware measurement.
- Output
- Observed constructs and adaptation choices grounded in measured responses.
- Validity risk
- Sampling, self-report, and construct validity limit generalization.
Gives AI evaluation an empirical reference instead of an intuition-only rubric.
02Decision pathways
Which latent factors and pathways shape adaptation decisions?
- Input
- Risk perception, efficacy, constraints, tenure, and reported actions.
- Method
- Confirmatory factor analysis and multi-group structural equation modeling.
- Output
- Interpretable pathway structures for overall and social-group comparisons.
- Validity risk
- Model fit does not establish causal identification by itself.
Moves evaluation from matching answers to comparing decision structure.
03Behavioral simulation
How do household decisions accumulate over time and space?
- Input
- Survey-grounded agent parameters, flood exposure, and insurance mechanics.
- Method
- Agent-based modeling coupled to a catastrophe model.
- Output
- Household trajectories across multiple census tracts from 2011–2023.
- Validity risk
- Results depend on calibration choices and scenario assumptions.
Tests behavior in a stateful system where choices create downstream effects.
04LLM behavior evaluation
Do generated decisions correspond to measured human pathways?
- Input
- Label-blind personas derived from approved survey attributes.
- Method
- Human-versus-LLM pathway comparison with repeated-run stability checks.
- Output
- A subgroup-aware evaluation design; manuscript in preparation.
- Validity risk
- Plausible language can mask structural disagreement or instability.
Separates surface fluency from evidence-grounded behavioral correspondence.
05Governance and repair
Can an agent action pass explicit domain and behavioral constraints?
- Input
- Structured proposed actions, ratings, and supporting rationale.
- Method
- Parsing, physical/financial/behavioral validation, targeted repair, and audit traces.
- Output
- Accepted, rejected, or repaired decisions before state updates.
- Validity risk
- Rules cover declared constraints, not every real-world failure mode.
Turns governance into inspectable system behavior instead of a policy slogan.
06External-system feedback
What changes after a decision enters the coupled environment?
- Input
- Adaptation action, hazard, property, insurance, and financial state.
- Method
- Deterministic state transitions and stochastic scenario ensembles.
- Output
- Updated exposure, damage, payout, and out-of-pocket burden.
- Validity risk
- Feedback quality is bounded by the external models and available data.
Evaluates agents by consequences in a changing system, not isolated text.
Contact
Build evaluations that connect model behavior to real evidence
I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.
View recruiter brief