Research program

From human decision evidence to governed agent action

My research asks where model-generated behavior corresponds to observed pathways, where it diverges, and what controls are needed before those decisions affect an external system.

Research toolkit

Methods and systems

  1. Survey design and psychometric measurement
  2. Confirmatory factor analysis and structural equation modeling
  3. Subgroup-aware and repeated-run LLM evaluation
  4. Agent-based and catastrophe-model coupling
  5. Structured outputs, validators, targeted repair, and audit logs
  6. Python research software, tests, CI, and versioned research artifacts

Validity boundaries

  • Survey evidence is specific to its population, place, constructs, and collection process.
  • Statistical model fit is evidence about a specified structure, not proof of causality.
  • LLM response correspondence does not make an agent equivalent to a person.
  • Governance rules expose and control declared constraints; they cannot cover unknown failure modes.
  • Simulation results remain conditional on model structure, calibration, and scenario assumptions.

The Behavioral Systems Observatory

Follow a decision through the full research system

Each stage exposes a different validity question: what was measured, how pathways were estimated, what the model changes, and where agent behavior can fail.

Open a stage to inspect its evidence, risk, and relevance.

01Human evidence

What do people report doing under flood risk?

Input
Household profiles from the Passaic River Basin survey.
Method
Survey design, psychometrics, and subgroup-aware measurement.
Output
Observed constructs and adaptation choices grounded in measured responses.
Validity risk
Sampling, self-report, and construct validity limit generalization.
For AI teams

Gives AI evaluation an empirical reference instead of an intuition-only rubric.

View case study
02Decision pathways

Which latent factors and pathways shape adaptation decisions?

Input
Risk perception, efficacy, constraints, tenure, and reported actions.
Method
Confirmatory factor analysis and multi-group structural equation modeling.
Output
Interpretable pathway structures for overall and social-group comparisons.
Validity risk
Model fit does not establish causal identification by itself.
For AI teams

Moves evaluation from matching answers to comparing decision structure.

View case study
03Behavioral simulation

How do household decisions accumulate over time and space?

Input
Survey-grounded agent parameters, flood exposure, and insurance mechanics.
Method
Agent-based modeling coupled to a catastrophe model.
Output
Household trajectories across multiple census tracts from 2011–2023.
Validity risk
Results depend on calibration choices and scenario assumptions.
For AI teams

Tests behavior in a stateful system where choices create downstream effects.

View case study
04LLM behavior evaluation

Do generated decisions correspond to measured human pathways?

Input
Label-blind personas derived from approved survey attributes.
Method
Human-versus-LLM pathway comparison with repeated-run stability checks.
Output
A subgroup-aware evaluation design; manuscript in preparation.
Validity risk
Plausible language can mask structural disagreement or instability.
For AI teams

Separates surface fluency from evidence-grounded behavioral correspondence.

View case study
05Governance and repair

Can an agent action pass explicit domain and behavioral constraints?

Input
Structured proposed actions, ratings, and supporting rationale.
Method
Parsing, physical/financial/behavioral validation, targeted repair, and audit traces.
Output
Accepted, rejected, or repaired decisions before state updates.
Validity risk
Rules cover declared constraints, not every real-world failure mode.
For AI teams

Turns governance into inspectable system behavior instead of a policy slogan.

View case study
06External-system feedback

What changes after a decision enters the coupled environment?

Input
Adaptation action, hazard, property, insurance, and financial state.
Method
Deterministic state transitions and stochastic scenario ensembles.
Output
Updated exposure, damage, payout, and out-of-pocket burden.
Validity risk
Feedback quality is bounded by the external models and available data.
For AI teams

Evaluates agents by consequences in a changing system, not isolated text.

View case study

Contact

Build evaluations that connect model behavior to real evidence

I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.

View recruiter brief