Articles

Methods, failure modes, and engineering judgment

Three practical notes on evaluating behavior, governing agent actions, and tracing decisions into system consequences. They explain methods without disclosing unpublished results.

01

Evaluating LLM Agents Against Measured Human Behavior

A behavioral evaluation should compare decision structure, subgroup patterns, and stability, not just answer similarity.

Read article
  1. 01Measured behavior
  2. 02Controlled persona
  3. 03Repeated decisions
  4. 04Structure + stability
Technical diagram
02

Why Governed Agents Need Validators Before State Changes

Treat an LLM output as a proposal. Deterministic checks decide whether it may alter the system.

Read article
  1. 01Structured proposal
  2. 02Deterministic checks
  3. 03Targeted repair
  4. 04Authorized update
Technical diagram
03

From Individual Decisions to System Consequences

A decision model becomes consequential when actions alter the environment that shapes the next decision.

Read article

Contact

Build evaluations that connect model behavior to real evidence

I am seeking a full-time Summer 2027 internship from approximately late May through mid-August. I am especially interested in LLM evaluation, agent systems, behavioral simulation, and AI for science.