Skip to content

Stage 7 — Agent Production Engineering: Testable, Observable, Stoppable, and Recoverable

Make an AI helper testable, observable, stoppable, and recoverable before sharing it.

🎯 What This Stage Does (Start Here)

Agent Production Engineering means making an Agent testable, observable, stoppable, and recoverable. Like a toy car going onto a road, it needs steering, brakes, and a dashboard. It need not serve millions.

The chapter uses one story: a research helper checks three sources, writes a summary, and asks before sending.

Remember this order:

Say what success looks like → keep a record of the work → ask a person before risky actions → prove it can continue after a fall → only then give it to other people.

Where you are stuck Do this first Evidence to produce
You do not know whether the summary is good Write fixed examples and success rules A check you can run again
You do not know which step failed Record every step, error, time, and cost One complete work record
It can send mail, pay, delete, or write data Stop before the action, ask a person, and save progress Who approved it and where to continue
The first three checks can rerun and pass Give it to other people Health, stop, continue-after-failure, and rollback instructions

Make one AI helper reliable first. Add helpers only for separable work or independent checks.

⏱ Expand: time, environment, cost, and safety notes
  • Split this stage into several short practice sessions. You do not need to finish it at once.
  • You need Python and Git. The deployment exercise also uses Docker.
  • Run the tests that need no API key first. Set a small budget before calling a paid model.
  • A work record may contain prompts, tool inputs, and model answers. Do not send passwords, personal data, or customer data directly to a tracing service.
  • Another Agent usually adds another model call, more latency, and more debugging. Do not assume Multi-Agent is automatically faster or more accurate.

📌 Learning Goals

After this stage, you can:

  1. Tell apart the AI helper's workspace, its repeat-and-check rhythm, and its branching route.
  2. Turn real failures into checks you can run again instead of trusting one pretty answer.
  3. Find every step, error, time, and cost from one task.
  4. Stop before risky actions and continue from the right saved point.
  5. Use the same evidence to decide whether the system is ready for other people.

🧩 Meet Nineteen Core Terms First

Use “plain-language meaning” to get oriented, then read “how this chapter uses it / technical boundary” to see what the term means here. Related terms share one group, so you do not need to read the same definition twice.

What to solve firstCore termPlain-language meaningHow this chapter uses it / technical boundary
Make the task runAgent HarnessThe room where the AI helper worksThe execution environment that holds the model, tools, permissions, state, error handling, and records; this chapter uses it to check sources and prepare a summary safely
Agent LoopDo one step, see the result, then choose the next stepThe model repeatedly chooses an action and reads the tool result until it finishes, reaches a limit, or must ask a person
Workflow GraphA route map with forksSteps, connections, conditions, and state that say which route to take; this chapter uses it for research, checking, and approval before sending
OrchestrationArrange who goes first and who goes nextControl steps, data flow, roles, retries, and stop conditions
Multi-AgentSeveral AI helpers share the workMultiple Agents complete a task with clear roles; it is an option, not a requirement
HandoffPass the baton and the notes togetherOne Agent passes control, needed data, and result evidence to another Agent
Prove it did the right thingEvaluation / EvalUse the same checklist each timeMeasure an Agent's result and process with fixed cases, environments, grading methods, and thresholds
OutcomeWhat really happened at the endThe externally verifiable state when the task ends; this chapter checks that the summary truly uses three valid sources
TrajectoryThe footprints left along the wayWhat happened during one run, including tool calls, intermediate results, errors, and output
GraderMark one answer using stated rulesA method, program, or model that scores one Eval Case against success criteria; this chapter keeps the rules and human spot checks visible
Evaluation HarnessThe exam room that gives the same test and keeps the scoreA test system that loads cases, reruns the Agent, calls graders, and saves results; it has a different responsibility from the Agent Harness used for daily work
TraceA notebook that collects the footprintsSteps, tool calls, errors, and results arranged by time for one task; this chapter uses it to find the failing step
ObservabilityPut a clear window on the systemUse traces, logs, and metrics to see internal state; this chapter uses it to find where a source went missing
Let it stop and continue safelyGuardrailBlock what must not happenRules that limit inputs, outputs, tool permissions, or risky actions
Human ApprovalAsk a person before a risky actionPause before a sensitive tool call so a person can approve, edit, or reject it
CheckpointSave before moving onSave recoverable workflow state and version information
ResumeContinue from the saved pointLoad a checkpoint with the same task or thread ID and continue execution
RecoveryCome back safely after fallingA strategy to stop, retry, compensate, or hand a failure to a person
IdempotencyPress twice, do it onceRetries with the same idempotency key do not duplicate external side effects

A Prompt is the instruction and material you give the model. Context is the information needed for this step. They still matter; this chapter adds execution, checking, and recovery around them.

🚪 Entry Conditions

You should have completed at least:

  • Stage 4: know what Agents, Tools, and Workflows are.
  • Stage 5: have seen tool permissions, Subagents, and development workflows.
  • Stage 6: know that Context, RAG, and Memory are different.

You can start even if Docker is new to you. Do the four core exercises first and learn Docker for Core Exercise 4.

📚 Required Reading

Read these six in production order:

  1. Anthropic — Demystifying evals for AI agents: distinguish Outcome from the complete Trajectory; an Agent saying “done” does not prove the outside result is done.
  2. OpenAI Agents SDK — Tracing: see how trace, span, tool, handoff, and guardrail events connect one run.
  3. OpenAI Agents SDK — Human-in-the-loop: pause before a sensitive tool, save RunState, approve or reject, and resume.
  4. LangGraph — Persistence: distinguish checkpoints from cross-thread stores; interruption, recovery, and long-term memory are not the same thing.
  5. LangGraph — Interrupts: see how human approval pauses and resumes, and why side effects before an interrupt must be idempotent.
  6. Anthropic — Building Effective Agents: start with simple compositions and add autonomy or Multi-Agent only when real division of labor needs it.
📖 Expand: further reading and purpose
  1. Anthropic — Develop tests and evaluations: define measurable success criteria before choosing a grader.
  2. OpenAI Agents SDK — Testing utilities: test with repeatable fake models instead of paying for every run.
  3. OpenAI Agents SDK — Running agents: see an Agent Loop repeat, and stop it with max_turns.
  4. OpenAI Agents SDK — Multi-agent orchestration: compare manager and Handoff patterns; this is an advanced option, not the first production step.
  5. LangGraph — Workflows and agents: distinguish fixed Workflows from Agents that choose their next step.
  6. Microsoft Agent Framework — Workflow concepts: see how executors, edges, events, and state form a Workflow Graph.
  7. OpenAI — Harness engineering: see how environments, feedback loops, and mechanical rules help Agents work reliably.
  8. OpenTelemetry — GenAI semantic conventions: learn portable tracing fields; the conventions are evolving, so do not assume every platform supports all of them.

🧭 Harness, Loop, Graph, and Eval: How They Work Together

They are not four product generations, and you do not choose only one. Think about the same research helper from four angles:

Responsibility Plain-language question Research-helper example
Agent Harness Where can it work safely? It may read sources, but it must stop before sending
Agent Loop Why should it take another round? If one source is missing, search again; stop at the limit
Workflow Graph Which route should it take now? Go back to research when sources are weak; otherwise ask for approval
Eval How do I know the result and process are acceptable? Check three valid sources, correct citations, and no skipped approval

Eval checks the Outcome and Trajectory, then uses a Grader to decide whether the run meets the stated rules. Eval can make the Loop retry, make the Graph choose another route, or make the Harness stop. Putting a Harness and Eval together still does not create the Loop's rules for repeating and stopping.

Agent Harness is the work environment, Agent Loop is the repeat-and-check rhythm, Workflow Graph is the branching route, and Eval uses a Grader to check Outcome and Trajectory
Open full-size image (new tab)

The learning order is Stage 3 Agent Loop → Stage 4 Workflow Graph / Agent Framework → safe production integration in this chapter. Loop Engineering is an emerging label used by IBM. Graph Engineering is even less settled. Learn the responsibilities first, then treat these labels as search terms used by the community. Sources: IBM — Loop Engineering, Anthropic — Agent harness and eval, and Microsoft Agent Framework — graph-based workflows.

🏗 Agent Harness: Prepare the Safe Workspace

Harness Engineering means designing the runtime that lets a model act as an Agent. The model produces decisions; the Harness processes input, tools, state, permissions, failures, and results, and it often runs the agent loop directly. An outer scheduler may call the Harness many times, so a Harness is not limited to “one short run.” Sources: OpenAI — Harness engineering, Anthropic — agent harness definition, and Anthropic — Managed agents.

The 8 Core Components of a Harness

These eight items are this project’s production checklist, not the world’s only official taxonomy.

Component Plain-language meaning Question before release
1. Orchestration / Run loop Decide what happens next Who starts, who stops, and what if a handoff fails?
2. Tool / Permission boundary Give it only the keys it needs Which tools may read, write, or delete?
3. Context / State / Checkpoint Save where it is now Can it resume from the correct point?
4. Retry / Recovery / Idempotency Try again without charging twice Could a retry repeat an email, payment, or database write?
5. Guardrail / Human approval Ask an adult before a risky action Which actions always require approval?
6. Telemetry / Observability Put a clear window on the system Can we see traces, errors, latency, and tokens?
7. Eval harness Retake the test after every change Are cases, scoring rules, and failure thresholds fixed?
8. Cost / Latency budget Decide the money and time limit first Above budget, should it stop, downgrade, or queue?
🔧 Expand: feedback, recovery, and cost details
  • Write tool errors as feedback an Agent can understand, not only a long stack trace.
  • Keep the grader separate from the worker when possible. Do not ask only, “How good was your own work?”
  • Design idempotency for every external side effect so a retry does not repeat a payment, email, or data write.
  • Prompt caching, batching, model routing, and smaller models may reduce cost, but results depend on the workload. Measure a baseline, change one thing, and measure again.
  • Anthropic prompt caching can be automatic or use explicit cache_control. Cache duration and read/write pricing depend on the option; use the official documentation.
  • Traces may capture sensitive inputs and outputs. Configure redaction, retention, and access before release.

🔁 Agent Loop: Act, Observe, Then Decide

First separate three kinds of Loop that sound similar but cover different scopes:

Name What repeats Example
Program loop The same code block for item in items; this is syntax, not the topic of this section
Agent Loop Model → tool → tool result → model Keep calling tools inside one run until completion or max_turns
Loop Engineering Goal → action → observation → adjustment Repeat work in one long run or across sessions/schedules, with verification, memory, budgets, and stop conditions on every round

IBM explains Loop Engineering as Goal → Action → Observation → Adjustment. The point is not to let an Agent run forever. Every round must answer: Is the goal still valid? Is the evidence enough? Should we continue, stop, or hand control to a person?

Therefore, Loop Engineering is not the next Harness product generation and does not automatically replace a Harness. In Anthropic's terminology, the Harness itself includes the loop that calls the model and routes tools. IBM's broader Loop Engineering practice includes goals, checking, tools, hooks, context, subagents, and persistent state. Sources draw the boundary differently, so remember the responsibilities instead of memorizing one universal layer diagram. Sources: IBM — Loop Engineering and Anthropic — Managed Agents.

When models improve, a particular patch may be removable. For example, Anthropic removed context resets used by an earlier harness for newer models. That only means the same Evals showed one workaround was no longer needed. It does not mean permissions, safety, logs, evals, or recovery automatically became obsolete. Source: Anthropic — Harness design for long-running applications. Stage 7.5 explains how to keep, simplify, or remove each part under Model–Harness Fit.

🗺 Workflow Graph: Know Where to Go at a Branch

The Agent Loop answers, “Should I do this again?” The Workflow Graph answers, “Where should I go next?” It is like a school map with forks. The map arranges the route, but it does not do the work at each stop.

Outside writing sometimes calls this engineering work Graph Engineering. The label is emerging. Official documents often call each stop a node, each connection an edge, and each choice a branch. Learn nodes, edges, branches, cycles, state, checkpoints, and human approval instead of memorizing a label that is not yet standardized.

A node can contain an Agent Loop; the Workflow Graph arranges the order between nodes.

🧠 Expand: choosing a Loop, Graph, or Multi-Agent design
  • Use a Loop when there is one path that may need many retries.
  • Use a Graph / Workflow when there are branches, parallel steps, human approvals, or a need to resume in the middle.
  • Add Multi-Agent only when parts can truly work independently or distinct roles must check one another.
  • A Graph node can be an Agent, a tool, fixed code, or “wait for human approval.” Not every box needs an Agent.

  • Optional official docs: OpenAI Responses Multi-agent is in Beta for GPT-6.1 Sol and all GPT-5.6 models. The model delegates to subagents with separate contexts. They share the request's model and tools; this is not SDK manager / handoff orchestration.

  • max_concurrent_subagents defaults to 3 active subagents across the tree and excludes the root. Concurrency settings, total agents, and tree depth have no fixed cap; delegation can add tokens. max_tool_calls, reasoning.summary, and /responses/compact are unsupported; server-side automatic compaction runs independently for each Agent.
  • The API executes hosted collaboration; your application executes custom function calls. Separate contexts do not isolate tool permissions: the application must still approve sensitive tools and enforce budgets and stop conditions.
  • Google Managed Agents offers Antigravity in Public Preview. antigravity-preview-09-2026 defaults to Gemini 3.8 Flash. It provides a managed Linux sandbox, files preserved across interactions, code execution, custom functions, and remote MCP.
  • Outbound network access is unrestricted by default; set an allowlist and minimal tool permissions. Search and URL fetching do not imply GUI browser control; computer_use is unsupported. A sandbox still needs this chapter's Evals, approval, and recovery.
  • Google's docs describe referencing secrets by managed credential ID: the egress proxy injects them without exposing them in the sandbox. An Agent can use the full scope of a supplied credential; grant only the minimum scope needed.

🧪 Eval: State What Good Means, Then Decide How to Grade

Eval is not one score or a report added after the system is done. First say what success means. Then use the same method to compare versions.

Start with the Outcome. The research helper's Outcome is not “the Agent says the summary is done.” It is “the summary really uses three valid sources, every citation opens, and human approval has not been skipped.”

Next, create a complete Eval Case. It is like an exam question that also includes its rules. Input is only one part.

Part of an Eval Case Research-helper example Why keep it
Input Summarize these three topics Tells the system what to do
Initial State Three candidate sources; not yet approved Fixes the starting environment
Success Criteria All three sources open; the summary has checkable citations Says what success means
Forbidden Actions Do not invent a source; do not send by itself Blocks bad behavior even when the answer looks good
Optional Reference Answer A summary checked by a person Gives comparison guidance when useful; not every case needs one
Grader Code checks links and counts; a person checks faithfulness Says who grades with which rule
Case Metadata Case ID, version, split, source, and labels Makes the same case rerunnable and traceable
A complete Eval Case includes Input, Initial State, Success Criteria, Forbidden Actions, Optional Reference Answer, Grader, and Case Metadata; Input is only one part
Open full-size image (new tab)

Several complete cases form an Eval Suite. Version the Suite so you know whether this run and the last run used the same exam.

This project calls a human-reviewed, reusable collection of complete cases a Reviewed Eval Set. Other sources may say Golden Set or Reference Set, but these labels do not have one shared definition across vendors. Check whether a source means questions, answers, criteria, or the whole collection.

A Golden / Reference Set is not input alone, and it is not the same as training data or few-shot examples. It usually contains complete cases, conditions, reference evidence, and grading methods. Its exact fields still come from the current project's definition.

Add these measurement terms last:

Term Plain-language meaning How this chapter uses it
Trial One actual attempt at one case Run a case more than once when model output can vary
Baseline Measure before changing anything Gives old and new versions the same starting point
Regression The new version becomes worse past a preset limit Check quality, cost, safety, and reliability together
Development Set Practice questions you may look at Rerun after changes and use failures to improve the system
Holdout Set The final exam you do not peek at Open only for a release candidate or final validation

Anthropic's Agent Eval guide separates task, trial, grader, trajectory, and outcome. OpenAI's Graders API lists several grader types. Tools may differ, but each report should keep dataset version, split, case ID, trial count, grader, Outcome, Trajectory, and baseline.

The system that loads cases, reruns an Agent, calls a grader, and saves results is the Evaluation Harness defined earlier. It can call an Agent Harness, but their jobs are different: one runs work safely, and one makes tests repeatable and comparable.

🔎 Observability: See Which Step Failed

Observability is like looking through the clear wall of a transparent box. It does not mean showing everything to everyone. It means keeping useful traces, logs, and metrics while hiding passwords, personal data, and customer data.

For the research helper, one Trace should answer: Which sources were checked? Which tool failed? How many retries happened? How long did it take? Why did it stop before approval? A Trace helps explain the Trajectory, but “we recorded a lot” does not mean the result is correct. Eval still checks the Outcome.

🛑 Approval, Checkpoint, Resume, and Recovery: Stop, Then Continue Safely

When the research helper is ready to send, it enters Human Approval. The person does not start from zero. They receive the summary, sources, and risks together.

Save a Checkpoint before approval. After a restart, Resume uses the same task ID to return to that point. If something failed, Recovery decides whether to retry, compensate, restore an older state, or hand the task to a person. Any action that sends mail, pays, or writes data needs an Idempotency key so the same retry produces the outside effect only once.

🛡 Complete Production Route: Eval → Observability → Approval / Recovery → Deploy

Deploy means giving a checked system to other people. It is like opening a shop. Opening the door is not proof of success; the tests, records, brakes, and recovery plan come first.

These four steps are not maturity badges; they are the check route for the same change:

Order Question to answer Minimum evidence to leave What to do if it fails
1. Eval Is the final result really correct? Did it take a dangerous shortcut? Anthropic suggests starting with 20–50 cases representing real work as a practical range, not a universal minimum. Also record Outcome, Trajectory, grader, cost, and a failure threshold Add cases or fix behavior; do not deploy
2. Observability Can you find the failed step? Task ID, trace/span, tool call, error type, latency, tokens, and sensitive-data redaction Make failures visible before changing the Prompt or model
3. Approval / Recovery Can a risky action stop first? Can it resume safely after interruption? Human approval point, versioned checkpoint, resume test, idempotency key, and reject/timeout/compensation route Fail closed, stop automation, and hand it to a person
4. Deploy Can the first three steps rerun on the new version? Health/readiness, rate limit, rollback, stop switch, version, and release record Keep the old version or roll back; “the service started” is not success

An Outcome Eval checks the outside-world result. For example, an Agent saying “the email was sent” is only text; the test environment having exactly one email sent to the right recipient is an Outcome pass. A Trajectory Eval checks which tools it used, how many attempts it made, whether it bypassed approval, and how many tokens it spent. Use both so a polished final sentence cannot pass by itself.

Build cases from real failures first: for each error, keep a de-identified input, expected Outcome, forbidden actions, and reproduction steps. Do not copy production data into a public repo; use structurally equivalent fake data when needed.

🧭 What Is the Difference Between OpenRouter, Pi, OpenCode, Orca, and QM?

They are not five versions of the same product. Put each one at the right layer:

Name What it is One-line memory aid
OpenRouter Model API gateway / router Connects software to different models; it is not a coding Agent
Pi Agent toolkit and coding-agent CLI Calls models and tools to finish a task
OpenCode Open-source coding Agent Reads, edits, and tests inside a code project
Orca Multi-Agent development environment Runs coding Agents in isolated worktrees for comparison
QM Team Multi-Agent harness Manages people, workspaces, permissions, schedules, and collaboration

Model gateway → Agent runtime → Multi-Agent collaboration platform. The three layers can work together, but they do not replace one another.

🛠 Hands-on Exercises

Start with the four core exercises. Do not rename files or copy everything into a blank file first. Run the test directly, then change one small thing.

Core Exercise 1: Eval

Result: fixed cases and rules reveal which behavior regressed.

cd examples/stage-7/02-eval
python test.py

Core Exercise 2: Observability

Result: see the steps, latency, tokens, and errors in one run.

cd examples/stage-7/03-observability
python test.py

Core Exercise 3: Approval, Checkpoint, and Recovery

Result: a sensitive action stops at human approval; after a restart it resumes from a checkpoint, and the same idempotency key does not repeat the action.

cd examples/stage-7/06-safe-execution
python test.py

Core Exercise 4: Deploy

Result: wrap an Agent in an API with /health and /chat, then test its error states.

cd examples/stage-7/05-deploy
python test.py
🛠 Expand: exercise order, paid paths, and what to observe
  1. Run python test.py first in every folder. It uses mocks and needs no API key.
  2. Only after Eval, Observability, and Deploy tests pass, choose the local Ollama or Anthropic path in that folder’s README; Safe Execution uses fake actions throughout and needs no model.
  3. Change only one thing: a grading rule, trace field, approval result, corrupted-checkpoint case, or API error response.
  4. Run the test again. Record what changed, which result moved, and whether it stayed within budget.
  5. Docker in Core Exercise 4 is optional at first. Verify behavior with FastAPI tests before starting a service.

🧭 Advanced Options (Keep the Entrances Visible)

Option A: Multi-Agent Debate

Result: two Agents make independent cases and a third Agent judges them with a rule. Do this only after a single-Agent baseline has an Eval and the roles truly need to be separate.

Open the Multi-Agent example

Option B: Streaming and Prompt caching

Result: compare streaming and prompt caching; measure the cost effect yourself. Cache is not a safety or recovery mechanism.

Open the advanced SDK example

🧪 Expand: direct test commands for both options
cd examples/stage-7/01-multi-agent-debate
python test.py

cd ../04-sdk-advanced
python test.py

Start with a single-Agent version:

  1. Find three sources and keep their URLs and retrieval times.
  2. Write a short summary using only the sources; say clearly when you do not know.
  3. Stop before “publishing the summary” so a person can approve, edit, or reject it.
  4. Save a checkpoint and simulate a restart followed by resume.
  5. Use an idempotency key to prove that rerunning one publication writes only once.

Produce an execution receipt: task ID, Outcome, Trajectory, tools, sources, elapsed time, tokens, errors, checkpoint version, and human approval. Start with five development cases for the baseline, then add real failures to a versioned suite. If results worsen, rerun enough trials, check predefined thresholds, and inspect failures; one random failure alone does not establish a regression.

Only after the single Agent is stable, consider separating “find sources” and “review” roles. Compare quality, cost, and latency.

📊 Agent Benchmark Landscape: How to read it, not just the leaderboard + ⚠ Reward-Hacking Warning

A Benchmark is like a shared exam. It helps comparison, but it cannot promise that your real work will perform the same way.

Ask five questions before trusting a score:

Check Plain-language question
Task Does the exam resemble my real work?
Environment Which tools, data, and permissions did the model receive?
Grader Who scored it, and can the rule be exploited?
Trajectory Did it really solve the task, or only stumble into a score?
Hold-out Did it pass my own tests that were not used for tuning?

Reward hacking means “getting the score without achieving the real goal.” It is like a child learning that pressing a bell earns candy, then pressing the bell repeatedly instead of doing the assigned task.

📊 Expand: useful Benchmarks and production evaluation

Do not copy one SOTA score into the page as a permanent fact. Release decisions should use your own cases, rubric, complete trajectories, cost, and latency. Whenever you change the model, Prompt, Tool, or Harness, rerun the development/reference cases first; do not tune repeatedly on the frozen holdout. Open the holdout only for a release candidate or final validation.

Choose by purpose; ratings are not GitHub stars. Compare the two new three-star readings after a single-Agent baseline. Their ratings reflect documented teaching value; they haven't been run against live APIs.

CategoryProject / documentTeaching fitBest forKnow this first
Orchestration / WorkflowAnthropic — Building Effective Agents⭐⭐⭐⭐⭐Learn simple workflows before AgentsA design guide, not a deployable framework
OpenAI Agents SDK orchestration⭐⭐⭐⭐⭐Compare manager and handoff patternsExamples center on OpenAI Agents SDK
OpenAI Responses Multi-agent (official docs)⭐⭐⭐Readers with a single-Agent baseline: model-directed independent tasksBeta; separate contexts, shared model/tools; differs from SDK manager / handoff
Microsoft Agent Framework orchestrations⭐⭐⭐⭐Sequence, concurrency, handoff, group chat, and approvalConfirm current package version and preview status
LangGraph⭐⭐⭐⭐⭐State, checkpointing, and human-in-the-loopMore abstraction than a first Agent needs
Eval / ObservabilityAnthropic — Develop tests and evaluations⭐⭐⭐⭐⭐Define success criteria and gradersYou must supply cases that represent real work
promptfoo⭐⭐⭐⭐⭐Put Evals in CIA config file cannot replace a good rubric
OpenTelemetry GenAI conventions⭐⭐⭐⭐Learn portable trace fieldsThe conventions evolve and support varies
Langfuse⭐⭐⭐⭐⭐Tracing, Eval, and prompt managementSelf-hosting still needs operations and data governance
Arize Phoenix⭐⭐⭐⭐OpenTelemetry and local analysisDesign sensitive-data redaction first
Anthropic — Demystifying evals for AI agents⭐⭐⭐⭐⭐Check Outcome, Trajectory, and graders togetherBuild cases from your own real work and failures
Harness / Sandbox / DeployClaude Agent SDK Python⭐⭐⭐⭐⭐Read tool loops, permissions, and subagent codeCenters on the Claude runtime
Google Antigravity agent (official docs)⭐⭐⭐Readers with a single-Agent baseline: sandbox, persistent files, and codePublic Preview; restrict network and tool permissions to manage safety
DeepSeek Harness⭐⭐⭐Read a plugin-based harness architectureDeveloper preview; breaking changes are possible
OpenAI Agents SDK — Human-in-the-loop⭐⭐⭐⭐⭐Pause sensitive tools, save RunState, and resumeSaved state may contain context and runtime metadata; manage it as sensitive data
LangGraph — Interrupts⭐⭐⭐⭐⭐Approval, checkpoints, resume, and idempotent side effectsProduction needs a durable checkpointer, not only memory
SandBase Harness⭐⭐⭐⭐See how a self-hosted runtime saves work, connects MCP, waits for approval, and keeps audit / replay recordsStill v0.x; isolation depends on the local, Docker, Kubernetes, or Worker backend and its deployment, not a fixed microVM guarantee
BentoML⭐⭐⭐⭐Package an application as a service and containerA deployment framework does not add Evals or Guardrails for you
Multi-Agent CasescrewAI⭐⭐⭐⭐Understand role-based task divisionMore roles do not guarantee a better answer
Orca⭐⭐⭐⭐Run coding Agents in isolated worktreesA person must still review and select parallel results
QM⭐⭐⭐⭐Study team workspaces, permissions, and schedulesOrganization-wide deployment is more complex than a personal CLI
LongHorizon-Harness⭐⭐⭐See Manager / Executor / Auditor rolesVery new, with limited long-term maintenance history
Edict⭐⭐⭐Learn planning, review, and execution roles from a Chinese-language caseIts special role names are a case design, not an industry standard

Existing resources reviewed: 2026-09-13 UTC; new docs reviewed: 2026-10-02 UTC

✅ Self-Check After Stage 7

  • I can distinguish Outcome and Trajectory and use both to check the same case.
  • I have fixed Eval cases built from real failures instead of one attractive output.
  • I can find the trace, error, latency, and token count for one run.
  • Risky tools have least privilege and human approval; without approval they fail closed.
  • I can resume from a checkpoint and prove the same idempotency key does not duplicate a side effect.
  • I can show an execution receipt and explain when to stop, recover, or roll back.
  • I can distinguish OpenRouter, an Agent runtime, and a Multi-Agent platform in one sentence, and I know a single Agent is the default.

Next, go to Stage 7.5 — Advanced Agentic Concept Map, then Stage 8 — Agent Interfaces. If one item is still unclear, return to its exercise, change one thing, and test again.