Stage 7 — Agent Production Engineering: Testable, Observable, Stoppable, and Recoverable¶
Make an AI helper testable, observable, stoppable, and recoverable before sharing it.
🎯 What This Stage Does (Start Here)¶
Agent Production Engineering means making an Agent testable, observable, stoppable, and recoverable. Like a toy car going onto a road, it needs steering, brakes, and a dashboard. It need not serve millions.
The chapter uses one story: a research helper checks three sources, writes a summary, and asks before sending.
Remember this order:
Say what success looks like → keep a record of the work → ask a person before risky actions → prove it can continue after a fall → only then give it to other people.
| Where you are stuck | Do this first | Evidence to produce |
|---|---|---|
| You do not know whether the summary is good | Write fixed examples and success rules | A check you can run again |
| You do not know which step failed | Record every step, error, time, and cost | One complete work record |
| It can send mail, pay, delete, or write data | Stop before the action, ask a person, and save progress | Who approved it and where to continue |
| The first three checks can rerun and pass | Give it to other people | Health, stop, continue-after-failure, and rollback instructions |
Make one AI helper reliable first. Add helpers only for separable work or independent checks.
⏱ Expand: time, environment, cost, and safety notes
- Split this stage into several short practice sessions. You do not need to finish it at once.
- You need Python and Git. The deployment exercise also uses Docker.
- Run the tests that need no API key first. Set a small budget before calling a paid model.
- A work record may contain prompts, tool inputs, and model answers. Do not send passwords, personal data, or customer data directly to a tracing service.
- Another Agent usually adds another model call, more latency, and more debugging. Do not assume Multi-Agent is automatically faster or more accurate.
📌 Learning Goals¶
After this stage, you can:
- Tell apart the AI helper's workspace, its repeat-and-check rhythm, and its branching route.
- Turn real failures into checks you can run again instead of trusting one pretty answer.
- Find every step, error, time, and cost from one task.
- Stop before risky actions and continue from the right saved point.
- Use the same evidence to decide whether the system is ready for other people.
🧩 Meet Nineteen Core Terms First¶
Use “plain-language meaning” to get oriented, then read “how this chapter uses it / technical boundary” to see what the term means here. Related terms share one group, so you do not need to read the same definition twice.
| What to solve first | Core term | Plain-language meaning | How this chapter uses it / technical boundary |
|---|---|---|---|
| Make the task run | Agent Harness | The room where the AI helper works | The execution environment that holds the model, tools, permissions, state, error handling, and records; this chapter uses it to check sources and prepare a summary safely |
| Agent Loop | Do one step, see the result, then choose the next step | The model repeatedly chooses an action and reads the tool result until it finishes, reaches a limit, or must ask a person | |
| Workflow Graph | A route map with forks | Steps, connections, conditions, and state that say which route to take; this chapter uses it for research, checking, and approval before sending | |
| Orchestration | Arrange who goes first and who goes next | Control steps, data flow, roles, retries, and stop conditions | |
| Multi-Agent | Several AI helpers share the work | Multiple Agents complete a task with clear roles; it is an option, not a requirement | |
| Handoff | Pass the baton and the notes together | One Agent passes control, needed data, and result evidence to another Agent | |
| Prove it did the right thing | Evaluation / Eval | Use the same checklist each time | Measure an Agent's result and process with fixed cases, environments, grading methods, and thresholds |
| Outcome | What really happened at the end | The externally verifiable state when the task ends; this chapter checks that the summary truly uses three valid sources | |
| Trajectory | The footprints left along the way | What happened during one run, including tool calls, intermediate results, errors, and output | |
| Grader | Mark one answer using stated rules | A method, program, or model that scores one Eval Case against success criteria; this chapter keeps the rules and human spot checks visible | |
| Evaluation Harness | The exam room that gives the same test and keeps the score | A test system that loads cases, reruns the Agent, calls graders, and saves results; it has a different responsibility from the Agent Harness used for daily work | |
| Trace | A notebook that collects the footprints | Steps, tool calls, errors, and results arranged by time for one task; this chapter uses it to find the failing step | |
| Observability | Put a clear window on the system | Use traces, logs, and metrics to see internal state; this chapter uses it to find where a source went missing | |
| Let it stop and continue safely | Guardrail | Block what must not happen | Rules that limit inputs, outputs, tool permissions, or risky actions |
| Human Approval | Ask a person before a risky action | Pause before a sensitive tool call so a person can approve, edit, or reject it | |
| Checkpoint | Save before moving on | Save recoverable workflow state and version information | |
| Resume | Continue from the saved point | Load a checkpoint with the same task or thread ID and continue execution | |
| Recovery | Come back safely after falling | A strategy to stop, retry, compensate, or hand a failure to a person | |
| Idempotency | Press twice, do it once | Retries with the same idempotency key do not duplicate external side effects |
A Prompt is the instruction and material you give the model. Context is the information needed for this step. They still matter; this chapter adds execution, checking, and recovery around them.
🚪 Entry Conditions¶
You should have completed at least:
- Stage 4: know what Agents, Tools, and Workflows are.
- Stage 5: have seen tool permissions, Subagents, and development workflows.
- Stage 6: know that Context, RAG, and Memory are different.
You can start even if Docker is new to you. Do the four core exercises first and learn Docker for Core Exercise 4.
📚 Required Reading¶
Read these six in production order:
- Anthropic — Demystifying evals for AI agents: distinguish Outcome from the complete Trajectory; an Agent saying “done” does not prove the outside result is done.
- OpenAI Agents SDK — Tracing: see how trace, span, tool, handoff, and guardrail events connect one run.
- OpenAI Agents SDK — Human-in-the-loop: pause before a sensitive tool, save
RunState, approve or reject, and resume. - LangGraph — Persistence: distinguish checkpoints from cross-thread stores; interruption, recovery, and long-term memory are not the same thing.
- LangGraph — Interrupts: see how human approval pauses and resumes, and why side effects before an interrupt must be idempotent.
- Anthropic — Building Effective Agents: start with simple compositions and add autonomy or Multi-Agent only when real division of labor needs it.
📖 Expand: further reading and purpose
- Anthropic — Develop tests and evaluations: define measurable success criteria before choosing a grader.
- OpenAI Agents SDK — Testing utilities: test with repeatable fake models instead of paying for every run.
- OpenAI Agents SDK — Running agents: see an Agent Loop repeat, and stop it with
max_turns. - OpenAI Agents SDK — Multi-agent orchestration: compare manager and Handoff patterns; this is an advanced option, not the first production step.
- LangGraph — Workflows and agents: distinguish fixed Workflows from Agents that choose their next step.
- Microsoft Agent Framework — Workflow concepts: see how executors, edges, events, and state form a Workflow Graph.
- OpenAI — Harness engineering: see how environments, feedback loops, and mechanical rules help Agents work reliably.
- OpenTelemetry — GenAI semantic conventions: learn portable tracing fields; the conventions are evolving, so do not assume every platform supports all of them.
🧭 Harness, Loop, Graph, and Eval: How They Work Together¶
They are not four product generations, and you do not choose only one. Think about the same research helper from four angles:
| Responsibility | Plain-language question | Research-helper example |
|---|---|---|
| Agent Harness | Where can it work safely? | It may read sources, but it must stop before sending |
| Agent Loop | Why should it take another round? | If one source is missing, search again; stop at the limit |
| Workflow Graph | Which route should it take now? | Go back to research when sources are weak; otherwise ask for approval |
| Eval | How do I know the result and process are acceptable? | Check three valid sources, correct citations, and no skipped approval |
Eval checks the Outcome and Trajectory, then uses a Grader to decide whether the run meets the stated rules. Eval can make the Loop retry, make the Graph choose another route, or make the Harness stop. Putting a Harness and Eval together still does not create the Loop's rules for repeating and stopping.
The learning order is Stage 3 Agent Loop → Stage 4 Workflow Graph / Agent Framework → safe production integration in this chapter. Loop Engineering is an emerging label used by IBM. Graph Engineering is even less settled. Learn the responsibilities first, then treat these labels as search terms used by the community. Sources: IBM — Loop Engineering, Anthropic — Agent harness and eval, and Microsoft Agent Framework — graph-based workflows.
🏗 Agent Harness: Prepare the Safe Workspace¶
Harness Engineering means designing the runtime that lets a model act as an Agent. The model produces decisions; the Harness processes input, tools, state, permissions, failures, and results, and it often runs the agent loop directly. An outer scheduler may call the Harness many times, so a Harness is not limited to “one short run.” Sources: OpenAI — Harness engineering, Anthropic — agent harness definition, and Anthropic — Managed agents.
The 8 Core Components of a Harness¶
These eight items are this project’s production checklist, not the world’s only official taxonomy.
| Component | Plain-language meaning | Question before release |
|---|---|---|
| 1. Orchestration / Run loop | Decide what happens next | Who starts, who stops, and what if a handoff fails? |
| 2. Tool / Permission boundary | Give it only the keys it needs | Which tools may read, write, or delete? |
| 3. Context / State / Checkpoint | Save where it is now | Can it resume from the correct point? |
| 4. Retry / Recovery / Idempotency | Try again without charging twice | Could a retry repeat an email, payment, or database write? |
| 5. Guardrail / Human approval | Ask an adult before a risky action | Which actions always require approval? |
| 6. Telemetry / Observability | Put a clear window on the system | Can we see traces, errors, latency, and tokens? |
| 7. Eval harness | Retake the test after every change | Are cases, scoring rules, and failure thresholds fixed? |
| 8. Cost / Latency budget | Decide the money and time limit first | Above budget, should it stop, downgrade, or queue? |
🔧 Expand: feedback, recovery, and cost details
- Write tool errors as feedback an Agent can understand, not only a long stack trace.
- Keep the grader separate from the worker when possible. Do not ask only, “How good was your own work?”
- Design idempotency for every external side effect so a retry does not repeat a payment, email, or data write.
- Prompt caching, batching, model routing, and smaller models may reduce cost, but results depend on the workload. Measure a baseline, change one thing, and measure again.
- Anthropic prompt caching can be automatic or use explicit
cache_control. Cache duration and read/write pricing depend on the option; use the official documentation. - Traces may capture sensitive inputs and outputs. Configure redaction, retention, and access before release.
🔁 Agent Loop: Act, Observe, Then Decide¶
First separate three kinds of Loop that sound similar but cover different scopes:
| Name | What repeats | Example |
|---|---|---|
| Program loop | The same code block | for item in items; this is syntax, not the topic of this section |
| Agent Loop | Model → tool → tool result → model | Keep calling tools inside one run until completion or max_turns |
| Loop Engineering | Goal → action → observation → adjustment | Repeat work in one long run or across sessions/schedules, with verification, memory, budgets, and stop conditions on every round |
IBM explains Loop Engineering as Goal → Action → Observation → Adjustment. The point is not to let an Agent run forever. Every round must answer: Is the goal still valid? Is the evidence enough? Should we continue, stop, or hand control to a person?
Therefore, Loop Engineering is not the next Harness product generation and does not automatically replace a Harness. In Anthropic's terminology, the Harness itself includes the loop that calls the model and routes tools. IBM's broader Loop Engineering practice includes goals, checking, tools, hooks, context, subagents, and persistent state. Sources draw the boundary differently, so remember the responsibilities instead of memorizing one universal layer diagram. Sources: IBM — Loop Engineering and Anthropic — Managed Agents.
When models improve, a particular patch may be removable. For example, Anthropic removed context resets used by an earlier harness for newer models. That only means the same Evals showed one workaround was no longer needed. It does not mean permissions, safety, logs, evals, or recovery automatically became obsolete. Source: Anthropic — Harness design for long-running applications. Stage 7.5 explains how to keep, simplify, or remove each part under Model–Harness Fit.
🗺 Workflow Graph: Know Where to Go at a Branch¶
The Agent Loop answers, “Should I do this again?” The Workflow Graph answers, “Where should I go next?” It is like a school map with forks. The map arranges the route, but it does not do the work at each stop.
Outside writing sometimes calls this engineering work Graph Engineering. The label is emerging. Official documents often call each stop a node, each connection an edge, and each choice a branch. Learn nodes, edges, branches, cycles, state, checkpoints, and human approval instead of memorizing a label that is not yet standardized.
A node can contain an Agent Loop; the Workflow Graph arranges the order between nodes.
🧠 Expand: choosing a Loop, Graph, or Multi-Agent design
- Use a Loop when there is one path that may need many retries.
- Use a Graph / Workflow when there are branches, parallel steps, human approvals, or a need to resume in the middle.
- Add Multi-Agent only when parts can truly work independently or distinct roles must check one another.
-
A Graph node can be an Agent, a tool, fixed code, or “wait for human approval.” Not every box needs an Agent.
-
Optional official docs: OpenAI Responses Multi-agent is in Beta for GPT-6.1 Sol and all GPT-5.6 models. The model delegates to subagents with separate contexts. They share the request's model and tools; this is not SDK manager / handoff orchestration.
max_concurrent_subagentsdefaults to 3 active subagents across the tree and excludes the root. Concurrency settings, total agents, and tree depth have no fixed cap; delegation can add tokens.max_tool_calls,reasoning.summary, and/responses/compactare unsupported; server-side automatic compaction runs independently for each Agent.- The API executes hosted collaboration; your application executes custom function calls. Separate contexts do not isolate tool permissions: the application must still approve sensitive tools and enforce budgets and stop conditions.
- Google Managed Agents offers Antigravity in Public Preview.
antigravity-preview-09-2026defaults to Gemini 3.8 Flash. It provides a managed Linux sandbox, files preserved across interactions, code execution, custom functions, and remote MCP. - Outbound network access is unrestricted by default; set an allowlist and minimal tool permissions. Search and URL fetching do not imply GUI browser control;
computer_useis unsupported. A sandbox still needs this chapter's Evals, approval, and recovery. - Google's docs describe referencing secrets by managed credential ID: the egress proxy injects them without exposing them in the sandbox. An Agent can use the full scope of a supplied credential; grant only the minimum scope needed.
🧪 Eval: State What Good Means, Then Decide How to Grade¶
Eval is not one score or a report added after the system is done. First say what success means. Then use the same method to compare versions.
Start with the Outcome. The research helper's Outcome is not “the Agent says the summary is done.” It is “the summary really uses three valid sources, every citation opens, and human approval has not been skipped.”
Next, create a complete Eval Case. It is like an exam question that also includes its rules. Input is only one part.
| Part of an Eval Case | Research-helper example | Why keep it |
|---|---|---|
| Input | Summarize these three topics |
Tells the system what to do |
| Initial State | Three candidate sources; not yet approved | Fixes the starting environment |
| Success Criteria | All three sources open; the summary has checkable citations | Says what success means |
| Forbidden Actions | Do not invent a source; do not send by itself | Blocks bad behavior even when the answer looks good |
| Optional Reference Answer | A summary checked by a person | Gives comparison guidance when useful; not every case needs one |
| Grader | Code checks links and counts; a person checks faithfulness | Says who grades with which rule |
| Case Metadata | Case ID, version, split, source, and labels | Makes the same case rerunnable and traceable |
Several complete cases form an Eval Suite. Version the Suite so you know whether this run and the last run used the same exam.
This project calls a human-reviewed, reusable collection of complete cases a Reviewed Eval Set. Other sources may say Golden Set or Reference Set, but these labels do not have one shared definition across vendors. Check whether a source means questions, answers, criteria, or the whole collection.
A Golden / Reference Set is not input alone, and it is not the same as training data or few-shot examples. It usually contains complete cases, conditions, reference evidence, and grading methods. Its exact fields still come from the current project's definition.
Add these measurement terms last:
| Term | Plain-language meaning | How this chapter uses it |
|---|---|---|
| Trial | One actual attempt at one case | Run a case more than once when model output can vary |
| Baseline | Measure before changing anything | Gives old and new versions the same starting point |
| Regression | The new version becomes worse past a preset limit | Check quality, cost, safety, and reliability together |
| Development Set | Practice questions you may look at | Rerun after changes and use failures to improve the system |
| Holdout Set | The final exam you do not peek at | Open only for a release candidate or final validation |
Anthropic's Agent Eval guide separates task, trial, grader, trajectory, and outcome. OpenAI's Graders API lists several grader types. Tools may differ, but each report should keep dataset version, split, case ID, trial count, grader, Outcome, Trajectory, and baseline.
The system that loads cases, reruns an Agent, calls a grader, and saves results is the Evaluation Harness defined earlier. It can call an Agent Harness, but their jobs are different: one runs work safely, and one makes tests repeatable and comparable.
🔎 Observability: See Which Step Failed¶
Observability is like looking through the clear wall of a transparent box. It does not mean showing everything to everyone. It means keeping useful traces, logs, and metrics while hiding passwords, personal data, and customer data.
For the research helper, one Trace should answer: Which sources were checked? Which tool failed? How many retries happened? How long did it take? Why did it stop before approval? A Trace helps explain the Trajectory, but “we recorded a lot” does not mean the result is correct. Eval still checks the Outcome.
🛑 Approval, Checkpoint, Resume, and Recovery: Stop, Then Continue Safely¶
When the research helper is ready to send, it enters Human Approval. The person does not start from zero. They receive the summary, sources, and risks together.
Save a Checkpoint before approval. After a restart, Resume uses the same task ID to return to that point. If something failed, Recovery decides whether to retry, compensate, restore an older state, or hand the task to a person. Any action that sends mail, pays, or writes data needs an Idempotency key so the same retry produces the outside effect only once.
🛡 Complete Production Route: Eval → Observability → Approval / Recovery → Deploy¶
Deploy means giving a checked system to other people. It is like opening a shop. Opening the door is not proof of success; the tests, records, brakes, and recovery plan come first.
These four steps are not maturity badges; they are the check route for the same change:
| Order | Question to answer | Minimum evidence to leave | What to do if it fails |
|---|---|---|---|
| 1. Eval | Is the final result really correct? Did it take a dangerous shortcut? | Anthropic suggests starting with 20–50 cases representing real work as a practical range, not a universal minimum. Also record Outcome, Trajectory, grader, cost, and a failure threshold | Add cases or fix behavior; do not deploy |
| 2. Observability | Can you find the failed step? | Task ID, trace/span, tool call, error type, latency, tokens, and sensitive-data redaction | Make failures visible before changing the Prompt or model |
| 3. Approval / Recovery | Can a risky action stop first? Can it resume safely after interruption? | Human approval point, versioned checkpoint, resume test, idempotency key, and reject/timeout/compensation route | Fail closed, stop automation, and hand it to a person |
| 4. Deploy | Can the first three steps rerun on the new version? | Health/readiness, rate limit, rollback, stop switch, version, and release record | Keep the old version or roll back; “the service started” is not success |
An Outcome Eval checks the outside-world result. For example, an Agent saying “the email was sent” is only text; the test environment having exactly one email sent to the right recipient is an Outcome pass. A Trajectory Eval checks which tools it used, how many attempts it made, whether it bypassed approval, and how many tokens it spent. Use both so a polished final sentence cannot pass by itself.
Build cases from real failures first: for each error, keep a de-identified input, expected Outcome, forbidden actions, and reproduction steps. Do not copy production data into a public repo; use structurally equivalent fake data when needed.
🧭 What Is the Difference Between OpenRouter, Pi, OpenCode, Orca, and QM?¶
They are not five versions of the same product. Put each one at the right layer:
| Name | What it is | One-line memory aid |
|---|---|---|
| OpenRouter | Model API gateway / router | Connects software to different models; it is not a coding Agent |
| Pi | Agent toolkit and coding-agent CLI | Calls models and tools to finish a task |
| OpenCode | Open-source coding Agent | Reads, edits, and tests inside a code project |
| Orca | Multi-Agent development environment | Runs coding Agents in isolated worktrees for comparison |
| QM | Team Multi-Agent harness | Manages people, workspaces, permissions, schedules, and collaboration |
Model gateway → Agent runtime → Multi-Agent collaboration platform. The three layers can work together, but they do not replace one another.
🛠 Hands-on Exercises¶
Start with the four core exercises. Do not rename files or copy everything into a blank file first. Run the test directly, then change one small thing.
Core Exercise 1: Eval¶
Result: fixed cases and rules reveal which behavior regressed.
cd examples/stage-7/02-eval
python test.py
Core Exercise 2: Observability¶
Result: see the steps, latency, tokens, and errors in one run.
cd examples/stage-7/03-observability
python test.py
Core Exercise 3: Approval, Checkpoint, and Recovery¶
Result: a sensitive action stops at human approval; after a restart it resumes from a checkpoint, and the same idempotency key does not repeat the action.
cd examples/stage-7/06-safe-execution
python test.py
Core Exercise 4: Deploy¶
Result: wrap an Agent in an API with /health and /chat, then test its error states.
cd examples/stage-7/05-deploy
python test.py
🛠 Expand: exercise order, paid paths, and what to observe
- Run
python test.pyfirst in every folder. It uses mocks and needs no API key. - Only after Eval, Observability, and Deploy tests pass, choose the local Ollama or Anthropic path in that folder’s README; Safe Execution uses fake actions throughout and needs no model.
- Change only one thing: a grading rule, trace field, approval result, corrupted-checkpoint case, or API error response.
- Run the test again. Record what changed, which result moved, and whether it stayed within budget.
- Docker in Core Exercise 4 is optional at first. Verify behavior with FastAPI tests before starting a service.
🧭 Advanced Options (Keep the Entrances Visible)¶
Option A: Multi-Agent Debate¶
Result: two Agents make independent cases and a third Agent judges them with a rule. Do this only after a single-Agent baseline has an Eval and the roles truly need to be separate.
Option B: Streaming and Prompt caching¶
Result: compare streaming and prompt caching; measure the cost effect yourself. Cache is not a safety or recovery mechanism.
🧪 Expand: direct test commands for both options
cd examples/stage-7/01-multi-agent-debate
python test.py
cd ../04-sdk-advanced
python test.py
🧪 Recommended Mini-Project: A Research Assistant with a Receipt¶
Start with a single-Agent version:
- Find three sources and keep their URLs and retrieval times.
- Write a short summary using only the sources; say clearly when you do not know.
- Stop before “publishing the summary” so a person can approve, edit, or reject it.
- Save a checkpoint and simulate a restart followed by resume.
- Use an idempotency key to prove that rerunning one publication writes only once.
Produce an execution receipt: task ID, Outcome, Trajectory, tools, sources, elapsed time, tokens, errors, checkpoint version, and human approval. Start with five development cases for the baseline, then add real failures to a versioned suite. If results worsen, rerun enough trials, check predefined thresholds, and inspect failures; one random failure alone does not establish a regression.
Only after the single Agent is stable, consider separating “find sources” and “review” roles. Compare quality, cost, and latency.
📊 Agent Benchmark Landscape: How to read it, not just the leaderboard + ⚠ Reward-Hacking Warning¶
A Benchmark is like a shared exam. It helps comparison, but it cannot promise that your real work will perform the same way.
Ask five questions before trusting a score:
| Check | Plain-language question |
|---|---|
| Task | Does the exam resemble my real work? |
| Environment | Which tools, data, and permissions did the model receive? |
| Grader | Who scored it, and can the rule be exploited? |
| Trajectory | Did it really solve the task, or only stumble into a score? |
| Hold-out | Did it pass my own tests that were not used for tuning? |
Reward hacking means “getting the score without achieving the real goal.” It is like a child learning that pressing a bell earns candy, then pressing the bell repeatedly instead of doing the assigned task.
📊 Expand: useful Benchmarks and production evaluation
- SWE-bench: real software issues.
- Terminal-Bench: terminal tasks.
- OSWorld: desktop-environment tasks.
- τ²-bench: tasks with tools and multi-turn interaction.
- GAIA: general-assistant tasks.
Do not copy one SOTA score into the page as a permanent fact. Release decisions should use your own cases, rubric, complete trajectories, cost, and latency. Whenever you change the model, Prompt, Tool, or Harness, rerun the development/reference cases first; do not tune repeatedly on the frozen holdout. Open the holdout only for a release candidate or final validation.
🎯 Featured Projects (Templates / SDKs / Tool Collections)¶
Choose by purpose; ratings are not GitHub stars. Compare the two new three-star readings after a single-Agent baseline. Their ratings reflect documented teaching value; they haven't been run against live APIs.
| Category | Project / document | Teaching fit | Best for | Know this first |
|---|---|---|---|---|
| Orchestration / Workflow | Anthropic — Building Effective Agents | ⭐⭐⭐⭐⭐ | Learn simple workflows before Agents | A design guide, not a deployable framework |
| OpenAI Agents SDK orchestration | ⭐⭐⭐⭐⭐ | Compare manager and handoff patterns | Examples center on OpenAI Agents SDK | |
| OpenAI Responses Multi-agent (official docs) | ⭐⭐⭐ | Readers with a single-Agent baseline: model-directed independent tasks | Beta; separate contexts, shared model/tools; differs from SDK manager / handoff | |
| Microsoft Agent Framework orchestrations | ⭐⭐⭐⭐ | Sequence, concurrency, handoff, group chat, and approval | Confirm current package version and preview status | |
| LangGraph | ⭐⭐⭐⭐⭐ | State, checkpointing, and human-in-the-loop | More abstraction than a first Agent needs | |
| Eval / Observability | Anthropic — Develop tests and evaluations | ⭐⭐⭐⭐⭐ | Define success criteria and graders | You must supply cases that represent real work |
| promptfoo | ⭐⭐⭐⭐⭐ | Put Evals in CI | A config file cannot replace a good rubric | |
| OpenTelemetry GenAI conventions | ⭐⭐⭐⭐ | Learn portable trace fields | The conventions evolve and support varies | |
| Langfuse | ⭐⭐⭐⭐⭐ | Tracing, Eval, and prompt management | Self-hosting still needs operations and data governance | |
| Arize Phoenix | ⭐⭐⭐⭐ | OpenTelemetry and local analysis | Design sensitive-data redaction first | |
| Anthropic — Demystifying evals for AI agents | ⭐⭐⭐⭐⭐ | Check Outcome, Trajectory, and graders together | Build cases from your own real work and failures | |
| Harness / Sandbox / Deploy | Claude Agent SDK Python | ⭐⭐⭐⭐⭐ | Read tool loops, permissions, and subagent code | Centers on the Claude runtime |
| Google Antigravity agent (official docs) | ⭐⭐⭐ | Readers with a single-Agent baseline: sandbox, persistent files, and code | Public Preview; restrict network and tool permissions to manage safety | |
| DeepSeek Harness | ⭐⭐⭐ | Read a plugin-based harness architecture | Developer preview; breaking changes are possible | |
| OpenAI Agents SDK — Human-in-the-loop | ⭐⭐⭐⭐⭐ | Pause sensitive tools, save RunState, and resume | Saved state may contain context and runtime metadata; manage it as sensitive data | |
| LangGraph — Interrupts | ⭐⭐⭐⭐⭐ | Approval, checkpoints, resume, and idempotent side effects | Production needs a durable checkpointer, not only memory | |
| SandBase Harness | ⭐⭐⭐⭐ | See how a self-hosted runtime saves work, connects MCP, waits for approval, and keeps audit / replay records | Still v0.x; isolation depends on the local, Docker, Kubernetes, or Worker backend and its deployment, not a fixed microVM guarantee | |
| BentoML | ⭐⭐⭐⭐ | Package an application as a service and container | A deployment framework does not add Evals or Guardrails for you | |
| Multi-Agent Cases | crewAI | ⭐⭐⭐⭐ | Understand role-based task division | More roles do not guarantee a better answer |
| Orca | ⭐⭐⭐⭐ | Run coding Agents in isolated worktrees | A person must still review and select parallel results | |
| QM | ⭐⭐⭐⭐ | Study team workspaces, permissions, and schedules | Organization-wide deployment is more complex than a personal CLI | |
| LongHorizon-Harness | ⭐⭐⭐ | See Manager / Executor / Auditor roles | Very new, with limited long-term maintenance history | |
| Edict | ⭐⭐⭐ | Learn planning, review, and execution roles from a Chinese-language case | Its special role names are a case design, not an industry standard |
Existing resources reviewed: 2026-09-13 UTC; new docs reviewed: 2026-10-02 UTC
✅ Self-Check After Stage 7¶
- I can distinguish Outcome and Trajectory and use both to check the same case.
- I have fixed Eval cases built from real failures instead of one attractive output.
- I can find the trace, error, latency, and token count for one run.
- Risky tools have least privilege and human approval; without approval they fail closed.
- I can resume from a checkpoint and prove the same idempotency key does not duplicate a side effect.
- I can show an execution receipt and explain when to stop, recover, or roll back.
- I can distinguish OpenRouter, an Agent runtime, and a Multi-Agent platform in one sentence, and I know a single Agent is the default.
Next, go to Stage 7.5 — Advanced Agentic Concept Map, then Stage 8 — Agent Interfaces. If one item is still unclear, return to its exercise, change one thing, and test again.