A3 — Connect a CLI agent to a safe team workflow¶
← Stage 5 — Track A core · Track A: CLI Power User Stop 3 (final core stop)
This stop has one goal: have a CLI agent perform a read-only check on a test PR. It may give feedback, but it must not merge, deploy, or obtain extra permissions by itself.
📌 Learning Goals¶
After finishing, you can:
- Give an MCP server only one safe scope.
- Have CI automatically produce a reviewable suggestion on a PR.
- Use Observability to understand the usage, time, and result left by one run.
- Hand the A2 Skill to a teammate and let them rerun it safely.
🧩 Three Core Terms First¶
| Core term | What it is, in plain language | How A3 uses it | What it is not |
|---|---|---|---|
| MCP (Model Context Protocol) | A standard adapter that connects an agent to external tools or data | Give a server only one demo folder or read-only tool | Not automatically safe; permissions still decide what it can touch |
| CI (Continuous Integration) | A checkpoint that runs automatically when a push or PR appears | Run one read-only review automatically on a test PR | Not an auto-merge button that skips human review |
| Observability | A receipt plus dashcam recording that keeps what happened | Record provider, model, usage, time, result, and failure reason | Not just one total-token number or a guess about unavailable cost |
The three terms appear together but are not the same thing: MCP connects tools, CI decides when to run automatically, and observability records the evidence after a run.
Follow the safety ladder first¶
- Read-only: let the agent see data first, without letting it change data.
- Least privilege: open only the folder, repo, tool, or token scope needed for this task.
- Demo repo: test first in a disposable practice environment.
- Human review: a person decides whether to use the agent’s suggestion.
- Only then consider writes: auto-merge, push, and deploy are outside this stop.
Expand for time, prerequisites, environment, and cost
- Time: finish the four smallest outcomes first. You can usually split them into several short practices; do not connect many services at once just to save time.
- Prerequisites: complete A1, A2, and the Stage 5 Track A core, sections 5.1–5.4, and be able to recognize the basic screens for
git status, PRs, and GitHub Actions. - Environment: a demo repo with no real secrets; use a GitHub-hosted Linux runner for the first round because a sandbox is easier to apply there.
- Cost: GitHub Actions, a CLI subscription, and model APIs may be billed separately. Check your own plan before running; do not treat someone else’s prices as yours.
If A2’s review-changes Skill cannot yet reliably output PASS or concrete issues, fix that first before starting A3.
📚 Required Reading¶
Required reading and learning resources checked: 2026-08-27 UTC
- First read MCP Connect to local servers to learn that a server can receive only the paths you give it.
- Then read GitHub Actions Security Hardening to understand least privilege and untrusted PRs first.
- Choose one CI path:
- Claude Code: official GitHub Actions documentation
- Codex: official GitHub Action documentation
- When you need trace, eval, or complete production theory, continue to Stage 7 and Stage 7.5.
🛠 Hands-on Exercises¶
Hands-on exercise CLI-9: Connect only one MCP server¶
Outcome: the agent can read a newly created demo folder, but it has not received access to your entire home directory, disk, real project, or secrets.
First copy the command for your computer to create a3-mcp-demo/hello.txt.
PowerShell:
New-Item -ItemType Directory -Force -Path a3-mcp-demo | Out-Null
Set-Content -LiteralPath a3-mcp-demo/hello.txt -Value 'hello from A3'
macOS/Linux:
mkdir -p a3-mcp-demo
printf 'hello from A3\n' > a3-mcp-demo/hello.txt
When connecting the official filesystem reference server to your CLI, pass only the absolute path to this folder.
When it works, the agent can read hello.txt; when asked to read a file outside the allowed scope, it should fail or ask you to grant authorization again.
Expand for CLI-9 installation, permission tests, and the GitHub MCP extension
- Open the configuration using the official MCP documentation for your main CLI; configuration files and commands differ between CLIs.
- Use the official package
@modelcontextprotocol/server-filesystem; its arguments should contain only the absolute path toa3-mcp-demo. Do not enter~, your home directory, the disk root, or the whole workspace. - Restart the CLI, ask it to list the demo folder, and then read
hello.txt. - Ask it to read an ordinary filename outside the demo scope. The correct result is a refusal or a request to add authorization; it must not read the file secretly.
- Remove the server configuration after practice and confirm that the CLI can no longer use it.
To read PRs or issues, use GitHub’s official github/github-mcp-server instead. Start with --read-only, then use toolsets or a tools allow-list to open only the capabilities you need. If you use a PAT, put it in a secure secret or environment variable, grant the smallest scope, and revoke it after practice; when OAuth is available, configure it through the host’s official process.
modelcontextprotocol/servers is useful for reading reference implementations, but its official description says they are not production-ready. The old github reference server has moved to the historical collection; do not use it as the current GitHub entry point.
Cost reminder: a local filesystem server usually has no separate charge, but the CLI or model may still cost money. A remote MCP may have its own plan too.
Hands-on exercise CLI-10: Give a PR one more read-only checker¶
Outcome: the test PR gets a review result; a person still decides whether to edit, merge, or deploy.
Choose Anthropic’s claude-code-action or OpenAI’s codex-action. For the first round, run it only in a demo repo and branch you control, reusing A2’s review-changes Skill.
The success standard is not “finish within a few minutes.” It is that the workflow finishes successfully and leaves a readable result in a PR comment, job summary, or artifact.
Expand for CLI-10 security settings and validation steps
- Build the workflow from the provider’s official example; do not copy YAML from an unknown source.
- Put the API key in a GitHub Actions secret. Do not write it into the workflow, prompt, repo, or log.
- Start
GITHUB_TOKENatcontents: read. Add only the necessary pull-request permission to that job when it needs to post a PR comment. - For Codex read-only work, use the currently supported
permission-profile: ":read-only"setting in the official action; do not also set mutually exclusive legacy sandbox fields. For Claude Code, restrict capabilities through the official action’s permissions and allowed tools. - Make the prompt ask only to read the diff, list issues, and output
PASSor concrete suggestions. State explicitly: do not edit, commit, push, merge, deploy, or send extra messages. - Start with a same-repo test branch that you created yourself. Do not use
pull_request_targetto check out untrusted PR code; this can expose secrets or write permissions to untrusted content. - Check the Actions log, review result, and repo diff. If there is any sign of secret leakage, immediately delete the log, revoke the secret, and rotate it.
GitHub recommends pinning third-party Actions in production workflows to a full commit SHA because a tag can move. The @v1 or @v5 forms in official documentation are useful for identifying product versions; before production use, verify and pin the trusted full SHA for that time.
Cost reminder: set a job timeout and concurrency to avoid hangs or duplicate triggers. Keep model APIs, provider plans, and GitHub Actions minutes separate.
Hands-on exercise CLI-11: Read the receipt for one run¶
Outcome: you record the provider/model, input usage, output usage, time, and result; fields you cannot obtain are clearly marked “unconfirmed” instead of guessed.
First distinguish whether you use a subscription plan or pay by API usage. When the official source provides token counts and prices, calculate cost only with this formula:
input tokens × input price + output tokens × output price
Expand for the CLI-11 record card, stop rules, and observability
Start with one small task and fill in this card:
| Field | What to record |
|---|---|
| Task | What you asked the agent to do |
| Provider/model | The provider and model actually used; write unconfirmed if unavailable |
| Usage | Input/output usage; do not write only a vague “total tokens” |
| Time | The actual duration shown by the workflow or CLI |
| Result | PASS, an issue list, or the reason for failure |
| Cost | Calculate only when it matches official prices; otherwise write the billing method or unconfirmed |
Then set a stop rule the tool really supports, such as a job timeout, maximum retries, provider spend limit, or human confirmation before each paid step. Do not create a setting a tool will not read just to create a false sense of safety.
For comparing multiple runs, you can choose Langfuse, Phoenix, Helicone, or promptfoo. First confirm where data will be sent and whether it contains the original prompt, code, or PII before deciding to connect it.
Prompt caching TTL, eligibility, and pricing vary by provider and model. Anthropic’s current documentation describes both a default 5-minute TTL and an optional 1-hour TTL; treat this as a product setting to check, not a fixed rule for every CLI.
Hands-on exercise CLI-12: Safely hand a Skill to a teammate¶
Outcome: a second clean demo repo can find the review-changes Skill, and running it makes no unexpected changes.
Put A2’s review-changes Skill in a version-controlled team repo and include four things: installation location, required permissions, test method, and removal method. Claude Code users can package it according to the official plugin format; other CLIs should follow their own Skill documentation.
Expand for CLI-12 sharing, installation, and revocation steps
- Before sharing, read
SKILL.mdand its attached scripts. Confirm that they do not download unfamiliar programs, read secrets, or change external systems. - Keep
skills/review-changes/SKILL.mdat the plugin root; do not package the project’s ownCLAUDE.md,AGENTS.md, or secrets with it. - Install it in a second clean demo repo according to the tool’s documentation. Claude Code users can refer to the Plugins documentation and
anthropics/claude-plugins-official. - Make a small document diff, run the Skill, and use
git status --shortto confirm that it only reviews and does not edit files. - Record the version or commit SHA. Read the diff before updating; when you stop using it, remove the plugin/Skill according to the documentation and confirm that the agent can no longer find it.
The core idea of a Skill can be shared, but folders, permissions, frontmatter, and installation methods may differ. Do not describe one tool’s plugin format as universal to all CLIs.
Cost reminder: sharing the files usually does not incur model charges, but each teammate’s Skill run may use their own subscription or API allowance.
Remember only this production safety loop¶
Define the scope → run read-only → leave a record → human judgment → recoverable
If you have no scope, evidence, or recovery method, do not increase permissions yet. This matters more than memorizing many tool names.
📋 Playbook 4: Dispatching subagents for independent tasks¶
Outcome: first list the agents the current tool really provides, then delegate an independent, verifiable task; do not assume every computer has an agent with the same name.
Expand for Playbook 4 and the other six advanced playbooks
Playbook 4 — subagent: a subagent is an independent helper dispatched by the main session. Claude Code currently has built-in subagents such as Explore, Plan, and general-purpose; the available list still depends on the version, session, and settings. code-reviewer is a custom example in the official documentation, not a built-in agent that every installation has. First run the tool’s agent list, then choose a read-only agent or create a restricted reviewer.
For other situations, remember one action and keep the theory in Stage 7.5:
- Unclear scope: write down paths that may and may not change; request a plan before changing files.
- Multiple people/agents in parallel: separate ownership and commits, then integrate at the end; do not change the same batch of files at the same time.
- Review-agent output: the reviewer provides evidence; it does not replace tests, branch protection, or human judgment.
- Running an agent in CI: start with read-only access and a trusted trigger; model fallback must be explicitly configured and revalidated, never switched silently.
- Cost control: use actual usage, timeouts, retries, and provider limits; say when data is unavailable.
- Preventing rule drift: deliberately make a small safe failure to confirm that the gate really blocks it; rule text by itself is not evidence.
Further reading: resources/subagent-cookbook.en.md and Stage 5.5. These pages will be rechecked in their own layer later; before using agent names, follow the official documentation and the actual list available to you.
🎯 Curated Projects¶
Editorial ratings are learning-map guidance, not GitHub stars. ⭐⭐⭐⭐⭐ marks a must-read or must-run entry for this path; it does not mean the tool is always safe or that production can skip its own threat model.
| Type | Resource | Read first | When to use | Rating | Source |
|---|---|---|---|---|---|
| Safe MCP connections | MCP Connect to local servers | Allowed directories and explicit authorization | Connecting a local server for the first time | ⭐⭐⭐⭐⭐ | Official docs |
| MCP Security Best Practices | Least privilege, scopes, and token handling | Before connecting an account or remote service | ⭐⭐⭐⭐⭐ | Official docs | |
github/github-mcp-server | --read-only, toolsets, and tools allow-list | Reading GitHub PRs/issues | ⭐⭐⭐⭐ | GitHub repo | |
modelcontextprotocol/servers | Reference implementations and the not-production-ready warning | Learning the protocol or reading example code | ⭐⭐⭐⭐⭐ | GitHub repo | |
| CI and PR review | GitHub Actions Secure Use | Least privilege, untrusted input, and pinning SHAs | Before writing any workflow with secrets | ⭐⭐⭐⭐⭐ | Official docs |
| Claude Code GitHub Actions | Official setup, permissions, and troubleshooting | Running Claude Code in CI | ⭐⭐⭐⭐⭐ | Official docs | |
anthropics/claude-code-action | Official examples and action inputs | Starting from an executable template | ⭐⭐⭐⭐⭐ | GitHub repo | |
| Codex GitHub Action | Permission profile, trigger, and output | Running Codex in CI | ⭐⭐⭐⭐⭐ | OpenAI official docs | |
openai/codex-action | :read-only and safety strategy | Checking the latest inputs and examples | ⭐⭐⭐⭐⭐ | GitHub repo | |
| Observability and evaluation | langfuse/langfuse | Traces, usage, and eval | Viewing multiple runs together | ⭐⭐⭐⭐⭐ | GitHub repo |
Arize-ai/phoenix | Tracing and evaluation | Observing an AI system with open source | ⭐⭐⭐⭐ | GitHub repo | |
Helicone/helicone | Proxy/gateway data flow and privacy boundary | Collecting request records from a gateway | ⭐⭐⭐⭐ | GitHub repo | |
promptfoo/promptfoo | Eval cases and CI regression | Comparing whether a change made things worse | ⭐⭐⭐⭐⭐ | GitHub repo | |
| Sharing Skills/plugins | Claude Code Plugins | Plugin structure, installation, and marketplace | Packaging for Claude Code | ⭐⭐⭐⭐ | Official docs |
anthropics/claude-plugins-official | Officially managed plugin directory | Finding readable official examples | ⭐⭐⭐⭐⭐ | GitHub repo | |
obra/superpowers-marketplace | Minimal marketplace shell | Understanding curator-only structure | ⭐⭐⭐ | GitHub repo | |
| Directories and complete examples | wong2/awesome-mcp-servers | Classify first, then check sources and permissions one by one | When official resources lack the server you need | ⭐⭐⭐⭐ | GitHub repo |
obra/superpowers | How Skills, rules, and workflows fit together | Looking at a complete example after the minimal workflow works | ⭐⭐⭐⭐ | GitHub repo |
The directory only helps you “find candidates”; it does not guarantee a candidate is safe. Before installing any MCP, Action, Skill, or plugin, check its source, permissions, recent maintenance, and removal method again.
✅ Track A Completion Check¶
- MCP received only the demo folder or a minimal read-only toolset.
- The PR workflow only gives feedback; it does not auto-merge, push, or deploy.
- Secrets are not in the repo, prompt, or log; the workflow uses least privilege.
- I can point to the result and usage for one run; unavailable data was not guessed.
- A teammate can run the Skill in a clean demo repo, and
git statusshows no unexpected changes afterward.
Once all five are true, the Track A core is complete. The recommended next stop is Stage 8 — Agent Interfaces, where you set safe boundaries for browsers, computers, and sandboxes. Stage 8 does not block Track A Capstone entry. If you want to build your own agent, return to Stage 3.