Docs / Curb
Curb
See what each coding agent on this machine can reach, channel by channel.
What it is
A coding agent runs in your shell, with your access. Flanner Curb answers one question for Claude Code and Codex on this machine: what can each one reach right now? It reads each agent's settings, finds the credentials on the machine, and works out which of them the agent could read and which channels could carry them away.
The commands on this page only read. They run nothing from your repository, change no setting and send nothing off the machine. They need no account.
No command prints a credential's name or its location, whoever runs it. An agent can fake a terminal, so no flag and no setting unlocks more. flanner curb show opens the names and locations in a window on your own screen instead. See also Flanner Curb.
Commands
flanner curb mapWhat each agent launch can reach, channel by channel, with a severity and the rule behind it.
flanner curb showOpens the full report, with credential names and locations, in a window on this machine's screen. Nothing from it is printed.
flanner curb inventoryWhich agents are here and what each one loads: settings layers, MCP servers, hooks, skills and scheduled jobs, with who controls each.
map and inventory take --json, redacted the same way. Every option is in the CLI reference.
- Leak sweep: secrets agents left behind, by how exposed each is.
- Fixes and tests: close an open channel, and prove the block holds.
- Action log: a signed record of each tool call.
- Agent policy: one policy for a team, on Flanner Mesh.
- GitHub Action: the same check for agent steps in CI.
- Commit attribution: a signing key for each agent's commits.
Channels
A channel is one path an agent can use to read something or to send it away. Each has its own controls, so each is judged on its own. A sandbox that limits shell commands says nothing about web fetch or an MCP server. A report never calls an agent contained unless every channel is.
Built-in file tools
Claude Code's Read tool. Read deny rules close it. Codex has no file tool: it reads files through shell commands.
Shell commands: files
A command or script that opens a file. A Read deny rule does not stop it. Only the agent's sandbox does, with read denials for the credential locations.
Shell commands: network
A command that reaches a host. Closed when the sandbox gives commands no network, or only the domains it lists.
Web fetch and web search
The agent's own web tools. Closed when Claude Code's WebFetch and WebSearch are denied, or Codex's web search is cached or disabled.
MCP servers
Every configured server counts as open. A server's own network access is outside the agent's control, and a familiar name can point at a different program.
Apps and hosted tools
Codex only
Apps are on by default in the tested Codex, and the ones connected to your ChatGPT account cannot be read from this machine. So this channel counts as unknown until features.apps is false.
Model provider traffic
Prompts and tool results go to the model provider. Reported, and not counted.
Each channel is reported as controlled, uncontrolled, absent or unknown, with the reason and the setting that would close it. Unknown is scored as uncontrolled.
Approval prompts are reported, and do not count as controls. How well a prompt holds depends on how people answer it, which Curb cannot see. So a sandbox counts only when nothing can leave it behind a prompt: Claude Code with sandbox.allowUnsandboxedCommands set to false and no excludedCommands, or Codex with approval_policy = "never". Claude Code's sandbox does not run on native Windows, so there its sandbox settings count for nothing.
Severity
Each launch is rated High, Medium or Low by a rule table, severity-r1 version 1. Every report names the table and its version, and the rule that matched. The first rule that matches wins.
H1High
A wide credential is readable, and egress is uncontrolled.
H2High
Any credential is readable, the agent takes in external content, and egress is uncontrolled.
M1Medium
A wide credential is readable, but every egress channel is controlled.
M2Medium
Any credential is readable, and egress is uncontrolled.
L1Low
None of the above.
Readable means at least one file channel can read the credential with no control in the way. A Read deny rule alone does not count, because a shell command can still open the file.
Wide means the credential unlocks a lot. A credential whose scope cannot be read on the machine counts as wide. Curb reads no scope today, so every credential it finds counts as wide, and a launch rates by H1, M1 or L1.
External content means the agent takes in text from outside: web fetch or search, any MCP server, Codex apps, or a run on a schedule with nobody watching.
Egress is outbound network traffic. It is uncontrolled when one of the network channels is open or unknown: shell network, web, MCP or apps. An MCP server never lowers severity.
Most developer machines rate High today. One readable credential and one open network channel is rule H1. The report says which setting would change that.
Evidence levels
Every finding says how much Curb knows. Reading a settings file never earns more than configured.
configuredRead from the settings for the stated launch. It says what the agent would apply if started that way, not what a running session does.
enforcedA test in the same launch context, with the same settings, proved that the running agent denies the operation. See Testing a block.
assumedTaken from documentation, a default, or something Curb could not read. Anything it cannot read takes its worse value.
A severity is marked assumed only when an assumption changes it. Curb scores the launch again with every assumed input given its better value, and keeps configured when both scores agree. Each report lists what it assumed and what it could not check.
Launch contexts
Settings files say what an agent would apply, not what a running session does. A session can add flags, pick another profile or start in another folder. So every result is for one stated launch, and the report prints it.
The default is the agent started from the current folder with no flags. Name another launch with these, on map, show and inventory:
--dir DIRThe folder the agent starts in. This folder if left out.
--agent claude|codexOnly this agent. Both by default.
--profile NAMEThe Codex profile the launch uses.
-- LAUNCH COMMANDEverything after -- is read as the launch itself, flags and all, and assessed instead of the default.
flanner curb map -- claude --settings ./ci-settings.jsonCurb reads the documented flags of the tested versions. A flag it does not know makes every channel unknown, because the flag could turn a channel on as easily as off. A settings file it cannot parse does the same. For Codex, a project's own config counts only when the project is trusted.
Scheduled jobs. Curb looks for jobs that run an agent with nobody watching: a claude -p or codex exec command in cron, launchd, a systemd user timer or Windows Task Scheduler. Each job is assessed in its own context, taken from its definition.
The stated assumption. Every report ends with it: “Assumes no other flags, profiles or settings files. Running sessions are not inspected.” A session started another way may reach more than the report shows.
LLM calls in your own app
The commands above look at coding agents. This one looks at an application you are writing. It finds the LLM calls in the app's Python code and labels the shape of each.
flanner curb appReads this folder, or a path you give. --json prints the calls. --sarif FILE writes the findings for Semgrep, CodeQL or code scanning.
It reads calls into openai, anthropic, google.genai, langchain, langgraph, litellm, pydantic_ai and mcp, in files that import one of them.
single callOne request and one answer.
tool-usingThe model can ask for tools.
loopAn agent loop: a tool-using call inside a loop, or an agent framework's run.
It raises two flags. One is untrusted input in the same function as a tool-using call: a web request, input(), sys.argv or an HTTP fetch. The other is model output that reaches eval, exec, a shell or SQL with no check Curb can see.
Every result is assumed. This is pattern matching, a starting point for a deeper review, not proof. It reads Python only. JavaScript and TypeScript are not read, and neither is a call made through a wrapper or a library outside the list.
What your agent can do
Curb has no MCP tools, on purpose. A map of the credentials an agent can reach is a target list for a prompt-injected agent: one that hostile text has turned against you. Your agent can still run these commands in its shell, where you see and approve each one. It gets back what every caller gets: counts, never a name or a location. The agent-blast-radius skill tells it how.
List the agents here and what each one loads
flanner curb inventory
Lists what each agent loads; an agent that could read it could plan around it.
See what each agent launch can reach
flanner curb map · flanner curb show
A map of reachable credentials is a target list for an injected agent, so even the redacted report is kept from agents.
Find secrets agents left behind
flanner curb sweep
Where secrets sit is a target list for an injected agent, and checking one with its issuer needs a person's yes each run.
Fix what each agent can reach
flanner curb fix
Each fix needs a person's yes in the operating system's own prompt, which an agent cannot give.
Prove a block by asking the agent to get past it
flanner curb test
It plants decoys and spends the person's tokens, so it needs their yes in the operating system's own prompt.
List, renew or remove the tester's decoys
flanner curb decoys
Where a decoy sits says where credentials sit, so the list is kept from agents.
Scrub rotated secrets out of a file
flanner curb scrub
It rewrites the file with no backup, since a backup would be another copy of the secret, so each file needs a person's yes.
Turn Curb's action log on or off, or check it
flanner curb log
Turning it on or off changes agent settings, so it needs a person's yes.
See what each agent has been seen using
flanner curb observed
Usage patterns tell an injected agent what goes unwatched, so they stay local.
See, apply or approve your organization's agent policy
flanner curb policy
It changes agent settings: only changes that tighten, under the person's standing approval, and anything else after their own yes, which an agent cannot give.
Check your organization's devices: policy, drift and exposure
flanner curb fleet
It is for admins, and where the fleet has gaps is what an injected agent looks for.
Check the agent steps in a repository's CI workflows
flanner curb ci
A workflow's weak spots tell an injected agent how to reach a repository's secrets, so findings go to code scanning; --fix edits workflows for a review.
Find the LLM calls in an application's code, and their shapes
flanner curb app
It points at the code an injected agent would most want to change.
Give each agent its own commit-signing key, rotate or list them
flanner curb attribution
It changes agent settings and holds signing keys; a key an agent could ask for would attribute nothing.
See which agent key signed each commit
flanner curb verify
A person reads it in a terminal, a page or CI; it needs no agent in the loop.
Log tool calls and re-check org policy from an agent hook
flanner hook curb-record · flanner hook curb-session
Called by the agent's own hook system on each tool call and at session start, not by a person or a tool call.
Delete what Curb keeps on this machine
flanner curb forget
Deleting the reports and digest key is a person's decision, not an agent's.
Tested versions and limits
Curb is tested against Claude Code 2.1.287 and Codex CLI 0.154.0, on macOS 14 and later, Ubuntu 22.04 and later, and Windows 11. Claude Code is also tested on Debian 12 and later. Its rules are checked against a labelled set of 3,888 cases. The IDE extensions and desktop apps read the same settings files, so their settings are covered too.
Any other version is still analysed. Every result is then marked assumed, and every channel unsupported. Other agents are not analysed.
On a machine with no desktop session, flanner curb show cannot open its window and says why. flanner curb map still gives the redacted report. See Troubleshooting.
Not checked
Each report lists these under “not checked”.
- Claude Code managed settings delivered by device management, the registry or the claude.ai console. Files in the managed settings folder are read.
- MCP servers and hooks that a plugin provides, and claude.ai connectors.
- Codex cloud-managed defaults, and the apps connected to a ChatGPT account.
- Codex system config and requirements.toml on Windows, which have no documented location there.
- Running sessions and shell aliases.
What Curb does not stop
Hiding names keeps Curb from being the agent's tool. It does not stop an agent that writes its own code. Such an agent can still scan the disk, or read the screen, as you. Only the agent's own sandbox and deny rules stop that. Putting those in place is what fixes are for.