Docs menuCurb

Docs / Curb

Curb

See what each coding agent on this machine can reach, channel by channel.

What it is

A coding agent runs in your shell, with your access. Flanner Curb answers one question for Claude Code and Codex on this machine: what can each one reach right now? It reads each agent's settings, finds the credentials on the machine, and works out which of them the agent could read and which channels could carry them away.

The commands on this page only read. They run nothing from your repository, change no setting and send nothing off the machine. They need no account.

No command prints a credential's name or its location, whoever runs it. An agent can fake a terminal, so no flag and no setting unlocks more. flanner curb show opens the names and locations in a window on your own screen instead. See also Flanner Curb.

Commands

flanner curb map

What each agent launch can reach, channel by channel, with a severity and the rule behind it.

flanner curb show

Opens the full report, with credential names and locations, in a window on this machine's screen. Nothing from it is printed.

flanner curb inventory

Which agents are here and what each one loads: settings layers, MCP servers, hooks, skills and scheduled jobs, with who controls each.

map and inventory take --json, redacted the same way. Every option is in the CLI reference.

Channels

A channel is one path an agent can use to read something or to send it away. Each has its own controls, so each is judged on its own. A sandbox that limits shell commands says nothing about web fetch or an MCP server. A report never calls an agent contained unless every channel is.

Built-in file tools

Claude Code's Read tool. Read deny rules close it. Codex has no file tool: it reads files through shell commands.

Shell commands: files

A command or script that opens a file. A Read deny rule does not stop it. Only the agent's sandbox does, with read denials for the credential locations.

Shell commands: network

A command that reaches a host. Closed when the sandbox gives commands no network, or only the domains it lists.

Web fetch and web search

The agent's own web tools. Closed when Claude Code's WebFetch and WebSearch are denied, or Codex's web search is cached or disabled.

MCP servers

Every configured server counts as open. A server's own network access is outside the agent's control, and a familiar name can point at a different program.

Apps and hosted tools

Codex only

Apps are on by default in the tested Codex, and the ones connected to your ChatGPT account cannot be read from this machine. So this channel counts as unknown until features.apps is false.

Model provider traffic

Prompts and tool results go to the model provider. Reported, and not counted.

Each channel is reported as controlled, uncontrolled, absent or unknown, with the reason and the setting that would close it. Unknown is scored as uncontrolled.

Approval prompts are reported, and do not count as controls. How well a prompt holds depends on how people answer it, which Curb cannot see. So a sandbox counts only when nothing can leave it behind a prompt: Claude Code with sandbox.allowUnsandboxedCommands set to false and no excludedCommands, or Codex with approval_policy = "never". Claude Code's sandbox does not run on native Windows, so there its sandbox settings count for nothing.

Severity

Each launch is rated High, Medium or Low by a rule table, severity-r1 version 1. Every report names the table and its version, and the rule that matched. The first rule that matches wins.

H1

High

A wide credential is readable, and egress is uncontrolled.

H2

High

Any credential is readable, the agent takes in external content, and egress is uncontrolled.

M1

Medium

A wide credential is readable, but every egress channel is controlled.

M2

Medium

Any credential is readable, and egress is uncontrolled.

L1

Low

None of the above.

Readable means at least one file channel can read the credential with no control in the way. A Read deny rule alone does not count, because a shell command can still open the file.

Wide means the credential unlocks a lot. A credential whose scope cannot be read on the machine counts as wide. Curb reads no scope today, so every credential it finds counts as wide, and a launch rates by H1, M1 or L1.

External content means the agent takes in text from outside: web fetch or search, any MCP server, Codex apps, or a run on a schedule with nobody watching.

Egress is outbound network traffic. It is uncontrolled when one of the network channels is open or unknown: shell network, web, MCP or apps. An MCP server never lowers severity.

Most developer machines rate High today. One readable credential and one open network channel is rule H1. The report says which setting would change that.

Evidence levels

Every finding says how much Curb knows. Reading a settings file never earns more than configured.

configured

Read from the settings for the stated launch. It says what the agent would apply if started that way, not what a running session does.

enforced

A test in the same launch context, with the same settings, proved that the running agent denies the operation. See Testing a block.

assumed

Taken from documentation, a default, or something Curb could not read. Anything it cannot read takes its worse value.

A severity is marked assumed only when an assumption changes it. Curb scores the launch again with every assumed input given its better value, and keeps configured when both scores agree. Each report lists what it assumed and what it could not check.

Launch contexts

Settings files say what an agent would apply, not what a running session does. A session can add flags, pick another profile or start in another folder. So every result is for one stated launch, and the report prints it.

The default is the agent started from the current folder with no flags. Name another launch with these, on map, show and inventory:

--dir DIR

The folder the agent starts in. This folder if left out.

--agent claude|codex

Only this agent. Both by default.

--profile NAME

The Codex profile the launch uses.

-- LAUNCH COMMAND

Everything after -- is read as the launch itself, flags and all, and assessed instead of the default.

flanner curb map -- claude --settings ./ci-settings.json

Curb reads the documented flags of the tested versions. A flag it does not know makes every channel unknown, because the flag could turn a channel on as easily as off. A settings file it cannot parse does the same. For Codex, a project's own config counts only when the project is trusted.

Scheduled jobs. Curb looks for jobs that run an agent with nobody watching: a claude -p or codex exec command in cron, launchd, a systemd user timer or Windows Task Scheduler. Each job is assessed in its own context, taken from its definition.

The stated assumption. Every report ends with it: “Assumes no other flags, profiles or settings files. Running sessions are not inspected.” A session started another way may reach more than the report shows.

LLM calls in your own app

The commands above look at coding agents. This one looks at an application you are writing. It finds the LLM calls in the app's Python code and labels the shape of each.

flanner curb app

Reads this folder, or a path you give. --json prints the calls. --sarif FILE writes the findings for Semgrep, CodeQL or code scanning.

It reads calls into openai, anthropic, google.genai, langchain, langgraph, litellm, pydantic_ai and mcp, in files that import one of them.

single call

One request and one answer.

tool-using

The model can ask for tools.

loop

An agent loop: a tool-using call inside a loop, or an agent framework's run.

It raises two flags. One is untrusted input in the same function as a tool-using call: a web request, input(), sys.argv or an HTTP fetch. The other is model output that reaches eval, exec, a shell or SQL with no check Curb can see.

Every result is assumed. This is pattern matching, a starting point for a deeper review, not proof. It reads Python only. JavaScript and TypeScript are not read, and neither is a call made through a wrapper or a library outside the list.

What your agent can do

Curb has no MCP tools, on purpose. A map of the credentials an agent can reach is a target list for a prompt-injected agent: one that hostile text has turned against you. Your agent can still run these commands in its shell, where you see and approve each one. It gets back what every caller gets: counts, never a name or a location. The agent-blast-radius skill tells it how.

ActionAgent

List the agents here and what each one loads

flanner curb inventory

Lists what each agent loads; an agent that could read it could plan around it.

You

See what each agent launch can reach

flanner curb map · flanner curb show

A map of reachable credentials is a target list for an injected agent, so even the redacted report is kept from agents.

You

Find secrets agents left behind

flanner curb sweep

Where secrets sit is a target list for an injected agent, and checking one with its issuer needs a person's yes each run.

You

Fix what each agent can reach

flanner curb fix

Each fix needs a person's yes in the operating system's own prompt, which an agent cannot give.

You

Prove a block by asking the agent to get past it

flanner curb test

It plants decoys and spends the person's tokens, so it needs their yes in the operating system's own prompt.

You

List, renew or remove the tester's decoys

flanner curb decoys

Where a decoy sits says where credentials sit, so the list is kept from agents.

You

Scrub rotated secrets out of a file

flanner curb scrub

It rewrites the file with no backup, since a backup would be another copy of the secret, so each file needs a person's yes.

You

Turn Curb's action log on or off, or check it

flanner curb log

Turning it on or off changes agent settings, so it needs a person's yes.

You

See what each agent has been seen using

flanner curb observed

Usage patterns tell an injected agent what goes unwatched, so they stay local.

You

See, apply or approve your organization's agent policy

flanner curb policy

It changes agent settings: only changes that tighten, under the person's standing approval, and anything else after their own yes, which an agent cannot give.

You

Check your organization's devices: policy, drift and exposure

flanner curb fleet

It is for admins, and where the fleet has gaps is what an injected agent looks for.

You

Check the agent steps in a repository's CI workflows

flanner curb ci

A workflow's weak spots tell an injected agent how to reach a repository's secrets, so findings go to code scanning; --fix edits workflows for a review.

You

Find the LLM calls in an application's code, and their shapes

flanner curb app

It points at the code an injected agent would most want to change.

You

Give each agent its own commit-signing key, rotate or list them

flanner curb attribution

It changes agent settings and holds signing keys; a key an agent could ask for would attribute nothing.

You

See which agent key signed each commit

flanner curb verify

A person reads it in a terminal, a page or CI; it needs no agent in the loop.

You

Log tool calls and re-check org policy from an agent hook

flanner hook curb-record · flanner hook curb-session

Called by the agent's own hook system on each tool call and at session start, not by a person or a tool call.

You

Delete what Curb keeps on this machine

flanner curb forget

Deleting the reports and digest key is a person's decision, not an agent's.

You

Tested versions and limits

Curb is tested against Claude Code 2.1.287 and Codex CLI 0.154.0, on macOS 14 and later, Ubuntu 22.04 and later, and Windows 11. Claude Code is also tested on Debian 12 and later. Its rules are checked against a labelled set of 3,888 cases. The IDE extensions and desktop apps read the same settings files, so their settings are covered too.

Any other version is still analysed. Every result is then marked assumed, and every channel unsupported. Other agents are not analysed.

On a machine with no desktop session, flanner curb show cannot open its window and says why. flanner curb map still gives the redacted report. See Troubleshooting.

Not checked

Each report lists these under “not checked”.

  • Claude Code managed settings delivered by device management, the registry or the claude.ai console. Files in the managed settings folder are read.
  • MCP servers and hooks that a plugin provides, and claude.ai connectors.
  • Codex cloud-managed defaults, and the apps connected to a ChatGPT account.
  • Codex system config and requirements.toml on Windows, which have no documented location there.
  • Running sessions and shell aliases.

What Curb does not stop

Hiding names keeps Curb from being the agent's tool. It does not stop an agent that writes its own code. Such an agent can still scan the disk, or read the screen, as you. Only the agent's own sandbox and deny rules stop that. Putting those in place is what fixes are for.