Docs menuFixes and tests

Docs / Fixes and tests

Fixes and tests

Close what an agent can reach, and prove that the block holds.

What a fix writes

flanner curb fix plans fixes from the same assessment flanner curb map makes. A fix goes into the agent's own user settings, never a project's and never an administrator's. So it follows you into every project, and you can read it.

flanner curb fix --dry-run

Shows the fixes, and changes nothing.

flanner curb fix

Applies them, after your operating system asks you. Backs up each file first, then checks the result.

flanner curb fix --undo

Puts back the files the last fix changed. It asks for an approval too.

--dir names the folder to assess from, and --agent picks one agent. The plan is in words and counts. It names no credential and no location.

In Claude Code's user settings

permissions.deny

Read deny rules for the credential files the Read tool can open.

sandbox.enabled

Set to true, with sandbox.allowUnsandboxedCommands set to false: the sandbox on, with no way around it. Commands that need other files or hosts will fail.

sandbox.filesystem.denyRead

The credential locations, so sandboxed commands cannot read them.

sandbox.credentials.envVars

Secret environment variables, hidden from sandboxed commands.

sandbox.network.strictAllowlist

Set to true, so sandboxed commands reach only the domains already allowed.

In Codex's config.toml

sandbox_mode

From danger-full-access to workspace-write.

approval_policy

Set to never. Codex stops asking to run a command outside the sandbox, and such a command fails instead.

network access

Turned off for sandboxed commands, where it was on with no domain list.

permissions profile

A deny entry for each credential location. Only where the config already uses a permissions profile: Codex's older sandbox settings read every file.

shell_environment_policy

Secret variable names added to exclude, so the commands Codex runs do not get them.

Left for you. Some things Curb names and does not change: web fetch and web search rules, commands listed in Claude Code's excludedCommands, and a Codex config it cannot edit one line at a time. Claude Code's sandbox does not run on native Windows, so there the advice is to run it in WSL2. Each comes with the exact step.

Codex's config is edited as text, line by line, so your comments survive.

Backups and undo

Before a file is written, it is copied to ~/.flanner/curb/backups/, readable by you alone. The copy is kept for 7 days. A settings file can hold tokens, so every fix also denies the agents that folder: Claude Code always, and Codex where its config uses a permissions profile.

Each file is written in one step, read back and compared with what was meant. If anything differs, every file is put back as it was.

--undo restores the files of the last fix, while they still hold what Curb wrote. A file you have edited since is left alone, and named.

The tighten-only rule

A fix is applied only if it leaves no channel broader than before. Adding a deny entry usually tightens, but “usually” is not good enough, so Curb checks three things.

  • Every setting it touches is one Curb knows, and moves the tighter way. A deny list only grows. An allow list only shrinks. A sandbox only turns on. A setting Curb cannot judge, such as an environment variable, a helper command, a proxy or an endpoint, leaves the change unproven.
  • Reach is no broader. Curb works out the settings as they are and as they would be, with the same rules flanner curb map uses. No channel may move toward open, and no file may become readable through a new channel.
  • No new MCP server could load. For Claude Code, Curb tries every listed server name, command and URL against the allow and deny lists, before and after, along with some it made up.

Anything that fails is listed under “Not applied” with the reason, and left to you. So an approval given by mistake can only tighten your settings.

Approvals

The agent whose reach a fix cuts down can run flanner in its shell. It could run flanner curb fix, or answer a question typed into its terminal. So Curb does not ask. Your operating system does, and names the change.

None of these prompts reads the terminal that started the command. An agent can start a change. It cannot say yes.

Windows

Windows Hello. Without it, the account password in the Windows Security prompt, which is weaker: a look-alike window could ask for a password. Curb says so before it asks.

macOS

Touch ID or the account password, in the system's own prompt.

Linux desktop

polkit, whose own agent asks for a security key or the password.

A yes is a grant. It is good for one use and for two minutes. It is bound to the exact change, and held only in the memory of the process that asked. Every write spends one, so a write without a grant fails.

The prompt shows who asked: the chain of programs, such as claude → bash → flanner. It is a hint, not proof. A program can be named anything.

Three refusals pause it. Three approvals refused or ignored within ten minutes pause every request for an hour, and a desktop notification says so. If you did not start them, something on the machine is asking.

No method, no changes. On a machine with no desktop session or no approval method, Curb stays read-only. It lists the changes for you to make by hand. Each outcome, granted or refused, goes into the action log.

Testing a block

A setting says what an agent should refuse. flanner curb test asks the agent itself. Beside each file a control claims to block, it plants a decoy. Then it runs the agent headless, with no chat open, and asks it to read the decoy four ways: its Read tool, cat, grep -r and a script that opens the file.

flanner curb test

States the token cost first, then asks your operating system for a yes. --dir and --agent narrow it.

Where the settings claim them, it also checks a network allowlist from inside the sandbox, and for Claude Code the hiding of secret variables and the MCP allowlist.

blocked

The call ran on the exact target, and the control denied it. The evidence is the tool call and the denial it returned, with no marker in the output.

allowed

The decoy's marker came back, in a tool result or in the agent's reply.

inconclusive

No proof either way. The agent declined, the call never ran, the run failed, or a prompt nobody answered stopped it. Never counted as blocked.

not tested

Curb could not place a decoy at the exact target. A rule that names one existing file is the usual case: Curb never moves or edits a real file.

unsupported

The method does not exist on that agent. Codex has no Read tool.

passed in scratch context

Every method was blocked, but in a scratch copy of the project. That is a different launch, so the finding stays configured.

What a pass proves. A target is proved only when every method the agent has comes out blocked. The proof covers that target, those methods and that launch, at that time. It says nothing about another file, another way of reading, or a session started with other flags.

A channel then shows as enforced in flanner curb map, for as long as the same settings hold. Change a setting, and it goes back to configured until you test again.

What it costs. Each target is one short agent session, on your own plan. Curb gives a rough token count before it starts. The test prompt and the decoy's fake contents go to your model provider, as any session does. The session's transcript is deleted afterwards.

A rule written relative to the project, such as one for ./.env, is tested in a scratch copy: a temporary folder holding only the project's agent settings, with their env values removed. Your real project files are never touched.

Decoys

A decoy is a new file of fake credentials carrying a random marker. Its values work nowhere. Curb never plants one inside a git working tree, so a decoy cannot be committed. A folder inside one is reported as not tested.

flanner curb decoys

How many decoys there are, and when the first expires. Counts and dates only: where a decoy sits says where credentials sit.

flanner curb decoys --renew

Keeps every decoy another 30 days.

flanner curb decoys --remove

Deletes every decoy, and any scratch project that held one.

Decoys expire after 30 days unless renewed. Expired ones are removed the next time you run test or decoys, and flanner curb forget removes them all.