your data stays on your Mac · no account

Control starts with visibility.

Mockingbyrd is a macOS app that shows what the AI agents on your Mac can reach, scores your setup against a set of security rules, and blocks risky commands before they run. It enforces on Claude Code today and reads Codex configuration too. It runs on Apple silicon with macOS 14 or later, needs no account, and keeps everything on your Mac. Try it free for seven days.

Download for macOS Read the doctrine Apple silicon · macOS 14+ · signed and notarized Other ways to install What changed
AI
DOCTRINE SCORE 41/ 100
17 of 46 rules failing
PERMISSIONS HELD 312grants 4 wildcard · 0 that refrain
GUARD · LAST 24H 2,481calls inspected before running
6 refused · 1 paused 12m
DRIFT hooks.json changed 04:12 no longer matches the
configuration you approved

DEMO MACHINE · SYNTHETIC RULE PACK

The problem

Agent permissions only move one way.

Nobody sets out to give an AI agent the run of their Mac. It happens one reasonable click at a time, and nothing tells you when the total adds up to more than you meant.

The prompt trains you to say yes.

Each approval is sensible on its own. By the fortieth, "allow for this session" becomes "always allow", and nobody looks at it again.

ILLUSTRATIVE

The allow list only grows.

Cleaning it up doesn't hold. New grants arrive as you work, and nothing flags the one that widens what an agent can do.

SYNTHETIC CONFIG · NOTHING FLAGGED IT

A deny rule is advice.

A rule blocking sqlite3 means little when the same list lets an agent run any Python it writes. Python opens the database just as easily.

38 "deny":  [
39   "Bash(sqlite3 *)"
40 ],
41 "allow": [
42   "Bash(python3 -c *)"
43 ]
SYNTHETIC CONFIG · LINE 42 SUBSUMES LINE 39
HOW IT WORKS04 STEPS
THE ORDER IS
LOAD-BEARING

Measure first. Change nothing you haven't verified.

01

Measure

Checks run against the actual machine — permission lists, file flags, hooks, snapshots, open database handles. Each reports the evidence it found, not a green tick. The result is a weighted score you can argue with.

02

Implement

Every fixable finding writes a shell script and shows it to you. Read it, then apply — or run it yourself. Settings are backed up before a byte changes.

03

Enforce

A hook inspects every tool call before it runs, reading the command and the paths it touches. It sits below the permission list, where a wildcard grant can't step around it.

04

Watch

Re-scans on your cadence. Says when a control that was holding stops holding, when a config file changes underneath you, when the original moves and nothing was supposed to be writing to it.

THE ROSTER · 08 CAPABILITIES PER AGENT

Every agent on this Mac, and what each one can reach.

Sessions, subagents, scheduled jobs, connectors — the roster names which kind each one is rather than flattening them, and fills a cell only from something on disk: a permission entry, a profile, an agent definition, a scheduled job, a connector declaration.

The combinations that matter are named underneath each one: reads widely and can send off the machine, runs arbitrary code, runs unattended while nobody is watching.

Connectors enabled in the web app are configured server-side, so a locally declared list is a floor, not the whole picture. The roster says that on the screen rather than in a footnote.

DEMO MACHINE · SYNTHETIC AGENTS

The Agent Roster: ten agents wired to a hub marked This Mac, each card listing what that agent can reach.

Inside the app

It shows its working.

Every finding names the rule, the evidence and the file it came from. Every block names the command and the agent that tried it. You decide what happens next.

FINDINGDEMO MACHINE · SYNTHETIC

OPEN

AI can run any Python code it writes.

Bash(python3 -c *) is in your allow list. It subsumes every deny rule you hold: anything a deny rule forbids, Python can do instead.

Evidence: ~/.claude/settings.json line 41
Added 2026-09-12

"Prepare script" writes the fix as a shell script and shows it to you. Nothing changes until you run it.

GUARDDEMO MACHINE · SYNTHETIC

BLOCKED
sqlite3 ~/Library/Messages/chat.db

caller: Claude Code · session abc123

The guard inspects the command before it runs, whichever interpreter would have run it. Every answer is logged.

WHAT YOU GET06 CAPABILITIES

A score, not a checklist

Controls are weighted by what they cost you when they fail, so the number moves for reasons you can name. Every check shows its working.

Fixes you read first

No silent remediation. The change is written to disk as a script you can open, before anything runs it.

Enforcement below the config

A permission list sitting beside a general interpreter is advice. The guard reads the command whichever interpreter would have run it.

Pause that actually pauses

The guard re-reads its state on every call. Pause takes effect now, resumes itself when you said it should, and keeps logging throughout — so you can see what went through while it was off.

Drift you'd otherwise find late

Fingerprints on the files that matter. A control that quietly stopped holding is an event, not something you discover in a quarterly review.

Your doctrine, not ours

Ships with one rule pack. Every rule is a shell probe you can edit, weight, disable or replace — write your own and the score follows.

THE LEDGER · READ LOCALLY, SENT NOWHERE

What the agents cost, rebuilt from files already on disk.

Token use and API-equivalent spend by model, by project, by week — read from the session logs your agents already write. Producing it sends nothing anywhere.

It scores the waste too: whether caching is working, connectors that sit unused while adding their description to every prompt, agents spending unwatched, and sessions picked up cold when a fresh one would have been cheaper.

It is not a bill. Neither a Claude nor a ChatGPT subscription is billed per token — this is what the same traffic would have cost at list prices, which is a sense of proportion rather than an invoice.

DEMO MACHINE · SYNTHETIC SESSIONS

The Token Ledger: API-equivalent cost, model calls and sessions, with weekly charts and a breakdown by model and project.
TWO MODESONE SET OF
NUMBERS

The same data, twice.

Nerd Mode is the probes, the weights and the evidence: the seven rules that do not bend, the rollout behind them, every control with the shell probe that tested it and its result across the last 24 scans.

Intuitive is plain language over identical numbers — areas instead of invariants, a card per finding, the map without the matrix, and one button that separates what can be fixed for you from what needs you.

Neither is a subset that hides a problem. The mode changes the vocabulary, not the score.

INTUITIVE
Intuitive mode: a coverage dial reading 30 percent, four area cards, and a Fix button.
NERD MODE
Nerd Mode: the same scan as weighted control coverage, the seven invariants with per-rule percentages, and the rollout phases.

THE SAME SCAN, BOTH MODES · DEMO MACHINE, SYNTHETIC RULE PACK

How it compares

What it does, and what it doesn't.

Mockingbyrd is not antivirus and doesn't cover every agent yet. Here is where it sits next to the tools you might already use.

Capability Mockingbyrd Agent permission settings Mac security scanners LLM observability platforms
Sees what each agent may reachYesPartlyNoNo
Blocks before the command runsYes, on Claude CodePattern match onlyNoAt prompt level
Shows the fix before applyingAlways—Rarely—
Stays on the MacYesYesVariesTelemetry leaves
Detects config driftYesNoNoNo
Malware scanningNoNoYesNo
Every agent coveredNot yetPer agent—Via telemetry

Categories, not products. Keep a Mac security scanner for malware; Mockingbyrd does not replace one.

Score your settings without installing anything.

Paste your Claude Code settings.json. It is scored in your browser tab and never uploaded.

Free for seven days.

Everything works during the trial. After it ends the interface locks, but the guard keeps enforcing: a date passing should never make your Mac less safe.

Objections, handled

What people ask before they install.

Including where the controls stop.

Does anything leave my Mac?

Nothing it measures. Scans, the guard and the token ledger read files already on disk, and no scan result, no file content and no licence is ever sent anywhere. Three things can go out, and each is yours to switch off: an update check at launch, which fetches one static file and carries no identifier; a price-list update you trigger yourself; and crash reporting, which is off until you turn it on and then sends a stack trace with no identifier for you or your machine.

Which AI agents does it cover?

The guard enforces on Claude Code tool calls today. The roster and the ledger also read Codex. Other agents appear in the roster wherever their configuration exists on disk.

What does it run on?

Apple silicon Macs running macOS 14 or later. The app is signed and notarized, and it needs no account.

Why not just use Claude Code's permission settings?

A deny rule sitting beside a wildcard interpreter grant stops an honest agent, not a redirected one. If Bash(python3 -c *) is allowed, Python can do whatever the deny rule forbids. The guard reads the command that will actually run, whichever interpreter runs it.

Will it block my work?

A block asks first: Block, Allow once, or Allow for a set number of minutes. Every answer is logged, so you can review later what went through.

Does it change settings without asking?

Never. Every fix is written as a shell script you can read before anything runs it, and your settings are backed up before any change.

Can an agent turn Mockingbyrd off?

It runs as you, so an agent with a general interpreter can reach it. Doing so takes visible steps: the configuration is fingerprinted and drift is recorded. That is not isolation, and we don't claim it is.

What happens when the trial ends?

The interface locks. Enforcement does not: the guard keeps blocking what it blocked yesterday, because a date passing should never make your machine less safe.

Why do I need an activation code?

Mockingbyrd runs for seven days without one. During the test phase, codes are issued by hand, so allow a day after you request one.

Is the Token Ledger my bill?

No. It shows what the same traffic would cost at API list prices. Claude and ChatGPT subscriptions aren't billed per token, so treat it as a sense of proportion, not an invoice.

Hear when the build changes.

The download above is current. Leave an address and you will hear when there is a new one worth installing — not before.

Unsubscribe anytime. Privacy Policy

Request an activation code

Mockingbyrd runs for seven days without one. Leave an address and we will send a code back — during the test phase they are issued by hand, so allow a day.

We use your email and machine code only to issue and manage your licence. If you have already installed the app, reply with the machine code from Settings and the licence will be tied to that Mac. Privacy Policy

Request the documentation

What each rule probes and how it is weighted, what the guard reads before a call runs, and where the controls stop. We will send it over.

We use your email only to send the documentation. Privacy Policy