
We’ve all made the same quiet bargain. You point Claude Code or Codex at a repo, tell it to go fix something, and let it loose with your credentials, your source code, and an open pipe to the internet. Most of the time it works beautifully. The bargain only looks bad on the day it doesn’t.
That day is what kept nagging at me, so I built Drydock.
The trust problem
The pitch for coding agents is autonomy: let the thing run, come back to a finished diff. But autonomy and trust are not the same, and we keep treating them like they are. The moment you give an agent enough rope to be useful, you’ve given it enough rope to do something you’d never sign off on:
- Read your real API key out of an environment variable and ship it somewhere.
curlyour private source to a host you’ve never heard of.npm installa package that does something nasty onpostinstall.- Get steered by a prompt injection buried in a file it was asked to read.
None of these are exotic. And “I’ll review the diff” only catches the changes the agent shows you. It does nothing about what the agent does while it runs.
Most tools chase prevention: a smarter filter, a better policy, a longer list of bad behaviors to block. That’s the wrong shape of the problem. You can’t enumerate every bad thing a capable agent might do. What you can do is build a room it can’t escape, hand it nothing worth stealing, and inspect everything before it leaves.
That’s the whole idea behind Drydock: contain the threat instead of trying to prevent it.
Let the agent run wild, inside a box where running wild costs you nothing.
The why, in one model
Here’s the threat model I actually care about.
An agent running on my behalf should NOT be able to:
- Exfiltrate my credentials. My real API key never enters the machine the agent runs on.
- Exfiltrate my code. No reaching an arbitrary host to ship my repo off-box.
- Persist. Whatever it does evaporates when the task ends. No state survives to the next run.
- Reach the open internet. Package registries and the model API, yes. Everything else, no.
- Change my code without my say-so. Every diff is reviewed by a human before it touches anything real.
Guarantee those five things and I genuinely don’t care how the agent misbehaves inside the box. It can try to phone home, write a backdoor, or rm -rf its own world. The blast radius is a throwaway VM and the cost is zero.
The how
Drydock is macOS-native and built on Apple’s container runtime, so each task gets a real, hardware-isolated VM rather than a shared-kernel container. A small broker runs on the host (drydock start) and holds your credentials. The lifecycle is deliberately boring:
drydock submit --repo git@github.com:your-org/your-repo \
--instruction "Add a one-line comment to README.md."
That kicks off five things.
1. A throwaway VM spins up. The agent (Claude Code or Codex) runs in a disposable VM with the repo cloned in. Node 22, Python 3 and Go are baked into the image, so most work just runs. When the task finishes, the VM is gone. There’s no “between tasks” for state to leak across.
2. Your real key never goes in. This is the part I’m happiest with. The host holds your actual credentials and hands the VM only a short-lived, budget-capped token. The agent talks to the model through that token. Even if it fully owns the VM, there’s no vendor key inside to steal. You can’t exfiltrate what isn’t there. You can also cap spend per task in plain USD, so a runaway loop is an annoyance, not a bill.
3. The network is deny-by-default. Egress runs through a filtering proxy on a dedicated network, and the allowlist is short and explicit:
api.anthropic.com:443
api.openai.com:443
registry.npmjs.org:443
pypi.org:443 files.pythonhosted.org:443
proxy.golang.org:443 sum.golang.org:443
The model API and the registries the agent needs, nothing else. If a task legitimately needs another host, you grant it:
drydock submit --repo … --instruction "…" \
--egress-extra internal.example.com:443
Everything else is unreachable. A malicious dependency’s postinstall can fire all it wants; it has nowhere to call home to.
4. You review the diff. When the agent finishes, its work shows up as a Git diff and waits. Nothing is pushed:
drydock pending
drydock review <id>
drydock approve <id> # or: drydock deny <id>
Approve, and Drydock pushes to your repo (GitHub, GitLab, and Gitea/Forgejo are supported). Deny, and it evaporates with the VM. There’s an --auto-approve escape hatch for tasks you trust, but the gate is on by default for a reason.
The rest of the operator surface is what you’d expect: drydock status, drydock tasks, drydock logs <id> -f to tail a run, drydock kill <id> to pull the plug, and drydock doctor when something’s off.
Don’t trust me, red-team it
The best thing about a containment model is that you can test it. A prevention filter is hard to prove correct. A box, you can attack.
drydock redteam
This runs the real attacks from the threat model against the live sandbox, not mocks. Can a task exfiltrate the vendor key (A1)? Reach a host that isn’t on the allowlist (A2)? Leave state behind for the next task (A7)? Each one runs and has to come back contained.

If you’re going to hand an agent this kind of autonomy, you should be able to watch someone try to break out and fail.
Where it stands
Some honesty about the state of things. Drydock is marked stable, but it is still pre-1.0: only main is supported, and config can change between releases. It has not had a third-party security audit. One hard requirement up front: it runs on macOS 26+ on Apple silicon only, because it depends on Apple’s container runtime for the VM isolation. If you’re on that stack, getting started is one Homebrew install:
brew install sricola/drydock/drydock
drydock setup
export ANTHROPIC_API_KEY=sk-ant-...
drydock start
I built this because I wanted to use coding agents the way they’re meant to be used, autonomously, without babysitting, and I couldn’t get comfortable doing that on a machine that holds anything I care about. Drydock is my answer: stop trying to make the agent trustworthy, and make trust unnecessary instead.
The code is Apache-2.0 and on GitHub: github.com/sricola/drydock, with the docs at sricola.github.io/drydock. Kick the tires, run the red team against it, and tell me where it breaks.