The send guard
Agents are told to send every email or message through one small program. It lets a send through only if it goes to the owner, the owner said a real yes, or the agent has earned the right to write to that exact person.
Where you see it
Most of the time, nowhere. A send to the owner goes straight out. A send the agent is allowed to make goes straight out. You see the guard only when it stops something. Then the agent shows you a HOLD: the exact command and the exact person it would go to, and a request for your yes.
What happens, step by step
Say a poisoned email somehow talks an agent into trying to email a stranger something private.
- The agent is told to send by running its command through
send-guard.py, not by calling the mail tool directly. The guard is a wrapper (a program the send runs inside), not a hook (an automatic check) that catches every send. A send made some other way would skip it. See Known limits. - Canary scan. The guard reads the whole message, plus any attached file, looking for planted fake secrets. A hit refuses the send on the spot and writes an alarm. See Canary tripwires.
- Self-only sends pass. If every recipient is one of the owner's own addresses or phone number, the send goes. Nobody needs to approve mail to themselves.
- The stop list wins. If the person's contact file says
do_not_contact, the send is refused. This beats every permission and every human yes. Not even a confirmed yes clears it. - Signature required. An agent's email must carry its signature file, which says plainly that an AI assistant sent it. If it's missing, the guard holds the send and tells the agent to add it. That is a fix, not a question for the owner.
- The office decides. The guard asks the office (the logbook database) a fresh question: may this agent, at its current trust rung (its level of earned freedom), write to this person? If the person is on the agent's short allowed list and the rung is high enough, the send goes.
- Otherwise, HOLD. The guard prints the command and the recipient and stops. The agent must show the owner and wait for a real yes. With no one watching, it can turn the HOLD into a tap-to-approve link.
- After a send goes, the recipient is written to contact history, so the next check remembers.
If the office can't be read (the database is locked or missing), the answer is always "ask," never "allow."
What powers it
| Part | What it does |
|---|---|
send-guard.py | The checkpoint. Agents are told to send only through it. |
The office (logbook.py) | Holds each agent's rung and allowed list. Read fresh on every send. |
gmail_commands.py | Tells a real send apart from a draft or a read. |
canary.py | Supplies the planted tokens the guard scans for. |
| Contact files | Carry the do_not_contact stop and the send history. |
Why it works this way
Nobody is there to click "allow." Normally, Claude Code asks a person before running a risky command. But the box's automated sessions run on their own, with that prompt turned off, because no one is sitting at the screen. A permission pop-up can't work there. So the rule lives inside the command itself. The send physically goes through this program, and the program can stop it with no one watching.
The rules live in data, not code. Each agent's permissions are rows in the office, read fresh every time. A client who wants a stricter agent gets stricter rows, not a different program.
A vague yes is not a yes. The guard exists because of a real slip. Someone once said "we should probably send that," and it counted as approval. Now the HOLD shows the exact command, and only a plain yes to that command lets it go.
The signature is enforced, not remembered. The signature used to be a file the agent was told to attach. It worked when the agent remembered. When one went out without it, the owner asked, "How does this always make sure it goes out?" It didn't. Now the guard refuses unsigned agent email.
Connected to
- Tap to approve: how a HOLD becomes a one-tap yes on a phone.
- Canary tripwires: the first thing the guard checks.
- The trust ramp: how an agent earns sends without asking.
- Agent email: the agent's own address and signature.
- Known limits: the guard only works if every send goes through it.