How Company OS stays safe
The agents read email from strangers, see private files, and can send messages. Any one of those is fine. All three together is where the danger lives. This section explains the layers that keep a stranger's words from turning into the agent's actions.
The thing that reads the email must not be the thing that can act, and the thing that can act must never see the email.
The three ingredients
An AI gets dangerous when three things are true at the same time.
- It can see private data. Files, notes, keys, the inbox.
- It reads text a stranger wrote. An email, a web page, a form answer.
- It can send something out. An email, a WhatsApp message, a post.
Take away any one and the risk drops a lot. A model that reads strangers' text but has nothing to leak is harmless. A model that sees secrets but never reads outside text can't be tricked by an outsider. The trouble starts when one model has all three. Security researchers call this "the lethal trifecta."
A system that reads the owner's mail has all three by default. The fix is not to make the AI smarter at spotting tricks. It is to split the three ingredients apart, so no single model holds all of them, and then to put checks at every door on the way out.
The layers
Each layer is a separate page. Read them in order, or jump to the one someone asked about.
| Layer | What it does, in one line |
|---|---|
| The email quarantine | A model with no tools reads outside email and can only fill in a small form. |
| The send guard | Nothing leaves the box unless it goes to the owner, the owner said yes, or the agent was allowed that person. |
| The memory guard | Outside text can never be written into the files that load every session. |
| Tap to approve | A one-use link lets the owner approve one exact, pre-written action from a phone. |
| Canary tripwires | Fake secrets planted in three private files set off an alarm if they ever leave. |
| Checking who sent an email | Mail that claims to be from the owner must prove it with a stamp a faker can't copy. |
| Server lockdown | Even a hijacked program on the box can't gain root or touch system folders. |
| Secrets and logins | The AI reuses saved logins and never types a password. Keys travel only encrypted. |
How the layers fit
Think of an attack as a path. A stranger's words come in, get read, and then try to cause an action. Each layer cuts that path at a different point.
- Coming in: sender checks decide whose mail is trusted at all.
- Being read: the quarantine keeps the reader away from any tools.
- Being remembered: the memory guard keeps outside text out of permanent memory.
- Going out: the send guard, the approval link, and the canary scan all stand at the exit.
- Underneath it all: the server lockdown limits what any program could do, even one that went wrong.
No layer is perfect
Every layer here has a gap, and the Known limits page lists them. That is on purpose. The point of stacking layers is that an attack has to beat all of them at once. A short command that slips past the quarantine still has to get past the send guard. A send that slips past the send guard with a planted secret inside still trips the canary. A hijacked process still can't get root.
A single clever filter can fail quietly. A stack of simple walls, each one easy to explain, fails loudly and only in part.
Connected to
- Prompt injection and the email quarantine: the core idea, with a real attack.
- The send guard: the last checkpoint before anything leaves.
- Known limits: what each layer misses, and what stands behind it.
- The trust ramp: how an agent earns the right to send without asking.