Proof it ran
A routine can say what "it worked" looks like. A small checker looks every 5 minutes. If a routine did not run, or ran and left no proof, the checker reports it once to the fixer. A routine with no proof set is marked cannot verify, never green.
Where you see it
Mostly as colors on the Working for you panel and the Health tab. Green means it ran and left its proof. Red means it did not. A routine that never said what its proof is shows cannot verify. It never shows a fake green. Not every routine has a proof yet. In the reference install, about half do.
The owner gets no message from the checker itself. It hands the break to Otto, the fixer, and he reports.
What happens, step by step
Each routine's row can carry a done field (the proof). It comes in two shapes:
- A fresh file. "This file must have been written in the last 26 hours."
- A command that must pass. A small test that has to finish without an error.
Every 5 minutes, a small program called routine-pulse.py (the checker) does this:
- It opens
routines.jsonand skips anything parked or not switched on yet. - For a timed job, it asks systemd (the server's job scheduler) when the job last ran and when it should have. It allows a grace window of about 10 minutes.
- For an always-on job, it checks that the job is running and has been up more than 10 minutes.
- If that passes, it checks the proof. Is the file fresh? Did the command pass?
- It gives one verdict: green, red, cannot verify, or unknown (it could not read the state at all).
- On red, it files one report to Otto. Then it says nothing more about that routine until it turns green again. On green, it closes any open report.
- It writes its own heartbeat file (a small "I ran" marker) and a status file the Health tab reads.
The checker never fixes anything. It never calls an AI model. It is plain code, so it cannot be talked into anything and costs nothing to run.
What powers it
| Part | What it does |
|---|---|
done field | The proof each routine promises, stored in routines.json. |
routine-pulse.py | The checker. Runs every 5 minutes, gives each routine a verdict. |
routine-pulse.timer | The clock that starts the checker. |
routine_report.py | The one helper every failure report goes through. One report per break. |
routine-failed@.service | Fires when an always-on job crashes and systemd's restart did not hold. |
healthcheck.py | The daily checkup. Files what it could not fix as one batched report. |
resident-wake.py | The agent alarm clock. Also watches the checker's heartbeat. |
Why it works this way
A morning prep run in the reference install once exited cleanly and logged DONE. The work never happened. A helper it started had died when the main program stopped. The owner lost a morning before anyone noticed.
Exit codes prove the process ended, not that the work happened. So a routine should name a real thing on disk that only a real run leaves behind. One that does not is shown as cannot verify.
Who checks the checker? The checker cannot notice its own death. So the agent alarm clock, resident-wake.py, reads the checker's heartbeat. If the checker has not run in 15 minutes, the alarm clock files the report itself: nothing is watching the other routines.
Three roads, one door. A break can reach the fixer three ways. The checker finds it. A crashed always-on job fires routine-failed@.service (a crash handler). Or the daily checkup finds something it could not repair. All three go through routine_report.py, so a break is reported once, never three times.
Connected to
- Otto, the fixer: who receives every report.
- Routines and the registry: where the
donefield lives. - The daily checkup: the third way a break is reported.
- The Health tab: where the verdicts show.