A turn is an attempt.
A case is a promise.
Done is a receipt.
Most assistants answer questions and forget you the moment the chat ends. John Bot owns outcomes. It watches, plans, works, verifies — around the clock — and interrupts exactly once, with one button, when a decision is genuinely yours. Everything it claims to have finished, it can prove.
Everything John Bot does — a message from you, a schedule firing, an email arriving, a server misbehaving — enters the same loop, and nothing leaves it without an ending.
Fourteen capabilities, one temperament: do the work, prove the work, spend John's attention like it's the scarcest resource in the system — because it is.
Ask for a follow-up on Friday and it doesn't just agree — it opens a case with an objective, a definition of done, and a wake time. The case survives restarts, failures, and weeks of silence. It cannot close without proof.
Outbound email used to mean copy, paste, send — your labor. Now the bot drafts, stages the exact payload, and hands you two buttons. Approve, and it sends, verifies the message ID, watches the thread, and closes the loop. Your part takes two seconds.
A reply from someone who matters, sitting unread for four business hours, becomes a case within the half hour — nights and weekends included. No more “luck of the next check-in.”
Before v3, a failed morning brief was simply skipped — forever, and silently. Now every kind of failure has exactly one owner that retries it, and anything that exhausts its retries lands in front of the right fixer instead of vanishing.
Twenty-nine health checks watch everything from disk space to expiring TLS certificates to its own watchers. Every possible red has a named owner: a mechanical fix it tries itself, an AI medic that investigates, or — rarely — you.
Every proactive run ends with a structured verdict — acted, silent, blocked, or needs-you — so delivery, retries, and metrics run on data instead of parsing prose. You only ever see the human sentence.
One message a day: what finished (with receipts), what fixed itself, what's staged and waiting, what it's watching, what it spent, and one batched list of the things only you can unblock. The invisible 99%, made visible.
Autonomy is measured nightly: verified completions over eligible work, human touches per outcome, silent losses (target: zero), cost per verified result — all against a baseline captured before the rebuild began.
When the bot says “I'll check tomorrow,” a deterministic audit verifies a real obligation was created — a schedule, a case, a watch. A promise with nothing behind it becomes a case within fifteen minutes.
React 👍 or 👎 to anything the bot sends and it asks one short question: why? Your answer is captured and routed — into the rule that produced the message, a standing preference, or a tuning case. Ignore the question and it expires silently; the thumb still counted.
You write authority in plain prose in your own notes. Fenced grant blocks inside it compile into standing permissions — “routine sends to the accountant are one-tap” — but only after you approve the exact diff, with a tap.
Two machines heartbeat each other. If the server goes quiet, the Mac notices. If the Mac goes quiet, the server says so. Whole-machine silence — the failure that hides every other failure — is no longer invisible.
Standing facts with provenance and expiry dates, full-text recall across every conversation ever, and a human-editable “mind” — identity, mandate, priorities — that you keep in plain markdown files.
Research, multi-file builds, long investigations — the bot spawns background workers that survive restarts and report back when done. It even builds parts of itself this way, with changes landing only through approved review.
Three layers, two services, one database. The AI is the mind; everything that must never be wrong is deterministic code with a test pinning it.
Every safety control is labeled for what it really is. Mistake-prevention stops a confused AI; adversary-resistance stops a manipulated one. The credential isolation, scrubbed environments, single-use approval tokens, and quarantine are the second kind. The rest is the first — and the documentation says so plainly, because a safety story you can't audit is theater. One rule survives everything, verbatim: nothing external-facing sends without John's explicit yes. A tap is a yes.
V3 began by measuring the before — so the after means something.
The goal was never zero involvement. It was zero labor. What remains for John is exactly the set of things that should never be automated: