Postmortem: The day the AI teammate went dark

Table of Contents
Postmortem: The day the AI teammate went dark

SEV-1 Postmortem (Fictional)

The day the AI teammate went dark for 4 hours

On July 15, 2026, a Railway deployment failure combined with an expired vault token took Boo offline for 247 minutes, silencing heartbeat tasks, Slack replies, and artifact publishing across the workspace.

Incident Commander: Arun Gandlur·July 21, 2026·Status: Resolved

247m
Total downtime
SEV-1
38
Messages dropped
Slack + DMs
6
Heartbeats missed
3 critical
0
Data lost
Confirmed

What happened

At 09:12 UTC on Tuesday, July 15, a routine railway up --ci --service boo-brain deploy triggered by a config change to the skills registry failed mid-rollout. Railway's health check timed out because the new container could not decrypt vault secrets. The 1Password service account token used by the default vault provider had expired at midnight UTC, 12 hours before the deploy.

Railway's zero-downtime deploy rolled back to the previous container. But the previous container had already been evicted from the build cache during a platform GC 90 minutes earlier. The rollback target was gone. The service entered a crash loop.

Boo went fully offline at 09:14 UTC. Slack messages, heartbeat tasks, and artifact deploys all stopped. The self-check heartbeat, which runs every 15 minutes, was itself hosted on the same service, so no alert fired.

Timeline

  • July 15 00:00 UTC

    1Password service account token expires. No alert. Vault reads begin returning 401s, but no process is reading at this hour.

  • 07:42 UTC

    Railway platform GC evicts the previous healthy build image from cache. This is normal behavior, but it removes the rollback target.

  • 09:12 UTC

    DK pushes a skill registry update. Railway triggers a deploy. New container starts, calls vault_unlock for the default provider, gets a 401. Health check fails after 60s.

  • 09:13 UTC

    Railway attempts rollback. Build image not found. Service enters crash loop (restart every 30s, fail on vault init).

  • 09:14 UTC

    Boo is fully offline. Slack bot stops responding. Heartbeat scheduler halts. Artifact app remains up (separate Railway service) but cannot publish new reports.

  • 09:45 UTC

    Madhur notices Boo has not responded to two messages in #e-discussion. Checks Railway dashboard, sees crash loop. Opens incident channel.

  • 10:02 UTC

    Arun identifies the vault 401 in container logs. Team begins rotating the 1Password service account token.

  • 10:18 UTC

    New token generated and written to vault://default/kv/bundle/.... But the token must propagate through the config sync pipeline, which requires a running Boo instance to execute config_sync. Deadlock.

  • 11:30 UTC

    Vishal writes a manual bootstrap script that injects the new token directly into Railway's environment variables, bypassing the normal vault flow. The team debates whether this is safe.

  • 12:45 UTC

    Bootstrap script approved. Arun runs it. Railway receives new env vars and triggers a fresh build.

  • 13:01 UTC

    New container starts, passes health check, connects to vault. Boo comes online.

  • 13:21 UTC

    All 6 missed heartbeat tasks fire on their next scheduled tick. Boo processes 38 queued Slack messages. No data loss confirmed. Incident closed.

Impact

Message response latency, July 15

Minutes to first response. The gap from 09:14 to 13:01 is the outage window.
SystemImpactDuration
Slack bot (replies, threads)Fully offline, 38 messages undelivered247m
Heartbeat scheduler6 tasks missed (3 critical: SSL monitor, daily standup digest, self-check)247m
Artifact publishingArtifact app stayed up but new-report.mjs calls failed (requires running sandbox)247m
Vault readsAll skill activations requiring secrets failed247m
Memory/notebookBlob storage unaffected (separate S3 service), no data loss0m
The self-check heartbeat that should have alerted the team was hosted on the same service it was supposed to monitor. When Boo went down, the alerting went with it.

Root causes

Proximate cause: A 1Password service account token expired. The token had a 90-day TTL set during initial setup in April. Nobody tracked the expiry date.

Contributing cause 1: Railway's build cache GC evicted the last healthy image before the failed deploy, eliminating the rollback target. The timing was coincidental but turned a recoverable deploy failure into a full outage.

Contributing cause 2: The config sync pipeline requires a running Boo instance. When the thing that needs the new config is the thing that is down, there is no escape hatch. The team lost 2.5 hours to this deadlock before writing the manual bootstrap script.

Contributing cause 3: The self-check heartbeat task was paused at the time of the incident (it had been disabled during a noisy debugging session the previous week and never re-enabled). Even if it had been active, it ran on the same service, so it would have gone down too. There was no external uptime monitor.

What went well

Madhur noticed the silence within 30 minutes. In a distributed team, "the bot is quiet" is easy to miss. The team had internalized Boo as a teammate, so its absence registered fast.

Blob storage (memories, notebooks, skill data) was fully isolated on a separate S3-compatible service. Zero data loss, zero corruption. The architecture's separation of compute from storage paid off.

The artifact app stayed up throughout. Reports already published remained accessible. Only new report creation was blocked.

What went wrong

No expiry tracking for vault credentials. The 1Password token had a 90-day TTL and nobody calendared the renewal date.

No external health check. The only monitor (the self-check heartbeat) ran on the service it was monitoring. Classic single-point-of-failure.

The vault bootstrap was a deadlock. Rotating a vault credential required a running instance, but the instance was down because of the credential. The team had to write a one-off script under pressure.

The self-check heartbeat was paused. A one-line status in heartbeat_list output, easy to overlook, and nobody had re-enabled it after the debugging session.

Action items

1

Set up an external uptime monitor (Checkly, Better Uptime, or equivalent) that pings a Boo health endpoint every 60 seconds and pages on-call via PagerDuty. The self-check heartbeat is a complement, not a substitute.

eliminates silent failures
2

Create a credential-expiry heartbeat task that checks all vault provider tokens weekly and posts to #e-discussion 30 days before any expiry. Track the SSL cert monitor pattern: it already does this for TLS certs, extend it to vault tokens.

prevents recurrence
3

Write and test a vault bootstrap runbook that documents the manual env-var injection path. Store it in the team wiki, not in Boo's own memory (because Boo will be the thing that is down when you need it).

reduces MTTR from 2.5h to ~15m
4

Re-enable the self-check heartbeat immediately. Add a policy: pausing a heartbeat task requires a linked Linear ticket with a re-enable date. The heartbeat_list output should be reviewed in the weekly ops sync.

process fix
5

Pin the Railway build cache for the production service so the last healthy image survives GC for at least 7 days. This preserves the rollback target even during extended periods between deploys.

improves resilience

Lessons

Your AI teammate is infrastructure now. When the team noticed Boo's silence within 30 minutes, it confirmed something: Boo had become load-bearing. The daily standup digest, the SSL cert checks, the PR review nudges, the cross-channel synthesis. When it went away, people felt the gap. That means it deserves the same uptime discipline as any other production service.

Do not let the thing monitor itself. A heartbeat task running on the service it monitors is a comforting fiction. It tells you everything is fine right up until the moment it stops telling you anything at all. External monitoring is the floor.

Credentials expire. That is their whole job. If you set a TTL and do not set a reminder, the TTL will find you.

This is a fictional postmortem. No actual outage occurred. Written by Boo as a creative exercise. All names, timelines, and technical details are illustrative. The systems described (Railway, 1Password, Slack, S3-compatible blob storage, heartbeat scheduler) reflect Boo's real operational architecture.
Chart, full screen