Why I’m building an AI SRE for a one-person company
I’ve been the entire on-call rotation since 2008. I’m building an AI agent to take some of that work off my plate.
It is never anything interesting. A disk at 94 percent. A job queue that has stopped moving. A server that wants a reboot. The alert arrives, I open a terminal, and for the next twenty minutes I am a systems administrator with no colleagues and no handover notes except the ones I wrote myself, years ago, in a wiki with pages called Queue overload mitigation and Recovering software RAID.
WebTranslateIt has been in production since 2008. These days I also run Trackberry, and Distill, an internal tool I’ve since written about. That makes three Rails apps on three servers at two hosting providers, looked after by a team the same size it was in 2008: me.
I want something to take part of that work off my plate. I thought about it a lot over the summer, and started building it this week. So treat this as a build log, not a success story.
Ops with a team of one
The monitoring is fine. I run a self-hosted SigNoz and the alerts fire when they should. But every alert still ends with me opening a terminal.
Outages aren’t actually the expensive part; they’re rare. It’s the little things that keep stealing my attention: the reboot a kernel update is waiting on, the SigNoz upgrade I do by hand every month. Each one is ten minutes, which is exactly why I’ve never automated any of them.
Why not more scripts
Scripts handle the things I already know how to handle. The incidents that cost me time need a little judgment. Is this disk growth normal? Which service do I restart first?
I’ve tried the obvious solution: hire someone who can share the on-call. I’ve never found a developer who wanted the job. I’ve been doing this for eighteen years, and I can hardly blame them. Being woken up to investigate a broken background job on someone else’s infrastructure isn’t exactly why most people become developers. Managed ops makes little sense for three servers either.
So I want something in between: something that reads the runbook, checks the signals, does the investigation, and eventually carries out a few repairs within limits I set.
What I’m building
It lives in one repository: checks, runbooks, rules and agent code together. I want to be exact about what exists, because posts like this one usually aren’t.
systemd timer, every 5 min
|
watcher -- read-only probe over SSH --> the three apps
|
| failing app + evidence
v
agent -- rules, runbook, SSH output --> Claude
| <-- commands to run, diagnosis ---
v
Slack -- later: approve --> allowlisted action
Running today: the watcher. A systemd timer on a box I already operate runs one check per app every five minutes. To reach the other hosts it uses an SSH key that each server restricts to a single read-only script. The key cannot open a shell. None of it runs in GitHub Actions, on purpose: scheduled workflows there are best-effort, and a watcher that silently skips ticks is worse than no watcher.
A check never stops at the first failure. It gathers every signal it can and reports the healthy ones too, so a failure arrives in Slack as one line like this:
wti: goodjob: worker heartbeat is 2840s stale (limit 120s)
[ok: http 200 from webtranslateit.com/up;
postgres active, 82/450 connections;
goodjob 0 failures in the last hour]
The site is up, the database is fine, nothing is erroring, and the worker hasn’t checked in for 47 minutes. That’s enough context for me to know where I’d start looking.
The agent exists, but I don’t trust it yet. It’s a small Python program that takes that line, loads the rules and the app’s runbook, and calls Claude with exactly one tool: run a command over SSH. The first version is read-only. The runbook has a row for this symptom: a heartbeat row that is present but stale means the worker may be hung, and restarting it is allowed with approval. The agent’s job is to gather the evidence and decide whether that row really applies, or whether this is something else wearing the same symptom. What I want back is something like:
Hypothesis worker container is up but hung
Evidence docker ps: jobs container running, up 9 days
good_job_processes: 1 row, updated 47 min ago
pg_locks: nothing waiting
oldest due job: 2790s and growing
Recommended restart the jobs container
(runbook: allowed, needs approval)
Confidence high
I wrote that report by hand. The agent hasn’t diagnosed a real incident yet, and the first few weeks will be me comparing what it says with what I would have said.
After that comes a short list of repairs: restart a service named in the runbook, clear listed temp directories, kill a runaway process that matches a known failure mode, retry a background job. Each one sits behind an approve button in Slack until I decide, action by action, to take the button away.
The chores I opened with aren’t on that list. Night reboots and the monthly upgrade are what I most want to stop doing, and for now both are forbidden: no rebooting a host, no installing packages. I’d rather earn my way to them.
It has already found something
Before the agent has fixed anything, the project has found something I had missed, and it had nothing to do with AI. The first phase was just me filling in runbooks and checking them against the live hosts. Doing that, I found that WebTranslateIt’s database backup had silently not run for three weeks. The container’s cron daemon ignores a crontab that isn’t owned by root, and the deploy tool mounted it as another user. When I looked in the storage bucket the nightly dumps are uploaded to, there was one file where there should have been about twenty. Nothing anywhere had said so.
I fixed the ownership problem, and every backup now pings an outside service that complains when it goes quiet. I am precisely the kind of operator who needs a second pair of eyes.
Where the guardrails are
The runbooks are the contract. If it isn’t written down, the agent doesn’t
do it; it escalates and never improvises. I care more about what it can’t do
than what it can. Nothing touches Postgres data, changes configuration, deploys code,
or manages users and keys. And those aren’t instructions to the model. The agent gets its own SSH account, separate from the watcher’s key, and the
account’s sudoers file lists the exact commands it may run. A prompt is a request. A sudoers file is a fact.
It hands over to me when the fix isn’t on the list, when a fix didn’t work, or when what it finds doesn’t match what the runbook describes. For now its confidence also has to be high. Every run reports, including the ones that find nothing, so I can always read what happened while I slept.
What I expect to go wrong
I don’t have to guess at all of it, because some has already happened.
The first time the probe ran against production it reported a worker heartbeat 7,209 seconds stale, against a limit of 120, on a healthy server. The job tables store UTC in a column with no timezone, that host’s database session is set to Berlin, and every age came out exactly two hours too large. The other two hosts run on UTC, so the code was right on two machines out of three, which is the worse way to be wrong. My self-test didn’t catch it; it feeds the checks fixed values, not SQL.
So false positives aren’t an edge case. They’re going to happen. Right now every failing run posts, with no deduplication, because I want to see how noisy it actually is before I decide how to deal with it. If I start ignoring the Slack channel because there are too many false alarms, the whole thing has failed.
What worries me more is an agent that finds a plausible explanation and then does the wrong thing. A model that says “I don’t know” is cheap. One that reads a log, builds a plausible story and restarts the wrong thing is expensive, and asking it to rate its own confidence is only a partial defence. That’s why version one can’t act, version two needs a button, and Postgres is off the table entirely.
And now I have another system to maintain. The watcher shares a host with Distill and SigNoz, so if I lose that one machine I lose an app, the alerting and the thing watching them, all together. I’m betting it breaks less often than what it watches. I don’t know yet whether that’s true.
Beyond my servers
I stayed small on purpose, and I’ve written at length about how that happened. The trade-off has always been that there’s a limit to how much I can operate myself. If an agent can take over some of the boring operational work without taking over the important decisions, I can keep more infrastructure running without adding another person to the rota. Maybe that changes what a company of one can operate. I don’t know yet. That’s what I’m trying to find out, on my own servers, with eighteen years of runbooks and an agent that isn’t allowed to touch the database.
What’s next
Read-only diagnosis first, then the button. I’ll write up what happens the first time it fixes something without me, and just as plainly if the first thing it does is get something wrong. If you’re building something similar for a company of one or two, I’d like to compare notes; I’m on LinkedIn.
Málaga, 18 September 2026.