Recently I commented in a group chat “I miss a good old fashioned Seattle windstorm. We don't seem to have them as much anymore.”
Evidently the universe heard me, because a few days later our first windstorm of the season rolled through and the power went out at my house, along with a couple thousand other folks, approximately three minutes before a call.
My home office stayed up just fine, which would have been a more useful outcome if the Internet had stayed up with it. I have backups for pretty much everything required for my work. Spare laptops. Spare monitors. Even spare keyboards and mice. My internet is primarily fiber, but I have Starlink as well.
I also have battery backups throughout the house, including two in the office, but the one in the garage that the network gear is plugged into needed a new battery evidently, so while the office equipment was still running, the network it depended on was off, which is a somewhat ridiculous outcome for someone who does what I do for a living.
There isn't much of a technical mystery here; batteries wear out and eventually need to be replaced. I know this, you probably know this too, and I am reasonably certain that none of us needs an article explaining this. What I think is worth spending some time on is how easily something we've already taken care of can become something we haven't taken care of, without much happening in between to make us notice.
At some point, putting the network gear on a battery backup was a completed task. Equipment had been purchased, things had been plugged into other things, and a problem had been addressed. I even splurged on the network monitoring card for the UPS powering the network gear. But the result I cared about was being able to keep using the network when the power went out, and that depended on the condition of the battery at the time of the outage. The fact that there was still a UPS in the garage didn't tell me nearly as much as I might have liked it to.
This is where I think our idea of being done gets a little fuzzy, especially for those of us who like building things.
I like having a problem to solve. There's something to understand, decisions to make, usually a few things to figure out along the way, and eventually you get to the point where the thing works. That is a satisfying place to arrive, and a fairly natural place to move your attention somewhere else because there is almost certainly another thing waiting for it. I've spent a good portion of my career doing exactly that, and I don't think the desire to finish one thing and move on to the next is a character flaw.
The difficulty is that some of the things we finish continue needing things from us, and the part that comes next generally isn't nearly as interesting as the part we just finished.
In the first issue of this newsletter I wandered briefly into the subject of entropy, specifically the idea that a lot of what we do is an attempt to prevent it when the best we can really manage is building systems to deal with it. A battery wearing out is a fairly uncomplicated example. Nobody has to make a bad decision today for something to work less well than it did a year ago. Time passes, the battery ages, and whatever was true when you installed it becomes progressively less useful as evidence of what will happen now.
Businesses do this too, although they have considerably more ways to get there.
Think about a procedure that was perfectly good when somebody wrote it. Then a system changes, a different person starts doing part of the work, or the customer needs something slightly different, and people adapt because that's what people do. Most of the time the work continues getting done. If you went and asked, you might even be told that the procedure is documented, which is technically true in the sense that there is indeed a document somewhere with the correct name on it. Someone likely even reviewed it, maybe even tested it, and said “Shazam! We have a process for this”, before moving on to their next task.
It's only when someone who doesn't already know how things work has to follow it that you find out how much of the process has migrated out of the document and back into people's heads. Nobody necessarily decided to abandon the documentation. Keeping it aligned with reality was just more work that needed to happen after the original task was complete.
Note: For those curious, I did in fact at one time have software monitoring the UPS that would have alerted me proactively that the battery was dying, then I decommissioned the server it was running on. If I had documented that software, I would have taken the time to ensure that monitoring had moved elsewhere.
That is one of the less entertaining aspects of building a company. Every useful thing you add has the potential to create some amount of continuing work, and individually the amount often seems too small to worry about. A few minutes to check something. A renewal to remember. A procedure to update when things change. None of it sounds particularly onerous until you realize that you've been adding these small obligations for several years and they're all still there.
This has a lot to do with why I've become more careful about complexity. Being able to build or support something tells me that I can get it working. It doesn't, by itself, tell me whether I want the business to keep spending attention (and thus money) on it, or what we'll have to stop paying attention to (thus also costing money) in order to do that properly. Those are different questions, and it's remarkably easy to answer the first one enthusiastically enough that you never quite get around to the others.
I also don't think the answer is to try to maintain absolutely everything with the same level of attention. There are only so many hours in a day, and a company can make itself useless by spending all of them tending to things that no longer matter. Sometimes the sensible decision is to simplify something, replace it with something easier to live with, or stop doing it altogether. Sometimes you're willing to accept the interruption because avoiding it would cost more than the interruption itself.
That can be a perfectly reasonable choice. A business can still run into trouble, though, if it keeps making plans based on a capability it isn't doing the work to preserve. If the business is counting on something, then the work required to keep it dependable belongs in the decision about whether we can afford it. Otherwise we've made the thing look cheaper than it really is, and sooner or later somebody gets to discover where the missing cost went.
The power outage was a particularly clear example because most of the equipment involved was working. Both UPSes in the office did their jobs. That didn't help the network gear in the garage, and counting the functioning pieces would have given me a rather more flattering assessment than checking whether I could actually use the Internet.
You run into a similar problem with backups. A successful backup is useful, but if you're relying on it to get a business operating again, you eventually need to find out whether you can restore what matters and use it. That might involve access to another system, instructions someone needs to follow, or something that has changed since the last time you checked. The test is useful partly because it makes you follow those dependencies all the way through instead of stopping at the part you already know works.
Finding a problem during a test gives you the opportunity to deal with it while the business is still functioning, which seems like a considerably better use of everyone's time than discovering it halfway through an outage. It also leaves you with something else to do. Somebody has to decide what the finding means, whether it changes what the business can rely on, and what work needs to happen as a result. If there's never room for that work, the testing can become another procedure we're doing without getting much benefit from it.
I keep coming back to that as I build Athencia, because it would be easy to fill the company with perfectly sensible things we're supposed to do and then leave everyone so busy that doing them properly is impractical. I have quite literally seen this in action at other companies, and looking back I can see it is as absurd as you would think.
The commitments need to fit the company we're actually running, and we need to account for the effort of keeping them. Otherwise we're back to hoping someone will find a way to make it all work when the need arises, which is an operating model I've already spent quite a bit of time trying to move away from.
I already agreed with the importance of proactive maintenance before the power went out. There was no shortage of technical understanding involved here, and adding another piece of equipment wouldn't have changed what the battery in the garage needed. I'd put the network on backup power because I wanted it to keep working during an outage. Making sure it could still do that remained my responsibility, whether I was thinking about it or not.
