It's 3am and a payment service is failing. The engineer who answers has never touched it, and the next hour goes to finding someone who can fix it.
In this episode, Lee goes further than his article "When Everything Is Critical, Nothing Is." The article named the two missing decisions. This episode is about how to make them: which services matter most, and which one team owns each of them.
In this episode:
Links:
Hello and welcome to Software Architecture Insights, your go-to
Lee:resource for empowering software architects and aspiring professionals
Lee:with the knowledge and tools they require to navigate the complex
Lee:landscape of modern software design.
Lee:Last month, I wrote about a 3:00 AM page.
Lee:A payment service was failing, and the engineer who answered
Lee:the page never touched it.
Lee:The next hour was spent finding someone who actually knew how to fix it.
Lee:That article was called "When Everything is Critical, Nothing Is."
Lee:Today, I want to go further than the article had room to discuss.
Lee:The article told you what's missing: two decisions, which
Lee:services matter and who owns them.
Lee:What it didn't tell you is how to make those decisions, and that's
Lee:where most teams tend to get stuck.
Lee:So that's this episode, how you add a service tier to your services,
Lee:and how you make ownership stick. Let's start with a quick recap
Lee:in case you missed the article.
Lee:By the way, if you'd like to read the original article, there's a link to
Lee:it in the show notes. When an outage runs long, we usually blame the
Lee:technology or we blame communication.
Lee:But look at where the time actually goes.
Lee:In a story from the article, the actual problem took 11 minutes to
Lee:find, and the other 62 minutes went to finding a person responsible.
Lee:No technical fix would have saved those 62 minutes.
Lee:What was missing were two key decisions.
Lee:The first is criticality.
Lee:Which of your services matter the most?
Lee:The second decision is ownership.
Lee:For each service, which one team, one single team, is accountable for
Lee:keeping that service up and running?
Lee:Now, most organizations have made neither decision, or even worse, they
Lee:think they've made the decision, and the answer on paper doesn't really match what
Lee:really happens at 3:00 in the morning.
Lee:So let's start with service tiers.
Lee:I've used a four-tier model for a long time now.
Lee:I wrote about it in my book, "Architecting for Scale."
Lee:Here's the short version of that.
Lee:A Tier 1 service is a service where if it goes down, the
Lee:business is hurt right away.
Lee:Customers can't buy, they can't log in, and money stops flowing into the business.
Lee:A Tier 1 service outage means the building is on fire.
Lee:A Tier 2 service is a service where failure hurts, but
Lee:the business can keep going.
Lee:Customers will notice something is wrong.
Lee:You know, maybe search is slow or not working, or recommendations are
Lee:missing from products, or something like that, but things generally still work.
Lee:And most importantly, the business can keep moving forward.
Lee:Money is still being made.
Lee:Commitments are still being kept. A Tier 3 service is a service
Lee:whose failure most customers wouldn't even notice right away.
Lee:Maybe eventually, but not right away.
Lee:Something in the background, like an email digest or a report that runs
Lee:overnight, or something like that.
Lee:Those are examples of Tier 3 services.
Lee:A Tier 4 service is internal and doesn't touch customers at all.
Lee:If a Tier 4 service goes down, somebody on your team might get annoyed, but that's
Lee:really the full extent of the impact.
Lee:Now, the definitions are the easy part.
Lee:Most teams can agree on those definitions in about 10 minutes.
Lee:The hard part, though, is applying them to your application.
Lee:In this article, I described an exercise where 53 services got sorted,
Lee:and 48 of them came back as Tier 1.
Lee:I've seen versions of that sort of problem occurring more than once.
Lee:And I want to be fair to the people that are doing it because
Lee:nobody's being dishonest here.
Lee:Each team gets asked, "Is your service important?"
Lee:And every team says, "Yes, of course," because it is
Lee:important, at least it is to them.
Lee:The problem is the question.
Lee:Is it important?
Lee:We'll always get a yes answer.
Lee:After all, if a service isn't important, well, why does it exist at all?
Lee:So change the question.
Lee:Here are three questions that I recommend asking instead of
Lee:the is it important question.
Lee:Start with this one: What happens to a customer in the first
Lee:hour this service goes down?
Lee:Not the first day, the first hour.
Lee:If the answer is they can't complete a purchase, then this
Lee:service is a Tier 1 service.
Lee:But if the answer is something more like a nightly report is late, then
Lee:this most definitely is not a Tier 1 service. Then ask this question.
Lee:If this service and the checkout service both fail at the exact same
Lee:time, which one do you fix first?
Lee:Everyone knows the answer.
Lee:The checkout service.
Lee:The checkout service is probably one of your most critical services, period, if
Lee:you're a e-commerce application at least.
Lee:Because without the checkout service, the business itself is dead. But the moment
Lee:you ask the question out loud, the moment you ask how does your service compare to
Lee:the checkout service, you've ranked the two services against each other. And the
last question:would you pay for a second cloud region for this service in order
last question:to improve availability, or would you put an engineer on call for it overnight?
last question:Making a service Tier 1 comes with a price tag.
last question:It costs money to duplicate a service for availability, and on-call engineers are
last question:expensive to your other projects as well.
last question:If nobody's willing to pay for those things, well, then the service just
last question:plain can't be a Tier 1 service.
last question:That last question changes the conversation completely because every tier
last question:assignment is now a budget decision. One more thing that helps is putting a limit
last question:on the number of Tier 1 services you're allowed to have in your application.
last question:Pick a number.
last question:Maybe it's five, maybe it's eight, maybe it's fifteen.
last question:The right number absolutely depends on your application.
last question:But pick one and write it down.
last question:Then, if a team wants to add a new service to the Tier 1 list and it,
last question:the list is already full, something else has to come off the list.
last question:Or they have to make the case in front of everyone why the
last question:limit needs to be increased.
last question:Now, this may sound bureaucratic, but in practice, it moves the ranking
last question:conversation into a meeting on a Tuesday afternoon, which is a whole
last question:lot better than a bridge call at three o'clock in the morning during a crisis.
last question:You're going to rank your services either way.
last question:You can do it calmly ahead of time, or you can let whoever answers the
last question:page do it in the middle of the night.
last question:It's your choice.
last question:Now, something this article didn't get into, and it trips up almost
last question:every team that tiers for the first time, and that is dependencies.
last question:Say the checkout service is a Tier 1 service.
last question:Now, checkout calls a tax calculation service.
last question:Somebody ranked the tax calculation service as a Tier 3 service
last question:because it's small and nobody really thinks that much about it.
last question:It was a poor ranking, but that's what they came up with.
last question:But what happens when the tax service goes down?
last question:Well, in most cases, checkout goes down with it.
last question:Your Tier 1 service is only as reliable as the least reliable
last question:thing that it can't live without.
last question:So here's the rule I use.
last question:When a Tier 1 service, such as the checkout service, depends on a lower
last question:tier service, you have two choices.
last question:You raise the dependent service to be a Tier 1 status service on its own right,
last question:with everything that comes with that and all the costs that are associated with
last question:that, or you make the Tier 1 service, like the checkout service, able to survive
last question:without the dependency. Maybe checkout can use a cached tax rate for a few minutes.
last question:Maybe it lets the order through and calculates the tax later.
last question:Maybe it turns off one feature instead of failing the whole page.
last question:Either raise the tier of the dependency or make the dependency itself optional.
last question:Either choice works.
last question:The one that fails is a Tier 1 service quietly depending on a Tier 3 service
last question:as a necessity with nobody noticing until the night it really matters.
last question:When you do your first tiering pass, walk the dependencies
last question:of every Tier 1 service.
last question:That's usually where some of the surprises are going to show up.
last question:Okay, let's talk about ownership, the second question.
last question:Every service has exactly one owning team.
last question:The team is named, and the team is current.
last question:That's part of a principle that I call STOSA, Single Team
last question:Oriented Service Architecture.
last question:One service, one owning team.
last question:The rule is simple, but living it is actually quite a bit harder.
last question:Owner is one of those words people use and nobody really defines.
last question:So let's try coming up with a definition of ownership.
last question:An owning team is the team that gets paged when a service breaks.
last question:It's the team that has the service on its roadmap, so
last question:upgrades and fixes get scheduled.
last question:And it's the team that can change the service and deploy it without
last question:asking anybody's permission.
last question:Paged, planned, permitted.
last question:If a team has all three of those attributes, it owns the service.
last question:If it's missing any one of them, it doesn't.
last question:So here's a test you can run.
last question:Pick a service, any service.
last question:Ask the team that's listed as the owner three questions.
last question:First, when this breaks at 3:00 AM, does your phone ring?
last question:Is there work for this service in your plan for this quarter?
last question:Could you deploy a change to it today without asking another team?
last question:If the answer to all three of those questions is yes,
last question:great, you have ownership.
last question:But what you'll often find is two yeses and a no.
last question:The team gets paged, but they can't deploy without also calling in the platform team.
last question:Or they can deploy it, but it's never on their roadmap, so nothing
last question:gets upgraded until it breaks.
last question:That no is where your next long outage is coming from.
last question:The hardest ownership cases are inherited services.
last question:Something written by a team that no longer exists, or by one person who left.
last question:Nobody wants those, and I understand why.
last question:Taking ownership means taking the 3:00 AM pages for code that you didn't
last question:write and don't fully understand.
last question:So don't hand it over and walk away.
last question:If you're asking a team to own an inherited service,
last question:give them time to learn it.
last question:Put that time on the roadmap.
last question:Let them fix the worst of the problems that occur in a service
last question:before they're on the hook overnight.
last question:And if a service isn't worth that level of investment, well,
last question:that tells you something as well.
last question:Maybe it belongs in Tier 4 instead of Tier 3 or two.
last question:Maybe the service should be retired.
last question:And that's where the two decisions start working together. The article made this
last question:point, but I'll make it again because it's the heart of the whole matter.
last question:A tier without an owner is a promise nobody made.
last question:You can call checkout a Tier 1 service, and if no single team is accountable
last question:for that, nobody is going to defend it when a deadline shows up. An owner
last question:without a tier is a team defending everything at equal priority, which means
last question:really defending nothing in particular.
last question:They'll spend their effort on whatever broke most recently, not
last question:what is most critical. Put those two decisions together and each one
last question:makes the other one work harder.
last question:The tier tells the team how much reliability you can buy, and
last question:ownership gives them the authority and the obligation to buy it. So
last question:here's what I'd do this week.
last question:Take your 10 busiest services.
last question:For each one, write down the tier and the owning team from memory,
last question:not from the service catalog.
last question:Then check the catalog.
last question:Then ask the owning teams the three questions: paged, planned, permitted.
last question:For your Tier 1 services, walk their dependencies.
last question:Find the lower tier service your Tier 1 service just can't live without.
last question:You don't have to fix everything you find, but you'll know where your
:00 AM page is probably going to come from, and you'll know who's
:going to answer it. The original article is linked in the show notes.
:So is the ownership gap diagnostic that I mentioned there.
:It's a spreadsheet, a simple spreadsheet that you can download that walks you
:through all of this for your own services.
Lee:Thank you for joining us on Software Architecture Insights.
Lee:If you found this episode interesting, please tell your friends and colleagues.
Lee:You can listen to Software Architecture Insights on all
Lee:of the major podcast platforms.
Lee:And if you want more from me, take a look at some of my many articles
Lee:at softwarearchitectureinsights.com.
Lee:And while you're there, join the 2000 people who have subscribed to my
Lee:newsletter, so you always get my latest content as soon as it's available.
Lee:Thank you for listening to Software Architecture Insights.