Artwork for podcast Software Architecture Insights
When Everything Is Critical, Nothing Is
Episode 13 • 29th September 2026 • Software Architecture Insights • Lee Atchison
00:00:00 00:16:01

Share Episode

Shownotes

It's 3am and a payment service is failing. The engineer who answers has never touched it, and the next hour goes to finding someone who can fix it.

In this episode, Lee goes further than his article "When Everything Is Critical, Nothing Is." The article named the two missing decisions. This episode is about how to make them: which services matter most, and which one team owns each of them.

In this episode:

  • The four service tiers, in plain terms
  • Three questions that keep everything from landing in Tier 1
  • Why every tier is a budget decision, and why you should cap Tier 1
  • The dependency trap: a Tier 1 service is only as reliable as the lower-tier service it can't live without
  • What "owner" means: paged, planned, permitted
  • Inherited services, and an exercise to run this week

Links:

Transcripts

Lee:

Hello and welcome to Software Architecture Insights, your go-to

Lee:

resource for empowering software architects and aspiring professionals

Lee:

with the knowledge and tools they require to navigate the complex

Lee:

landscape of modern software design.

Lee:

Last month, I wrote about a 3:00 AM page.

Lee:

A payment service was failing, and the engineer who answered

Lee:

the page never touched it.

Lee:

The next hour was spent finding someone who actually knew how to fix it.

Lee:

That article was called "When Everything is Critical, Nothing Is."

Lee:

Today, I want to go further than the article had room to discuss.

Lee:

The article told you what's missing: two decisions, which

Lee:

services matter and who owns them.

Lee:

What it didn't tell you is how to make those decisions, and that's

Lee:

where most teams tend to get stuck.

Lee:

So that's this episode, how you add a service tier to your services,

Lee:

and how you make ownership stick. Let's start with a quick recap

Lee:

in case you missed the article.

Lee:

By the way, if you'd like to read the original article, there's a link to

Lee:

it in the show notes. When an outage runs long, we usually blame the

Lee:

technology or we blame communication.

Lee:

But look at where the time actually goes.

Lee:

In a story from the article, the actual problem took 11 minutes to

Lee:

find, and the other 62 minutes went to finding a person responsible.

Lee:

No technical fix would have saved those 62 minutes.

Lee:

What was missing were two key decisions.

Lee:

The first is criticality.

Lee:

Which of your services matter the most?

Lee:

The second decision is ownership.

Lee:

For each service, which one team, one single team, is accountable for

Lee:

keeping that service up and running?

Lee:

Now, most organizations have made neither decision, or even worse, they

Lee:

think they've made the decision, and the answer on paper doesn't really match what

Lee:

really happens at 3:00 in the morning.

Lee:

So let's start with service tiers.

Lee:

I've used a four-tier model for a long time now.

Lee:

I wrote about it in my book, "Architecting for Scale."

Lee:

Here's the short version of that.

Lee:

A Tier 1 service is a service where if it goes down, the

Lee:

business is hurt right away.

Lee:

Customers can't buy, they can't log in, and money stops flowing into the business.

Lee:

A Tier 1 service outage means the building is on fire.

Lee:

A Tier 2 service is a service where failure hurts, but

Lee:

the business can keep going.

Lee:

Customers will notice something is wrong.

Lee:

You know, maybe search is slow or not working, or recommendations are

Lee:

missing from products, or something like that, but things generally still work.

Lee:

And most importantly, the business can keep moving forward.

Lee:

Money is still being made.

Lee:

Commitments are still being kept. A Tier 3 service is a service

Lee:

whose failure most customers wouldn't even notice right away.

Lee:

Maybe eventually, but not right away.

Lee:

Something in the background, like an email digest or a report that runs

Lee:

overnight, or something like that.

Lee:

Those are examples of Tier 3 services.

Lee:

A Tier 4 service is internal and doesn't touch customers at all.

Lee:

If a Tier 4 service goes down, somebody on your team might get annoyed, but that's

Lee:

really the full extent of the impact.

Lee:

Now, the definitions are the easy part.

Lee:

Most teams can agree on those definitions in about 10 minutes.

Lee:

The hard part, though, is applying them to your application.

Lee:

In this article, I described an exercise where 53 services got sorted,

Lee:

and 48 of them came back as Tier 1.

Lee:

I've seen versions of that sort of problem occurring more than once.

Lee:

And I want to be fair to the people that are doing it because

Lee:

nobody's being dishonest here.

Lee:

Each team gets asked, "Is your service important?"

Lee:

And every team says, "Yes, of course," because it is

Lee:

important, at least it is to them.

Lee:

The problem is the question.

Lee:

Is it important?

Lee:

We'll always get a yes answer.

Lee:

After all, if a service isn't important, well, why does it exist at all?

Lee:

So change the question.

Lee:

Here are three questions that I recommend asking instead of

Lee:

the is it important question.

Lee:

Start with this one: What happens to a customer in the first

Lee:

hour this service goes down?

Lee:

Not the first day, the first hour.

Lee:

If the answer is they can't complete a purchase, then this

Lee:

service is a Tier 1 service.

Lee:

But if the answer is something more like a nightly report is late, then

Lee:

this most definitely is not a Tier 1 service. Then ask this question.

Lee:

If this service and the checkout service both fail at the exact same

Lee:

time, which one do you fix first?

Lee:

Everyone knows the answer.

Lee:

The checkout service.

Lee:

The checkout service is probably one of your most critical services, period, if

Lee:

you're a e-commerce application at least.

Lee:

Because without the checkout service, the business itself is dead. But the moment

Lee:

you ask the question out loud, the moment you ask how does your service compare to

Lee:

the checkout service, you've ranked the two services against each other. And the

last question:

would you pay for a second cloud region for this service in order

last question:

to improve availability, or would you put an engineer on call for it overnight?

last question:

Making a service Tier 1 comes with a price tag.

last question:

It costs money to duplicate a service for availability, and on-call engineers are

last question:

expensive to your other projects as well.

last question:

If nobody's willing to pay for those things, well, then the service just

last question:

plain can't be a Tier 1 service.

last question:

That last question changes the conversation completely because every tier

last question:

assignment is now a budget decision. One more thing that helps is putting a limit

last question:

on the number of Tier 1 services you're allowed to have in your application.

last question:

Pick a number.

last question:

Maybe it's five, maybe it's eight, maybe it's fifteen.

last question:

The right number absolutely depends on your application.

last question:

But pick one and write it down.

last question:

Then, if a team wants to add a new service to the Tier 1 list and it,

last question:

the list is already full, something else has to come off the list.

last question:

Or they have to make the case in front of everyone why the

last question:

limit needs to be increased.

last question:

Now, this may sound bureaucratic, but in practice, it moves the ranking

last question:

conversation into a meeting on a Tuesday afternoon, which is a whole

last question:

lot better than a bridge call at three o'clock in the morning during a crisis.

last question:

You're going to rank your services either way.

last question:

You can do it calmly ahead of time, or you can let whoever answers the

last question:

page do it in the middle of the night.

last question:

It's your choice.

last question:

Now, something this article didn't get into, and it trips up almost

last question:

every team that tiers for the first time, and that is dependencies.

last question:

Say the checkout service is a Tier 1 service.

last question:

Now, checkout calls a tax calculation service.

last question:

Somebody ranked the tax calculation service as a Tier 3 service

last question:

because it's small and nobody really thinks that much about it.

last question:

It was a poor ranking, but that's what they came up with.

last question:

But what happens when the tax service goes down?

last question:

Well, in most cases, checkout goes down with it.

last question:

Your Tier 1 service is only as reliable as the least reliable

last question:

thing that it can't live without.

last question:

So here's the rule I use.

last question:

When a Tier 1 service, such as the checkout service, depends on a lower

last question:

tier service, you have two choices.

last question:

You raise the dependent service to be a Tier 1 status service on its own right,

last question:

with everything that comes with that and all the costs that are associated with

last question:

that, or you make the Tier 1 service, like the checkout service, able to survive

last question:

without the dependency. Maybe checkout can use a cached tax rate for a few minutes.

last question:

Maybe it lets the order through and calculates the tax later.

last question:

Maybe it turns off one feature instead of failing the whole page.

last question:

Either raise the tier of the dependency or make the dependency itself optional.

last question:

Either choice works.

last question:

The one that fails is a Tier 1 service quietly depending on a Tier 3 service

last question:

as a necessity with nobody noticing until the night it really matters.

last question:

When you do your first tiering pass, walk the dependencies

last question:

of every Tier 1 service.

last question:

That's usually where some of the surprises are going to show up.

last question:

Okay, let's talk about ownership, the second question.

last question:

Every service has exactly one owning team.

last question:

The team is named, and the team is current.

last question:

That's part of a principle that I call STOSA, Single Team

last question:

Oriented Service Architecture.

last question:

One service, one owning team.

last question:

The rule is simple, but living it is actually quite a bit harder.

last question:

Owner is one of those words people use and nobody really defines.

last question:

So let's try coming up with a definition of ownership.

last question:

An owning team is the team that gets paged when a service breaks.

last question:

It's the team that has the service on its roadmap, so

last question:

upgrades and fixes get scheduled.

last question:

And it's the team that can change the service and deploy it without

last question:

asking anybody's permission.

last question:

Paged, planned, permitted.

last question:

If a team has all three of those attributes, it owns the service.

last question:

If it's missing any one of them, it doesn't.

last question:

So here's a test you can run.

last question:

Pick a service, any service.

last question:

Ask the team that's listed as the owner three questions.

last question:

First, when this breaks at 3:00 AM, does your phone ring?

last question:

Is there work for this service in your plan for this quarter?

last question:

Could you deploy a change to it today without asking another team?

last question:

If the answer to all three of those questions is yes,

last question:

great, you have ownership.

last question:

But what you'll often find is two yeses and a no.

last question:

The team gets paged, but they can't deploy without also calling in the platform team.

last question:

Or they can deploy it, but it's never on their roadmap, so nothing

last question:

gets upgraded until it breaks.

last question:

That no is where your next long outage is coming from.

last question:

The hardest ownership cases are inherited services.

last question:

Something written by a team that no longer exists, or by one person who left.

last question:

Nobody wants those, and I understand why.

last question:

Taking ownership means taking the 3:00 AM pages for code that you didn't

last question:

write and don't fully understand.

last question:

So don't hand it over and walk away.

last question:

If you're asking a team to own an inherited service,

last question:

give them time to learn it.

last question:

Put that time on the roadmap.

last question:

Let them fix the worst of the problems that occur in a service

last question:

before they're on the hook overnight.

last question:

And if a service isn't worth that level of investment, well,

last question:

that tells you something as well.

last question:

Maybe it belongs in Tier 4 instead of Tier 3 or two.

last question:

Maybe the service should be retired.

last question:

And that's where the two decisions start working together. The article made this

last question:

point, but I'll make it again because it's the heart of the whole matter.

last question:

A tier without an owner is a promise nobody made.

last question:

You can call checkout a Tier 1 service, and if no single team is accountable

last question:

for that, nobody is going to defend it when a deadline shows up. An owner

last question:

without a tier is a team defending everything at equal priority, which means

last question:

really defending nothing in particular.

last question:

They'll spend their effort on whatever broke most recently, not

last question:

what is most critical. Put those two decisions together and each one

last question:

makes the other one work harder.

last question:

The tier tells the team how much reliability you can buy, and

last question:

ownership gives them the authority and the obligation to buy it. So

last question:

here's what I'd do this week.

last question:

Take your 10 busiest services.

last question:

For each one, write down the tier and the owning team from memory,

last question:

not from the service catalog.

last question:

Then check the catalog.

last question:

Then ask the owning teams the three questions: paged, planned, permitted.

last question:

For your Tier 1 services, walk their dependencies.

last question:

Find the lower tier service your Tier 1 service just can't live without.

last question:

You don't have to fix everything you find, but you'll know where your

:

00 AM page is probably going to come from, and you'll know who's

:

going to answer it. The original article is linked in the show notes.

:

So is the ownership gap diagnostic that I mentioned there.

:

It's a spreadsheet, a simple spreadsheet that you can download that walks you

:

through all of this for your own services.

Lee:

Thank you for joining us on Software Architecture Insights.

Lee:

If you found this episode interesting, please tell your friends and colleagues.

Lee:

You can listen to Software Architecture Insights on all

Lee:

of the major podcast platforms.

Lee:

And if you want more from me, take a look at some of my many articles

Lee:

at softwarearchitectureinsights.com.

Lee:

And while you're there, join the 2000 people who have subscribed to my

Lee:

newsletter, so you always get my latest content as soon as it's available.

Lee:

Thank you for listening to Software Architecture Insights.

Chapters

Video

More from YouTube