The New Failure Mode: Everything Is Working

Why healthy components can still produce an unhealthy system

There is a particularly frustrating kind of production incident.

The API is returning 200.

The Kubernetes pods are healthy.

CPU and memory are normal.

The database is responding.

The message broker has no significant backlog.

The monitoring dashboard is mostly green.

And yet, the user cannot complete the most important action in the application.

Nothing appears to be broken.

Everything is working.

Except the system.

This is becoming one of the most important reliability problems in modern distributed software: the difference between component health and system health.


The Green Dashboard Illusion

Traditional monitoring often starts with a simple assumption:

If the important components are healthy, the application should be healthy.

That assumption worked reasonably well when applications were relatively self-contained.

Modern systems are different.

A single user workflow can cross:

Browser
   ↓
CDN
   ↓
API Gateway
   ↓
Authentication
   ↓
Service A
   ↓
Message Broker
   ↓
Service B
   ↓
Database
   ↓
External API

Each component can report a perfectly acceptable status while the complete business transaction fails.

Consider a simple checkout flow.

The following services may all be healthy:

  • Authentication: healthy
  • Product service: healthy
  • Payment service: healthy
  • Inventory service: healthy
  • Database: healthy

But suppose the inventory reservation is delayed by 8 seconds.

Technically, nothing may be “down”.

From the customer’s perspective, however:

checkout is broken.


Component Health Is Not System Health

A useful distinction is:

Component health

Is this individual technical component functioning within its expected parameters?

System health

Can the system successfully perform the business capabilities it is expected to provide?

These are related, but they are not equivalent.

A service can have:

CPU       → Normal
Memory    → Normal
Latency   → Normal
Errors    → Low
Availability → 99.99%

while the business workflow has:

Checkout success rate → 72%

The infrastructure dashboard may still look excellent.


The Missing Layer: Workflow Health

Modern observability needs another level above infrastructure and services.

Think of reliability as three layers:

                 BUSINESS
▲
│
Workflow Health
▲
│
Service Health
▲
│
Infrastructure Health

Infrastructure health answers:

Is the machine, container, database or network behaving correctly?

Service health answers:

Is this service behaving correctly?

Workflow health answers:

Can the user actually accomplish what they came to do?

The third question is often the most important.


Why Distributed Systems Make This Worse

Distributed systems introduce more ways for a workflow to degrade without producing an obvious failure.

Partial failure

One dependency becomes slow but does not completely fail.

Semantic failure

An API returns a technically valid response containing an unexpected business state.

Ordering failure

Events arrive in an order that the application did not anticipate.

Configuration failure

Every component is healthy but they are configured inconsistently.

Dependency degradation

An external service remains available but its latency or behavior changes.

Data inconsistency

Different components see different versions of the same business state.

None of these necessarily produces a classic:

HTTP 500

And that is exactly why they are difficult.


Stop Monitoring Services. Start Monitoring Capabilities.

This does not mean abandoning traditional monitoring.

It means adding another dimension.

Instead of monitoring only:

API latency
Database CPU
Pod status
Queue depth
HTTP errors

also monitor:

Login success rate
Order completion rate
Payment completion rate
Document generation success
Search success
User registration completion

These metrics describe capabilities, not infrastructure.


The Synthetic User as a Reliability Signal

One of the simplest ways to measure system health is to periodically execute critical workflows.

For example:

Open application
      ↓
Authenticate
      ↓
Create test transaction
      ↓
Validate result
      ↓
Clean up

This is more powerful than checking whether every service returns 200.

Why?

Because it tests the interaction between components.

A service-level health check might say:

Payment API: UP

A synthetic workflow might say:

Payment workflow: FAILED
Reason:
Authentication token accepted
Payment request accepted
Payment confirmation not received
Timeout: 8 seconds

The second signal is much closer to the user’s reality.


Business SLOs Should Complement Technical SLOs

Technical SLO:

99.9% of API requests complete within 500 ms.

Business SLO:

99.5% of checkout attempts successfully create an order.

These are not interchangeable.

A service can meet its technical SLO while the business SLO is failing.

That gap is where many modern incidents hide.


A Practical System Health Model

A useful model is to track four dimensions.

1. Availability

Can the capability be used?

Success / Attempt

2. Latency

How long does it take?

p50
p95
p99

3. Correctness

Did the system produce the expected result?

This is often harder than measuring HTTP status.

A request returning 200 can still produce the wrong business result.

4. Completeness

Did the entire workflow finish?

For example:

Order created
Payment confirmed
Inventory reserved
Confirmation sent

If only two of four steps complete, the system is not healthy from the user’s perspective.


The “200 OK” Trap

One of the most dangerous assumptions in modern monitoring is:

HTTP success means application success.

It does not.

Consider:

HTTP/1.1 200 OK

The response could still represent:

{
  "status": "pending",
  "result": null
}

Technically successful.

Operationally useless if the user expected a completed transaction.

This is why reliability engineering increasingly needs semantic validation, not just transport-level monitoring.


Designing Better Health Checks

A good health strategy should exist at multiple levels.

Level 1 — Infrastructure

CPU
Memory
Disk
Network
Node health

Level 2 — Service

Availability
Latency
Error rate
Dependency status

Level 3 — Transaction

Request accepted
Business operation completed
Expected state reached

Level 4 — User journey

User can complete the intended task

The higher the level, the closer the signal is to business reality.


What Should Trigger an Incident?

Not every technical degradation should become an incident.

The important question is:

Is a critical capability becoming unavailable, incorrect, or materially degraded?

For example:

Database CPU: 90%

might be concerning.

But if:

Checkout success: 99.8%

the business impact may currently be limited.

Conversely:

Database CPU: 40%

looks healthy.

But:

Checkout success: 65%

is an incident.

This changes how teams prioritize alerts.


From Alerting to Capability Monitoring

A mature observability architecture could therefore look like:

Infrastructure
     │
     ▼
Service Metrics
     │
     ▼
Distributed Traces
     │
     ▼
Business Transactions
     │
     ▼
User Journeys
     │
     ▼
Business SLOs

The objective is not to create more dashboards.

It is to create better signals.


A Practical Implementation Pattern

Start with your top five business-critical workflows.

For each workflow, define:

Workflow
├── Entry condition
├── Critical steps
├── Expected result
├── Maximum acceptable latency
├── Correctness criteria
├── Dependencies
└── Failure impact

Then connect those workflows to observability.

For example:

Checkout
 ├── Authentication
 ├── Cart validation
 ├── Inventory reservation
 ├── Payment
 ├── Order creation
 └── Confirmation

Now an incident can be described as:

Checkout success rate dropped from 99.4% to 91.2%, primarily due to payment confirmation latency.

That is much more actionable than:

Payment API is showing elevated latency.


The Bigger Shift

The industry has spent years improving infrastructure observability.

The next step is not simply collecting more telemetry.

It is understanding what that telemetry means to the system and its users.

The important question is moving from:

“Is the service healthy?”

to:

“Can the system still deliver what it is supposed to deliver?”

That distinction becomes critical as applications become more distributed, dynamic and dependent on external services.


Final Takeaway

A green dashboard can tell you that your components are alive.

It cannot necessarily tell you that your product works.

Modern reliability therefore needs to measure three different realities:

Is the infrastructure healthy?
          ↓
Are the services healthy?
          ↓
Can users successfully complete their workflows?

The last question is the one that matters most.

Because the most dangerous production incident may no longer be:

Everything is down.

It may be:

Everything is up — and the customer still can’t use it.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top