Why healthy components can still produce an unhealthy system

There is a particularly frustrating kind of production incident.
The API is returning 200.
The Kubernetes pods are healthy.
CPU and memory are normal.
The database is responding.
The message broker has no significant backlog.
The monitoring dashboard is mostly green.
And yet, the user cannot complete the most important action in the application.
Nothing appears to be broken.
Everything is working.
Except the system.
This is becoming one of the most important reliability problems in modern distributed software: the difference between component health and system health.
The Green Dashboard Illusion
Traditional monitoring often starts with a simple assumption:
If the important components are healthy, the application should be healthy.
That assumption worked reasonably well when applications were relatively self-contained.
Modern systems are different.
A single user workflow can cross:
Browser
↓
CDN
↓
API Gateway
↓
Authentication
↓
Service A
↓
Message Broker
↓
Service B
↓
Database
↓
External API
Each component can report a perfectly acceptable status while the complete business transaction fails.
Consider a simple checkout flow.
The following services may all be healthy:
- Authentication: healthy
- Product service: healthy
- Payment service: healthy
- Inventory service: healthy
- Database: healthy
But suppose the inventory reservation is delayed by 8 seconds.
Technically, nothing may be “down”.
From the customer’s perspective, however:
checkout is broken.
Component Health Is Not System Health
A useful distinction is:
Component health
Is this individual technical component functioning within its expected parameters?
System health
Can the system successfully perform the business capabilities it is expected to provide?
These are related, but they are not equivalent.
A service can have:
CPU → Normal
Memory → Normal
Latency → Normal
Errors → Low
Availability → 99.99%
while the business workflow has:
Checkout success rate → 72%
The infrastructure dashboard may still look excellent.
The Missing Layer: Workflow Health
Modern observability needs another level above infrastructure and services.
Think of reliability as three layers:
BUSINESS
▲
│
Workflow Health
▲
│
Service Health
▲
│
Infrastructure Health
Infrastructure health answers:
Is the machine, container, database or network behaving correctly?
Service health answers:
Is this service behaving correctly?
Workflow health answers:
Can the user actually accomplish what they came to do?
The third question is often the most important.
Why Distributed Systems Make This Worse
Distributed systems introduce more ways for a workflow to degrade without producing an obvious failure.
Partial failure
One dependency becomes slow but does not completely fail.
Semantic failure
An API returns a technically valid response containing an unexpected business state.
Ordering failure
Events arrive in an order that the application did not anticipate.
Configuration failure
Every component is healthy but they are configured inconsistently.
Dependency degradation
An external service remains available but its latency or behavior changes.
Data inconsistency
Different components see different versions of the same business state.
None of these necessarily produces a classic:
HTTP 500
And that is exactly why they are difficult.
Stop Monitoring Services. Start Monitoring Capabilities.
This does not mean abandoning traditional monitoring.
It means adding another dimension.
Instead of monitoring only:
API latency
Database CPU
Pod status
Queue depth
HTTP errors
also monitor:
Login success rate
Order completion rate
Payment completion rate
Document generation success
Search success
User registration completion
These metrics describe capabilities, not infrastructure.
The Synthetic User as a Reliability Signal
One of the simplest ways to measure system health is to periodically execute critical workflows.
For example:
Open application
↓
Authenticate
↓
Create test transaction
↓
Validate result
↓
Clean up
This is more powerful than checking whether every service returns 200.
Why?
Because it tests the interaction between components.
A service-level health check might say:
Payment API: UP
A synthetic workflow might say:
Payment workflow: FAILED
Reason:
Authentication token accepted
Payment request accepted
Payment confirmation not received
Timeout: 8 seconds
The second signal is much closer to the user’s reality.
Business SLOs Should Complement Technical SLOs
Technical SLO:
99.9% of API requests complete within 500 ms.
Business SLO:
99.5% of checkout attempts successfully create an order.
These are not interchangeable.
A service can meet its technical SLO while the business SLO is failing.
That gap is where many modern incidents hide.
A Practical System Health Model
A useful model is to track four dimensions.
1. Availability
Can the capability be used?
Success / Attempt
2. Latency
How long does it take?
p50
p95
p99
3. Correctness
Did the system produce the expected result?
This is often harder than measuring HTTP status.
A request returning 200 can still produce the wrong business result.
4. Completeness
Did the entire workflow finish?
For example:
Order created
Payment confirmed
Inventory reserved
Confirmation sent
If only two of four steps complete, the system is not healthy from the user’s perspective.
The “200 OK” Trap
One of the most dangerous assumptions in modern monitoring is:
HTTP success means application success.
It does not.
Consider:
HTTP/1.1 200 OK
The response could still represent:
{
"status": "pending",
"result": null
}
Technically successful.
Operationally useless if the user expected a completed transaction.
This is why reliability engineering increasingly needs semantic validation, not just transport-level monitoring.
Designing Better Health Checks
A good health strategy should exist at multiple levels.
Level 1 — Infrastructure
CPU
Memory
Disk
Network
Node health
Level 2 — Service
Availability
Latency
Error rate
Dependency status
Level 3 — Transaction
Request accepted
Business operation completed
Expected state reached
Level 4 — User journey
User can complete the intended task
The higher the level, the closer the signal is to business reality.
What Should Trigger an Incident?
Not every technical degradation should become an incident.
The important question is:
Is a critical capability becoming unavailable, incorrect, or materially degraded?
For example:
Database CPU: 90%
might be concerning.
But if:
Checkout success: 99.8%
the business impact may currently be limited.
Conversely:
Database CPU: 40%
looks healthy.
But:
Checkout success: 65%
is an incident.
This changes how teams prioritize alerts.
From Alerting to Capability Monitoring
A mature observability architecture could therefore look like:
Infrastructure
│
▼
Service Metrics
│
▼
Distributed Traces
│
▼
Business Transactions
│
▼
User Journeys
│
▼
Business SLOs
The objective is not to create more dashboards.
It is to create better signals.
A Practical Implementation Pattern
Start with your top five business-critical workflows.
For each workflow, define:
Workflow
├── Entry condition
├── Critical steps
├── Expected result
├── Maximum acceptable latency
├── Correctness criteria
├── Dependencies
└── Failure impact
Then connect those workflows to observability.
For example:
Checkout
├── Authentication
├── Cart validation
├── Inventory reservation
├── Payment
├── Order creation
└── Confirmation
Now an incident can be described as:
Checkout success rate dropped from 99.4% to 91.2%, primarily due to payment confirmation latency.
That is much more actionable than:
Payment API is showing elevated latency.
The Bigger Shift
The industry has spent years improving infrastructure observability.
The next step is not simply collecting more telemetry.
It is understanding what that telemetry means to the system and its users.
The important question is moving from:
“Is the service healthy?”
to:
“Can the system still deliver what it is supposed to deliver?”
That distinction becomes critical as applications become more distributed, dynamic and dependent on external services.
Final Takeaway
A green dashboard can tell you that your components are alive.
It cannot necessarily tell you that your product works.
Modern reliability therefore needs to measure three different realities:
Is the infrastructure healthy?
↓
Are the services healthy?
↓
Can users successfully complete their workflows?
The last question is the one that matters most.
Because the most dangerous production incident may no longer be:
Everything is down.
It may be:
Everything is up — and the customer still can’t use it.
