If you are learning DevOps, Site Reliability Engineering (SRE), Google Cloud, or cloud monitoring, you will frequently come across terms such as SLA, SLI, SLO, Error Budget, and Alerting.
These terms are closely related, but they have different meanings.
Understanding how they work together is essential for designing reliable production systems and building effective monitoring and alerting strategies.
In this guide, we’ll explain everything using simple examples.

What Are SLA, SLI, SLO, Error Budget, and Alerts?
The easiest way to understand these concepts is to think about a production application.
Imagine you operate an online shopping application.
Customers expect:
- The website to be available.
- Pages to load quickly.
- Payments to work.
- Orders to be processed successfully.
- The application to recover quickly when something goes wrong.
To manage this reliability, SRE teams use several concepts:
SLI
↓
What are we measuring?
SLO
↓
What reliability target do we want?
SLA
↓
What do we promise our customers?
Error Budget
↓
How much failure can we tolerate?
Alert
↓
When does someone need to take action?
Let’s understand each one.
1. What Is an SLI?
SLI stands for Service Level Indicator.
An SLI is a measurement that tells us how well a service is performing from the user’s perspective.
In simple terms:
SLI tells us what actually happened.
For example, suppose your API receives 10,000 requests.
Out of those requests:
- 9,950 succeed.
- 50 fail.
Your availability SLI can be calculated as:
SLI = Good Events / Total Valid Events
Therefore:
SLI = 9,950 / 10,000
SLI = 99.5%
So your service had an availability SLI of 99.5%.
Common SLI Examples
Different aspects of a service can have different SLIs.
| Reliability Aspect | Example SLI |
|---|---|
| Availability | Percentage of successful requests |
| Latency | Percentage of requests completed within 500 ms |
| Correctness | Percentage of requests returning correct results |
| Durability | Percentage of data successfully stored |
For example:
Availability SLI:
Successful requests
────────────────────
Total valid requests
Or:
Latency SLI:
Requests completed within 500 ms
─────────────────────────────────
Total valid requests
2. Why Is an SLI Based on User Experience?
Not every monitoring metric is a good SLI.
For example, you might monitor:
CPU = 90%
Memory = 80%
Disk = 70%
These metrics are useful, but they don’t necessarily tell you whether users are experiencing a problem.
Your CPU could be at 90% while:
API requests → Successful
Response time → Fast
Users → Happy
Therefore, CPU utilization isn’t necessarily a good user-facing SLI.
Compare that with:
99.9% of API requests succeeded
That directly tells you something about the user’s experience.
A good SLI should have a strong relationship with what users actually experience.
3. What Is an SLO?
SLO stands for Service Level Objective.
An SLO combines an SLI with a target.
In simple terms:
SLI = What actually happened.
SLO = What we want to achieve.
For example:
SLI = 99.95% successful requests
SLO = 99.90% successful requests
The actual performance is above the target, so the service is meeting its SLO.
Example of an SLO
Suppose your company decides:
“At least 99.9% of valid API requests should succeed.”
Then:
SLO = 99.9% availability
Your monitoring system measures the actual availability.
Suppose it reports:
Actual SLI = 99.95%
Then:
99.95% > 99.9%
SLO → Achieved
But if the actual SLI becomes:
99.5%
then:
99.5% < 99.9%
SLO → Not achieved
4. What Does “Three Nines” Mean?
You will often hear SRE engineers talk about “nines.”
For example:
99% → Two nines
99.9% → Three nines
99.99% → Four nines
99.999% → Five nines
For example:
99.9% availability is called three nines.
It might look like a tiny difference, but every additional nine significantly reduces the allowed downtime.
For approximately 30 days:
| Availability | Approximate downtime |
|---|---|
| 99% | 7 hours 12 minutes |
| 99.9% | 43 minutes |
| 99.99% | 4 minutes 19 seconds |
| 99.999% | 26 seconds |
This is why improving availability from 99.9% to 99.99% can require significant engineering effort and cost.
5. What Is an SLA?
SLA stands for Service Level Agreement.
An SLA is a formal agreement between a service provider and its customer.
It usually describes the level of service the customer can expect.
For example:
“The service provider guarantees 99.9% availability.”
An SLA may also include consequences if the provider fails to meet the agreed level, such as service credits.
The important distinction is:
SLI → Measurement
SLO → Internal reliability target
SLA → External/customer commitment
6. SLI vs SLO vs SLA
A simple example makes this much easier.
Imagine you operate a cloud API.
SLI
You measure:
99.95% successful requests
This is what actually happened.
SLO
Your engineering target is:
99.9% successful requests
This is what you want to achieve.
SLA
Your customer contract says:
99.9% availability
This is what you promise the customer.
Therefore:
SLI
"What happened?"
↓
SLO
"What do we want?"
↓
SLA
"What did we promise?"
7. What Is an Error Budget?
Now we get to one of the most important SRE concepts.
An Error Budget represents the amount of unreliability that your service can tolerate while still meeting its SLO.
The basic formula is:
Error Budget = 100% - SLO
For example:
SLO = 99.9%
Error Budget = 100% - 99.9%
Error Budget = 0.1%
So if your SLO is 99.9%, you have an error budget of 0.1%.
8. Error Budget Example
Suppose your application receives:
1,000,000 requests per month
Your SLO is:
99.9%
Therefore:
Error Budget = 0.1%
Calculate the allowed failures:
1,000,000 × 0.1%
= 1,000
So your application can have approximately:
1,000 failed requests
and still remain within the SLO.
9. Error Budget as Downtime
Error budgets can also be expressed as downtime.
Suppose:
SLO = 99.9%
Therefore:
Error Budget = 0.1%
For approximately 30 days:
30 days × 0.1%
≈ 43 minutes
So your error budget is approximately:
43 minutes of downtime per month.
If an incident causes 10 minutes of downtime:
43 minutes
-10 minutes
───────────
33 minutes remaining
You still have 33 minutes of error budget remaining.
10. Why Is an Error Budget Important?
Error budgets help balance:
Reliability vs Innovation
Without an error budget, teams can fall into two extremes.
Extreme 1: Deploy everything
Developers want to release features quickly.
More releases
↓
More changes
↓
Potentially more failures
Extreme 2: Never deploy
Operations teams may become extremely cautious.
Avoid changes
↓
Avoid risk
↓
Slow innovation
Error budgets provide a measurable compromise.
If the service is performing reliably and plenty of error budget remains:
The team can generally take more deployment risk.
If the error budget is almost exhausted:
The team should prioritize reliability and reduce risky changes.
11. What Is Error Budget Burn?
Error budget burn refers to how quickly your service is consuming its error budget.
Imagine:
Monthly error budget = 43 minutes
After 10 days:
30 minutes already consumed
That is a warning sign.
You haven’t necessarily violated the SLO yet, but you’re consuming the budget too quickly.
This leads to another important concept:
Burn rate
Burn rate tells you how quickly you’re consuming your available error budget.
12. What Is an Alert?
Now that we understand SLI, SLO, and error budgets, we can talk about alerting.
An alert is an automated notification that tells someone:
“Something needs attention.”
A typical flow looks like this:
Application
↓
Monitoring
↓
Metric / SLI
↓
Alert condition
↓
Notification
↓
Engineer
Notifications can be sent through channels such as:
- Slack
- SMS
- PagerDuty
- Webhooks
- Pub/Sub
- Ticketing systems
13. The Goal of Alerting
goal of alerting is not to send as many notifications as possible.
The goal is:
A person is notified when they need to take action.
This is extremely important.
Suppose you have hundreds of alerts every day.
Eventually engineers may start ignoring them.
This is called:
Alert fatigue
A good alerting strategy should reduce unnecessary alerts while still detecting important incidents.
14. What Is a Time Series?
Monitoring systems usually collect data over time.
For example:
Time Errors
10:00 2
10:01 3
10:02 5
10:03 10
10:04 50
10:05 100
This is called a time series.
A time series allows monitoring systems to analyze:
- Trends
- Error rates
- Response times
- Traffic
- Resource usage
- Availability
and determine whether an alert should be triggered.
15. What Is an Alert Window?
An alert window is the period over which the monitoring system evaluates a condition.
For example:
Last 5 minutes
or:
Last 60 minutes
Suppose your alert condition is:
“Alert if the error rate exceeds 1% during the last 10 minutes.”
Then the monitoring system continuously evaluates that 10-minute period.
16. Why Alert Based on Error Budget?
One of the best times to generate an alert is when the system is on track to consume its error budget too quickly.
Suppose:
SLO = 99.9%
Error Budget = 0.1%
You don’t want to wait until:
Error Budget = 0%
That’s too late.
Instead, you want to detect:
Current error rate
↓
Current burn rate
↓
Projected budget consumption
↓
Budget likely to be exhausted
↓
🚨 ALERT
The idea is similar to a financial budget.
If you have:
Monthly budget = ₹30,000
and you spend ₹20,000 during the first week, you don’t wait until you reach ₹30,000.
You realize:
“At this spending rate, I’m going to run out of money.”
SRE alerting works similarly.
17. Small Alert Windows vs Long Alert Windows
One of the important alerting decisions is choosing the right window.
Small Windows
Example:
5-minute window
Advantages:
- Faster detection
- Faster response
- Shorter reset time
Disadvantages:
- More false positives
- Temporary spikes can trigger alerts
For example:
10:00 → Error rate 5%
10:01 → Error rate 0%
10:02 → Error rate 0%
A short window may detect this as a problem even though it was temporary.
18. Longer Alert Windows
Example:
60-minute window
Advantages:
- Better signal
- Fewer false positives
- More stable alerts
Disadvantages:
- Slower detection
- A short incident might be missed
Therefore:
Small window
↓
Fast detection
↓
Potentially more false positives
Whereas:
Long window
↓
More stable detection
↓
Potentially slower response
19. Precision and Recall in Alerting
Two important concepts when evaluating alerting systems are:
Precision and Recall.
Precision
It asks:
When we generate an alert, how often is it a real problem?
For example:
10 alerts generated
9 were real problems
1 was a false alarm
That’s relatively high precision.
Recall
Recall asks:
Of all the real problems that occurred, how many did our alerting system detect?
For example:
10 real incidents occurred
Monitoring detected 9
Recall = high
If monitoring detected only 3:
Recall = low
20. Precision vs Recall Trade-Off
There is usually a trade-off.
If you make alerts extremely sensitive:
More alerts
↓
Detect more problems
↓
Higher recall
↓
But more false positives
If you make alerts less sensitive:
Fewer alerts
↓
Fewer false positives
↓
Higher precision
↓
But you may miss some problems
The goal is to find a practical balance.
21. Using Successive Failure Counts
One strategy is to require multiple failures before triggering an alert.
Instead of:
1 failed window → ALERT
you could use:
3 consecutive failed windows → ALERT
For example:
Window 1 → wrong
Window 2 → correct
Window 3 → wrong
No alert.
But:
Window 1 → wrong
Window 2 → wrong
Window 3 → wrong
🚨 Alert.
This helps avoid alerts caused by temporary anomalies.
22. The Problem With Successive Failures
There is a trade-off.
Suppose your system goes down for only 5 minutes.
If your alert requires three consecutive failed windows, you may detect the incident too late or miss it entirely.
Therefore:
Wait longer
↓
Better precision
↓
But slower detection
This is why alerting needs to be designed carefully.
23. Use Multiple Conditions
Instead of relying on only one condition, you can combine several signals.
For example:
Error rate > 5%
AND
Traffic > 100 requests/sec
AND
Condition persists for 5 minutes
Only when the conditions are satisfied:
🚨 ALERT
This can improve the quality of alerts.
24. Alerts Don’t Always Have to Go Directly to Humans
A sophisticated monitoring architecture might look like:
Monitoring
↓
Alert
↓
Pub/Sub
↓
Cloud Run
↓
Additional logic
↓
Decision
↓
Human notification
For example, monitoring detects:
Error rate = 5%
Instead of immediately waking an engineer, an automated system can check:
Is this production?
Is traffic high?
Is the problem lasting more than 5 minutes?
Is the error budget being consumed rapidly?
Are customers affected?
If the answer is yes:
🚨 PagerDuty
🚨 SMS
🚨 Slack
Otherwise, the event could simply be logged or sent to a ticket.
25. Prioritize Alerts Based on Customer Impact
Not every alert deserves the same priority.
Consider this:
CPU = 95%
That sounds serious.
But if:
API = healthy
Users = unaffected
Latency = normal
you might not need to wake someone immediately.
Now consider:
CPU = 50%
Checkout API = DOWN
This is much more important.
Why?
Because customers are directly affected.
Therefore:
Alert severity should be strongly influenced by customer impact.
26. Example Alert Severity Levels
A practical alerting system might use:
INFO
CPU = 70%
Action:
Log
Dashboard
Email
No immediate human response.
WARNING
Error rate increasing
Action:
Slack
Ticket
Email
Someone should investigate.
CRITICAL
Production API unavailable
Action:
PagerDuty
SMS
Phone
Slack
Immediate response may be required.
27. A Complete SRE Reliability Model
Now we can connect everything.
USER
│
↓
APPLICATION
│
↓
SLI
"What happened?"
│
↓
SLO
"What do we want?"
│
↓
ERROR BUDGET
"How much can we fail?"
│
↓
BURN RATE
"How fast are we failing?"
│
↓
ALERT
"Does someone need to act?"
│
┌────────┴────────┐
↓ ↓
Low impact High impact
↓ ↓
Ticket/Email PagerDuty/SMS
28. Real-World Example
Imagine an online banking API.
Your reliability requirements are:
SLI:
Percentage of successful API requests
SLO:
99.9% successful requests
Error Budget:
0.1%
Now monitoring detects:
Error rate increasing
↓
Error budget being consumed rapidly
↓
Projected budget exhaustion
↓
Alert triggered
The alert system checks:
Is this production? YES
Is customer traffic high? YES
Is error rate increasing? YES
Is the problem persistent? YES
Then:
🚨 CRITICAL ALERT
Banking API error rate is increasing.
Error budget is being consumed rapidly.
Customer impact detected.
Action: Investigate API/database errors.
The engineer investigates the incident.
After the issue is fixed:
Errors return to normal
↓
Alert condition clears
↓
Alert resets
This is the complete reliability cycle.
29. The Complete Mental Model
If you’re preparing for a DevOps/SRE interview, remember this:
SLI
↓
A measurement of actual service reliability
SLO
↓
The reliability target
SLA
↓
The customer/business commitment
Error Budget
↓
The amount of unreliability we can tolerate
Burn Rate
↓
How quickly we're consuming that budget
Alert
↓
A notification when action is required
And the most important principle:
Don’t alert simply because a metric looks unusual. Alert when the condition indicates meaningful customer impact or a significant risk to the SLO, and someone needs to take action.
30. Interview Example
If an interviewer asks:
“How would you design an alerting strategy for a production service?”
A strong answer would be:
“I would start by defining user-focused SLIs such as availability and latency, then establish appropriate SLOs. From the SLO, I would calculate the error budget and use burn-rate-based alerting to detect when we’re consuming the budget too quickly. I would use appropriate alert windows and potentially multiple conditions to balance precision, recall, and detection time. Alerts would then be prioritized based on customer impact and severity, with critical issues routed to an on-call system such as PagerDuty and lower-severity issues going to Slack, email, or ticketing systems. The goal would be to ensure that humans are only paged when they need to take action.”
Conclusion
SRE is not about trying to make a service 100% perfect.
Instead, it is about defining a realistic reliability target and managing the trade-off between reliability, engineering effort, and product velocity.
The complete relationship is:
SLI
"What are we measuring?"
↓
SLO
"What reliability do we want?"
↓
SLA
"What have we promised customers?"
↓
Error Budget
"How much failure can we tolerate?"
↓
Burn Rate
"How quickly are we consuming that budget?"
↓
Alert
"Does someone need to take action?"
Once you understand this chain, concepts such as Google Cloud Monitoring, SRE alerting, burn-rate alerts, incident management, and production reliability become much easier to understand.
Quick Cheat Sheet
| Concept | Simple Meaning |
|---|---|
| SLI | What actually happened? |
| SLO | What reliability do we want? |
| SLA | What did we promise the customer? |
| Error Budget | How much failure can we tolerate? |
| Burn Rate | How quickly are we using the budget? |
| Alert | When does someone need to take action? |
| Precision | How many alerts are real problems? |
| Recall | How many real problems did we detect? |
| Alert Window | How long do we evaluate the condition? |
| Severity | How urgently should we respond? |
The golden rule:
A good alert is actionable, meaningful, and connected to customer impact.
Leave a Reply