SLA vs SLI vs SLO vs Error Budget vs Alerting: A Complete SRE Guide

If you are learning DevOps, Site Reliability Engineering (SRE), Google Cloud, or cloud monitoring, you will frequently come across terms such as SLA, SLI, SLO, Error Budget, and Alerting.

These terms are closely related, but they have different meanings.

Understanding how they work together is essential for designing reliable production systems and building effective monitoring and alerting strategies.

In this guide, we’ll explain everything using simple examples.


What Are SLA, SLI, SLO, Error Budget, and Alerts?

The easiest way to understand these concepts is to think about a production application.

Imagine you operate an online shopping application.

Customers expect:

  • The website to be available.
  • Pages to load quickly.
  • Payments to work.
  • Orders to be processed successfully.
  • The application to recover quickly when something goes wrong.

To manage this reliability, SRE teams use several concepts:

SLI
↓
What are we measuring?

SLO
↓
What reliability target do we want?

SLA
↓
What do we promise our customers?

Error Budget
↓
How much failure can we tolerate?

Alert
↓
When does someone need to take action?

Let’s understand each one.


1. What Is an SLI?

SLI stands for Service Level Indicator.

An SLI is a measurement that tells us how well a service is performing from the user’s perspective.

In simple terms:

SLI tells us what actually happened.

For example, suppose your API receives 10,000 requests.

Out of those requests:

  • 9,950 succeed.
  • 50 fail.

Your availability SLI can be calculated as:

SLI = Good Events / Total Valid Events

Therefore:

SLI = 9,950 / 10,000

SLI = 99.5%

So your service had an availability SLI of 99.5%.


Common SLI Examples

Different aspects of a service can have different SLIs.

Reliability AspectExample SLI
AvailabilityPercentage of successful requests
LatencyPercentage of requests completed within 500 ms
CorrectnessPercentage of requests returning correct results
DurabilityPercentage of data successfully stored

For example:

Availability SLI:

Successful requests
────────────────────
Total valid requests

Or:

Latency SLI:

Requests completed within 500 ms
─────────────────────────────────
Total valid requests

2. Why Is an SLI Based on User Experience?

Not every monitoring metric is a good SLI.

For example, you might monitor:

CPU = 90%
Memory = 80%
Disk = 70%

These metrics are useful, but they don’t necessarily tell you whether users are experiencing a problem.

Your CPU could be at 90% while:

API requests → Successful
Response time → Fast
Users → Happy

Therefore, CPU utilization isn’t necessarily a good user-facing SLI.

Compare that with:

99.9% of API requests succeeded

That directly tells you something about the user’s experience.

A good SLI should have a strong relationship with what users actually experience.


3. What Is an SLO?

SLO stands for Service Level Objective.

An SLO combines an SLI with a target.

In simple terms:

SLI = What actually happened.

SLO = What we want to achieve.

For example:

SLI = 99.95% successful requests

SLO = 99.90% successful requests

The actual performance is above the target, so the service is meeting its SLO.


Example of an SLO

Suppose your company decides:

“At least 99.9% of valid API requests should succeed.”

Then:

SLO = 99.9% availability

Your monitoring system measures the actual availability.

Suppose it reports:

Actual SLI = 99.95%

Then:

99.95% > 99.9%

SLO → Achieved

But if the actual SLI becomes:

99.5%

then:

99.5% < 99.9%

SLO → Not achieved

4. What Does “Three Nines” Mean?

You will often hear SRE engineers talk about “nines.”

For example:

99%     → Two nines
99.9%   → Three nines
99.99%  → Four nines
99.999% → Five nines

For example:

99.9% availability is called three nines.

It might look like a tiny difference, but every additional nine significantly reduces the allowed downtime.

For approximately 30 days:

AvailabilityApproximate downtime
99%7 hours 12 minutes
99.9%43 minutes
99.99%4 minutes 19 seconds
99.999%26 seconds

This is why improving availability from 99.9% to 99.99% can require significant engineering effort and cost.


5. What Is an SLA?

SLA stands for Service Level Agreement.

An SLA is a formal agreement between a service provider and its customer.

It usually describes the level of service the customer can expect.

For example:

“The service provider guarantees 99.9% availability.”

An SLA may also include consequences if the provider fails to meet the agreed level, such as service credits.

The important distinction is:

SLI → Measurement

SLO → Internal reliability target

SLA → External/customer commitment

6. SLI vs SLO vs SLA

A simple example makes this much easier.

Imagine you operate a cloud API.

SLI

You measure:

99.95% successful requests

This is what actually happened.

SLO

Your engineering target is:

99.9% successful requests

This is what you want to achieve.

SLA

Your customer contract says:

99.9% availability

This is what you promise the customer.

Therefore:

SLI
"What happened?"

      ↓

SLO
"What do we want?"

      ↓

SLA
"What did we promise?"

7. What Is an Error Budget?

Now we get to one of the most important SRE concepts.

An Error Budget represents the amount of unreliability that your service can tolerate while still meeting its SLO.

The basic formula is:

Error Budget = 100% - SLO

For example:

SLO = 99.9%

Error Budget = 100% - 99.9%

Error Budget = 0.1%

So if your SLO is 99.9%, you have an error budget of 0.1%.


8. Error Budget Example

Suppose your application receives:

1,000,000 requests per month

Your SLO is:

99.9%

Therefore:

Error Budget = 0.1%

Calculate the allowed failures:

1,000,000 × 0.1%

= 1,000

So your application can have approximately:

1,000 failed requests

and still remain within the SLO.


9. Error Budget as Downtime

Error budgets can also be expressed as downtime.

Suppose:

SLO = 99.9%

Therefore:

Error Budget = 0.1%

For approximately 30 days:

30 days × 0.1%

≈ 43 minutes

So your error budget is approximately:

43 minutes of downtime per month.

If an incident causes 10 minutes of downtime:

43 minutes
-10 minutes
───────────
33 minutes remaining

You still have 33 minutes of error budget remaining.


10. Why Is an Error Budget Important?

Error budgets help balance:

Reliability vs Innovation

Without an error budget, teams can fall into two extremes.

Extreme 1: Deploy everything

Developers want to release features quickly.

More releases
      ↓
More changes
      ↓
Potentially more failures

Extreme 2: Never deploy

Operations teams may become extremely cautious.

Avoid changes
      ↓
Avoid risk
      ↓
Slow innovation

Error budgets provide a measurable compromise.

If the service is performing reliably and plenty of error budget remains:

The team can generally take more deployment risk.

If the error budget is almost exhausted:

The team should prioritize reliability and reduce risky changes.


11. What Is Error Budget Burn?

Error budget burn refers to how quickly your service is consuming its error budget.

Imagine:

Monthly error budget = 43 minutes

After 10 days:

30 minutes already consumed

That is a warning sign.

You haven’t necessarily violated the SLO yet, but you’re consuming the budget too quickly.

This leads to another important concept:

Burn rate

Burn rate tells you how quickly you’re consuming your available error budget.


12. What Is an Alert?

Now that we understand SLI, SLO, and error budgets, we can talk about alerting.

An alert is an automated notification that tells someone:

“Something needs attention.”

A typical flow looks like this:

Application
     ↓
Monitoring
     ↓
Metric / SLI
     ↓
Alert condition
     ↓
Notification
     ↓
Engineer

Notifications can be sent through channels such as:

  • Email
  • Slack
  • SMS
  • PagerDuty
  • Webhooks
  • Pub/Sub
  • Ticketing systems

13. The Goal of Alerting

goal of alerting is not to send as many notifications as possible.

The goal is:

A person is notified when they need to take action.

This is extremely important.

Suppose you have hundreds of alerts every day.

Eventually engineers may start ignoring them.

This is called:

Alert fatigue

A good alerting strategy should reduce unnecessary alerts while still detecting important incidents.


14. What Is a Time Series?

Monitoring systems usually collect data over time.

For example:

Time       Errors

10:00       2
10:01       3
10:02       5
10:03       10
10:04       50
10:05       100

This is called a time series.

A time series allows monitoring systems to analyze:

  • Trends
  • Error rates
  • Response times
  • Traffic
  • Resource usage
  • Availability

and determine whether an alert should be triggered.


15. What Is an Alert Window?

An alert window is the period over which the monitoring system evaluates a condition.

For example:

Last 5 minutes

or:

Last 60 minutes

Suppose your alert condition is:

“Alert if the error rate exceeds 1% during the last 10 minutes.”

Then the monitoring system continuously evaluates that 10-minute period.


16. Why Alert Based on Error Budget?

One of the best times to generate an alert is when the system is on track to consume its error budget too quickly.

Suppose:

SLO = 99.9%
Error Budget = 0.1%

You don’t want to wait until:

Error Budget = 0%

That’s too late.

Instead, you want to detect:

Current error rate
        ↓
Current burn rate
        ↓
Projected budget consumption
        ↓
Budget likely to be exhausted
        ↓
🚨 ALERT

The idea is similar to a financial budget.

If you have:

Monthly budget = ₹30,000

and you spend ₹20,000 during the first week, you don’t wait until you reach ₹30,000.

You realize:

“At this spending rate, I’m going to run out of money.”

SRE alerting works similarly.


17. Small Alert Windows vs Long Alert Windows

One of the important alerting decisions is choosing the right window.

Small Windows

Example:

5-minute window

Advantages:

  • Faster detection
  • Faster response
  • Shorter reset time

Disadvantages:

  • More false positives
  • Temporary spikes can trigger alerts

For example:

10:00 → Error rate 5%
10:01 → Error rate 0%
10:02 → Error rate 0%

A short window may detect this as a problem even though it was temporary.


18. Longer Alert Windows

Example:

60-minute window

Advantages:

  • Better signal
  • Fewer false positives
  • More stable alerts

Disadvantages:

  • Slower detection
  • A short incident might be missed

Therefore:

Small window
↓
Fast detection
↓
Potentially more false positives

Whereas:

Long window
↓
More stable detection
↓
Potentially slower response

19. Precision and Recall in Alerting

Two important concepts when evaluating alerting systems are:

Precision and Recall.

Precision

It asks:

When we generate an alert, how often is it a real problem?

For example:

10 alerts generated

9 were real problems
1 was a false alarm

That’s relatively high precision.


Recall

Recall asks:

Of all the real problems that occurred, how many did our alerting system detect?

For example:

10 real incidents occurred

Monitoring detected 9

Recall = high

If monitoring detected only 3:

Recall = low

20. Precision vs Recall Trade-Off

There is usually a trade-off.

If you make alerts extremely sensitive:

More alerts
    ↓
Detect more problems
    ↓
Higher recall
    ↓
But more false positives

If you make alerts less sensitive:

Fewer alerts
    ↓
Fewer false positives
    ↓
Higher precision
    ↓
But you may miss some problems

The goal is to find a practical balance.


21. Using Successive Failure Counts

One strategy is to require multiple failures before triggering an alert.

Instead of:

1 failed window → ALERT

you could use:

3 consecutive failed windows → ALERT

For example:

Window 1 → wrong
Window 2 → correct
Window 3 → wrong

No alert.

But:

Window 1 → wrong
Window 2 → wrong
Window 3 → wrong

🚨 Alert.

This helps avoid alerts caused by temporary anomalies.


22. The Problem With Successive Failures

There is a trade-off.

Suppose your system goes down for only 5 minutes.

If your alert requires three consecutive failed windows, you may detect the incident too late or miss it entirely.

Therefore:

Wait longer
    ↓
Better precision
    ↓
But slower detection

This is why alerting needs to be designed carefully.


23. Use Multiple Conditions

Instead of relying on only one condition, you can combine several signals.

For example:

Error rate > 5%

AND

Traffic > 100 requests/sec

AND

Condition persists for 5 minutes

Only when the conditions are satisfied:

🚨 ALERT

This can improve the quality of alerts.


24. Alerts Don’t Always Have to Go Directly to Humans

A sophisticated monitoring architecture might look like:

Monitoring
     ↓
Alert
     ↓
Pub/Sub
     ↓
Cloud Run
     ↓
Additional logic
     ↓
Decision
     ↓
Human notification

For example, monitoring detects:

Error rate = 5%

Instead of immediately waking an engineer, an automated system can check:

Is this production?
Is traffic high?
Is the problem lasting more than 5 minutes?
Is the error budget being consumed rapidly?
Are customers affected?

If the answer is yes:

🚨 PagerDuty
🚨 SMS
🚨 Slack

Otherwise, the event could simply be logged or sent to a ticket.


25. Prioritize Alerts Based on Customer Impact

Not every alert deserves the same priority.

Consider this:

CPU = 95%

That sounds serious.

But if:

API = healthy
Users = unaffected
Latency = normal

you might not need to wake someone immediately.

Now consider:

CPU = 50%

Checkout API = DOWN

This is much more important.

Why?

Because customers are directly affected.

Therefore:

Alert severity should be strongly influenced by customer impact.


26. Example Alert Severity Levels

A practical alerting system might use:

INFO

CPU = 70%

Action:

Log
Dashboard
Email

No immediate human response.


WARNING

Error rate increasing

Action:

Slack
Ticket
Email

Someone should investigate.


CRITICAL

Production API unavailable

Action:

PagerDuty
SMS
Phone
Slack

Immediate response may be required.


27. A Complete SRE Reliability Model

Now we can connect everything.

                    USER
                      │
                      ↓
                  APPLICATION
                      │
                      ↓
                     SLI
              "What happened?"
                      │
                      ↓
                     SLO
              "What do we want?"
                      │
                      ↓
                ERROR BUDGET
           "How much can we fail?"
                      │
                      ↓
                 BURN RATE
          "How fast are we failing?"
                      │
                      ↓
                    ALERT
           "Does someone need to act?"
                      │
             ┌────────┴────────┐
             ↓                 ↓
          Low impact       High impact
             ↓                 ↓
        Ticket/Email       PagerDuty/SMS

28. Real-World Example

Imagine an online banking API.

Your reliability requirements are:

SLI:
Percentage of successful API requests

SLO:
99.9% successful requests

Error Budget:
0.1%

Now monitoring detects:

Error rate increasing
        ↓
Error budget being consumed rapidly
        ↓
Projected budget exhaustion
        ↓
Alert triggered

The alert system checks:

Is this production?       YES
Is customer traffic high? YES
Is error rate increasing? YES
Is the problem persistent? YES

Then:

🚨 CRITICAL ALERT

Banking API error rate is increasing.
Error budget is being consumed rapidly.
Customer impact detected.

Action: Investigate API/database errors.

The engineer investigates the incident.

After the issue is fixed:

Errors return to normal
        ↓
Alert condition clears
        ↓
Alert resets

This is the complete reliability cycle.


29. The Complete Mental Model

If you’re preparing for a DevOps/SRE interview, remember this:

SLI
↓
A measurement of actual service reliability

SLO
↓
The reliability target

SLA
↓
The customer/business commitment

Error Budget
↓
The amount of unreliability we can tolerate

Burn Rate
↓
How quickly we're consuming that budget

Alert
↓
A notification when action is required

And the most important principle:

Don’t alert simply because a metric looks unusual. Alert when the condition indicates meaningful customer impact or a significant risk to the SLO, and someone needs to take action.


30. Interview Example

If an interviewer asks:

“How would you design an alerting strategy for a production service?”

A strong answer would be:

“I would start by defining user-focused SLIs such as availability and latency, then establish appropriate SLOs. From the SLO, I would calculate the error budget and use burn-rate-based alerting to detect when we’re consuming the budget too quickly. I would use appropriate alert windows and potentially multiple conditions to balance precision, recall, and detection time. Alerts would then be prioritized based on customer impact and severity, with critical issues routed to an on-call system such as PagerDuty and lower-severity issues going to Slack, email, or ticketing systems. The goal would be to ensure that humans are only paged when they need to take action.”


Conclusion

SRE is not about trying to make a service 100% perfect.

Instead, it is about defining a realistic reliability target and managing the trade-off between reliability, engineering effort, and product velocity.

The complete relationship is:

SLI
"What are we measuring?"

        ↓

SLO
"What reliability do we want?"

        ↓

SLA
"What have we promised customers?"

        ↓

Error Budget
"How much failure can we tolerate?"

        ↓

Burn Rate
"How quickly are we consuming that budget?"

        ↓

Alert
"Does someone need to take action?"

Once you understand this chain, concepts such as Google Cloud Monitoring, SRE alerting, burn-rate alerts, incident management, and production reliability become much easier to understand.

Quick Cheat Sheet

ConceptSimple Meaning
SLIWhat actually happened?
SLOWhat reliability do we want?
SLAWhat did we promise the customer?
Error BudgetHow much failure can we tolerate?
Burn RateHow quickly are we using the budget?
AlertWhen does someone need to take action?
PrecisionHow many alerts are real problems?
RecallHow many real problems did we detect?
Alert WindowHow long do we evaluate the condition?
SeverityHow urgently should we respond?

The golden rule:

A good alert is actionable, meaningful, and connected to customer impact.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *