What Is an SLA? (SLA vs SLO vs SLI Explained Simply)

If you run any kind of online service, you have probably encountered the term SLA. A Service Level Agreement (SLA) is a formal contract between a service provider and a customer that defines the expected level of service, including uptime guarantees, response times, and the consequences when those targets are missed.

But SLAs are just one piece of the reliability puzzle. Here is how the three key terms relate to each other:

  • SLI (Service Level Indicator) — The actual metric you measure. For example, "the percentage of HTTP requests that returned a 2xx response code in the last 30 days." This is your raw data.
  • SLO (Service Level Objective) — An internal target for your SLI. For example, "99.9% of requests should succeed each month." SLOs guide your engineering priorities but are not contractually binding.
  • SLA (Service Level Agreement) — The customer-facing promise, usually backed by financial penalties. For example, "We guarantee 99.9% uptime. If we miss it, you receive a 10% service credit."

Think of it this way: your SLI is what you measure, your SLO is what you aim for, and your SLA is what you promise. Smart teams set their internal SLO higher than their public SLA to create a safety buffer. If your SLA guarantees 99.9%, your SLO should target 99.95% so you have room to breathe before breaching the contract.

The Uptime Percentage Table: What the Nines Really Mean

People throw around "five nines" like it is a casual goal, but the differences between each level of uptime are dramatic. Here is exactly how much downtime each percentage allows:

Uptime %Downtime per MonthDowntime per Year
99%7 hours 18 minutes3.65 days
99.5%3 hours 39 minutes1.83 days
99.9%43 minutes 50 seconds8.76 hours
99.95%21 minutes 55 seconds4.38 hours
99.99%4 minutes 23 seconds52.6 minutes
99.999%26 seconds5.26 minutes

The jump from 99.9% to 99.99% means going from roughly 44 minutes of monthly downtime to just over 4 minutes. That single extra nine demands a fundamentally different approach to infrastructure, deployment practices, and incident response. Every additional nine roughly multiplies your operational complexity and cost by ten.

How to Calculate Your Actual Uptime

The formula for calculating uptime is straightforward:

Uptime % = ((Total Time - Downtime) / Total Time) x 100

For example, if your service experienced 45 minutes of downtime in a 30-day month (43,200 total minutes):

((43,200 - 45) / 43,200) x 100 = 99.896%

That is below 99.9%, which would breach a "three nines" SLA even though 45 minutes might feel insignificant. The details matter.

When calculating uptime, you need to decide on your measurement window:

  • Monthly calculation — The most common approach. Each calendar month is evaluated independently. A bad January does not affect February's numbers.
  • Annual calculation — Smooths out individual bad months but can hide recurring problems.
  • Rolling window — Measures the last 30 or 90 days continuously, regardless of calendar boundaries. This gives the most honest picture of current reliability and is what most modern monitoring tools report.

Most SLAs use monthly windows because they align naturally with billing cycles, but rolling windows are more useful for internal engineering decisions because they always reflect recent performance.

Scheduled Maintenance: Does It Count as Downtime?

This is one of the most debated questions in SLA management. The answer depends on your agreement, but here are the industry norms:

  • Enterprise SLAs typically exclude scheduled maintenance from uptime calculations, provided the customer receives advance notice (usually 48-72 hours) and the maintenance window falls during off-peak hours.
  • Modern SaaS companies increasingly do not exclude maintenance. If users cannot access your service, that is downtime regardless of whether it was planned. Zero-downtime deployments are now an expected standard.
  • Planned vs unplanned is the key distinction. Planned maintenance with proper notice is different from an unexpected outage. However, from the user's perspective, the result is the same: they cannot use your service.

The best practice is to invest in deployment strategies (blue-green deployments, rolling updates, feature flags) that eliminate the need for maintenance windows entirely. If you must have scheduled downtime, define exactly how it is handled in your SLA to avoid disputes.

Using Monitoring Data for SLA Reporting

You cannot manage what you do not measure. Accurate SLA reporting starts with reliable, continuous monitoring from multiple locations. Here is what you need:

  • Frequent checks — 30-second intervals give you granular data. A 5-minute interval means you could miss short outages entirely or overcount downtime for brief blips.
  • Multi-location monitoring — A single monitoring location might report downtime due to a regional network issue that does not affect your actual users. Cross-reference from multiple regions to confirm real outages.
  • Historical data retention — You need at least 12 months of data to generate annual SLA reports and identify long-term trends.
  • Automated reporting — Manual spreadsheets invite errors. Use a tool that calculates and exports uptime reports automatically.

GoPinger's SLO tracking lets you set your target uptime percentage and continuously measures your actual performance against it. When you need to generate an SLA report for a customer or stakeholder, the data is already there, accurate down to the second. Learn more about how this works on our SLA page.

Setting Realistic Uptime Targets by Business Type

Not every service needs 99.99% uptime. Chasing unnecessary nines wastes engineering time and money. Here are realistic targets based on what the service actually demands:

  • E-commerce (99.95%) — Every minute of downtime is lost revenue. Customers will buy from a competitor if your checkout is down. High uptime is directly tied to business outcomes.
  • SaaS platforms (99.9%) — Your users depend on your product for their daily work. A three-nines target balances reliability with the engineering cost of achieving it.
  • Marketing websites (99.5%) — A few hours of downtime per month is annoying but rarely catastrophic. Invest your reliability budget elsewhere.
  • Internal tools (99.5%) — Your team can tolerate brief outages and knows who to contact. Over-engineering internal tool uptime diverts resources from customer-facing services.
  • Financial services (99.99%) — Regulatory requirements and the high cost of transaction failures justify the investment in four-nines infrastructure.
  • Healthcare platforms (99.99%) — Patient safety and compliance requirements demand the highest levels of availability.

Choose a target that reflects the actual business impact of downtime, not an aspirational number that sounds good in a sales pitch. Then set your internal SLO one step above your public SLA to give yourself a buffer.

SLA Breach Response: What Happens When You Miss the Target

No matter how well you engineer your systems, SLA breaches will eventually happen. What matters is how you respond. A well-handled breach can actually strengthen customer trust.

Service Credits

Most SLAs include a service credit structure. A common model:

  • 99.0% - 99.9% — 10% credit on the affected month's invoice
  • 95.0% - 99.0% — 25% credit
  • Below 95.0% — 50% credit or the option to terminate the contract

Define these tiers clearly in your SLA before they are needed. Ambiguity during a breach leads to disputes and eroded trust.

Communication

When a breach occurs, proactive communication is essential. Acknowledge the issue publicly, explain the root cause in plain language, and detail the specific steps you are taking to prevent recurrence. Customers forgive outages far more easily than they forgive silence.

Post-Incident Review

Every SLA breach should trigger a blameless post-incident review within 48 hours. Document what happened, build a timeline, identify what worked and what did not, and assign concrete action items with owners and deadlines. This review becomes part of your SLA compliance record and demonstrates due diligence to customers.

Build Your SLA Dashboard

An SLA is only useful if both you and your customers can see the current status at any time. The most effective approach combines two layers:

  • Public status page — Shows real-time component status, uptime history (30/90 day), active incidents, and scheduled maintenance. This is your customer-facing transparency layer.
  • Internal SLA dashboard — Shows detailed metrics, SLO burn rate, error budgets, and trend analysis. This drives your engineering decisions.

With GoPinger, your status page updates automatically based on your monitoring data. There is no manual step between detecting an outage and reflecting it on your status page. Combine this with historical uptime charts and you have a complete SLA reporting pipeline without building anything custom.

An SLA is not just a number in a contract. It is a shared understanding between you and your customers about what they can expect. Measure it accurately, report it transparently, and treat every breach as an opportunity to improve.

Ready to start tracking your uptime with precision? GoPinger gives you 30-second monitoring on Pro, SLO tracking from Starter up, and automated status pages on every plan. Check our pricing to find the right fit, or see how we compare to Better Stack.