What is a service level agreement (SLA)?
A service level agreement (SLA) is a formal commitment between a service provider and its customer that defines the level of service to be delivered, how it will be measured, and what happens when targets are missed. SLAs translate vague expectations, such as “high availability,” into specific, contractually binding numbers, such as uptime percentages, response times, and resolution targets.
Also called a service-level agreement or service-level contract.
Why SLAs matter
SLAs are how the business holds IT and its vendors accountable for service quality. They turn abstract reliability promises into measurable obligations. When a payments platform commits to 99.99% availability, that is not aspirational language; it is a contractual ceiling on the amount of downtime acceptable in a year.
The financial stakes can be material. SLA breaches trigger service credits, fee reductions, or renegotiation rights. In regulated industries, repeated breaches can also feed into audit findings and customer churn analyses. Internally, SLA attainment is one of the most common metrics that IT leaders report to executives and boards.
SLAs also drive operational behavior. They shape on-call rotations, alerting thresholds, incident severity definitions, and how aggressively teams pursue root cause analysis. An organization that does not know its SLAs is, by definition, operating without a clear definition of “good enough”.
Key components of an SLA
- Service scope: The specific services covered, including any exclusions for planned maintenance or third-party dependencies.
- Service-level objectives (SLOs): The measurable targets, such as uptime percentage, response time, or resolution time, that define acceptable performance.
- Service-level indicators (SLIs): The actual metrics collected, such as the successful HTTP request ratio or the mean time to acknowledge.
- Measurement window: The period over which performance is averaged, typically monthly or quarterly.
- Remedies and penalties: Service credits, fee adjustments, or termination rights triggered by a breach.
- Reporting and review: How performance is reported, how often it’s reviewed, and how disputes are handled.
SLA vs. OLA vs. UC vs. KPI
SLA is one of several related but distinct constructs. Each governs a different relationship and serves a different purpose.
| Term | Between | Purpose |
|---|---|---|
| SLA (Service Level Agreement) | Provider and customer | External commitment with contractual remedies |
| OLA (Operational Level Agreement) | Internal teams within the provider | Aligns internal teams to meet customer-facing SLAs |
| UC (Underpinning Contract) | Provider and third-party supplier | Ensures vendor performance supports the SLA |
| KPI (Key Performance Indicator) | Internal | Tracks performance against goals; not contractually binding |
Uptime tiers and what they mean
Most availability SLAs are expressed as a percentage of uptime over a year. Each additional “nine” sharply compresses the amount of allowable downtime.
| Uptime | Allowed downtime per year | Allowed downtime per month |
|---|---|---|
| 99% | 3 days 15 hours | 7 hours 18 min |
| 99.9% | 8 hours 46 min | 43.8 min |
| 99.95% | 4 hours 23 min | 21.9 min |
| 99.99% | 52.6 min | 4.38 min |
| 99.999% | 5.26 min | 26.3 sec |
The cost of an SLA breach
Breach costs come in three forms. The first is direct: service credits, refunds, or contractual penalties tied to the agreement itself. The second is reputational: customer churn, executive scrutiny, and pipeline impact when prospects discover the breach. The third is operational: incident response, post-mortems, remediation projects, and follow-on changes that consume engineering capacity for weeks after the event.
Sustained SLA performance depends less on heroics during incidents and more on the upstream investments that keep them rare and short: noise reduction in monitoring, fast incident correlation, well-rehearsed runbooks, and disciplined change risk control.
SLA use cases in IT operations
- Hosting and SaaS contracts: Defining availability and support response commitments to enterprise customers.
- Internal IT services: Setting expectations between IT and business units for shared services such as email, identity, or VPN.
- MSP and outsourcing agreements: Governing performance from managed service providers and outsourced NOC operations.
- Incident response targets: Driving severity-based response and resolution targets that flow into on-call rotations and escalation paths.
- AIOps-supported SLA attainment: Using event correlation and incident automation to reduce MTTD and MTTR, which directly improves SLA compliance.
Frequently asked questions about service level agreement (SLA)
What is the difference between an SLA and an SLO?
An SLA is the external, contractual commitment to a customer, including remedies for missed targets. An SLO is the specific performance objective being measured, such as 99.95% availability. SLAs typically reference one or more SLOs and add commercial terms around them.
What is the difference between SLA and OLA?
An SLA is between a service provider and a customer. An OLA is an internal agreement between teams inside the provider that ensures the customer-facing SLA can be met. If a customer SLA promises a four-hour response, an OLA might require the database team to acknowledge escalations within 30 minutes.
How is SLA uptime calculated?
Uptime is the percentage of time a service is available within a measurement window, usually a month or a quarter. It is calculated as total time minus downtime, divided by total time. Most SLAs exclude scheduled maintenance and certain force-majeure events from the calculation.
What happens when an SLA is breached?
The remedies are defined in the agreement itself, typically as service credits applied to the next invoice or fee reductions. Serious or repeated breaches can also trigger termination rights, escalation to executive review, or renegotiation. Internally, a breach usually triggers a post-incident review.
What is a realistic SLA target?
It depends on the service and the cost of meeting it. Each additional “nine” of availability roughly multiplies the engineering and infrastructure investment required. 99.9% is typical for SaaS; 99.99% is typical for mission-critical platforms; 99.999% is rare and expensive, reserved for systems like core telecom and payments switching.
How does AIOps help with SLA compliance?
AIOps reduces alert noise, correlates events into incidents, and accelerates triage, which lowers MTTD and MTTR. Because SLA performance is sensitive to both detection and resolution speed, AIOps tends to translate directly into improved SLA attainment for incident-driven services.
See also
- ITSM
- ITIL
- MTTR
- MTTD
- IT incident management
- Error Budget
Check out more related content
PLATFORM
BigPanda Agentic ITOps
See how BigPanda uses agentic AI in IT operations.