Site Monitoring

Uptime SLAs: What to Include in Client Maintenance Contracts

The agency gets a call on a Saturday morning. The client’s e-commerce site has been down for three hours. Lost revenue, angry customers, a panicked owner demanding to know what you’re doing about it — and what you’re going to do to make it right. In that moment, what does your maintenance contract actually say?

For most agencies, the honest answer is: not enough. Maintenance retainers tend to grow organically, starting as an informal “we’ll keep an eye on things” arrangement and evolving into a significant monthly fee — often £200–£600/mo — with contract language that was drafted once and never updated. The result is a relationship governed by assumptions rather than agreements, and assumptions fall apart the moment something goes wrong.

This guide covers exactly what a well-drafted uptime SLA should include, which commitments you should never make, and how to structure your monitoring so you can actually stand behind what you’ve signed. It is deliberately practical — aimed at agency owners who need to update contracts and set up better processes this quarter, not in the abstract future.

The Problem With Vague Uptime Commitments

Most agency maintenance contracts contain language along the lines of “we will monitor your website and respond to issues in a timely manner.” Every word in that sentence is ambiguous. What counts as monitoring — manual checks once a week, or automated checks every minute? What counts as an issue — full downtime only, or slow load times, expired SSL certificates, failed contact forms? What is timely — four hours, or four minutes?

That vagueness feels safe when you’re drafting the contract. It gives you flexibility. In practice, it means every incident becomes a negotiation about what was actually promised, and that negotiation takes place at the worst possible time — when the client is already upset. Without a defined standard, the client’s expectation (immediate response, full resolution within the hour) will almost always exceed your operational capability (check email within four hours, resolve within a working day).

The fix is specificity. A well-written SLA doesn’t expose you to more risk — it reduces it, because it replaces the client’s imagined standard (perfect) with an agreed one (realistic). That shift in framing is worth more than any dispute clause you could add.

What Your SLA Must Define

Before you can promise a response time or a resolution target, you need to define the terms that those promises depend on. The following are the non-negotiable definitions every maintenance contract should include.

Uptime measurement methodology

Uptime should be defined as a percentage of total minutes in a calendar month, calculated from automated monitoring data. Specify the monitoring interval (every 1 minute or every 5 minutes are both common), the monitoring provider or tool, and what constitutes a confirmed outage versus a transient failure. A single failed check in an otherwise-operational month is not an outage — most tools require two or three consecutive failed checks before raising an alert, and your contract should mirror that logic. Without this definition, a client can point to a single monitoring tool screenshot and dispute your figures.

The monitoring scope

List exactly what is being monitored. At minimum this should cover HTTP availability (the site loads a valid 200 response), SSL certificate validity (with advance notice of expiration — 30 days is standard), and domain expiration where relevant. If you include WordPress plugin updates, define that too: does monitoring mean checking for available updates, or does it mean applying them? These are very different things. Whatever is not on the scope list is not included in the SLA, and that needs to be explicit.

Incident severity tiers

A site that is completely inaccessible is not the same as a site where one image is broken. Your SLA should define at least two severity tiers — Critical and Standard are sufficient — with separate response and resolution targets for each. Critical might mean: complete downtime, SSL certificate expired, or payment gateway non-functional. Standard might mean: single page error, minor display issue, plugin update available. The distinction matters because it sets appropriate expectations and lets you prioritise correctly when you are managing incidents across multiple clients simultaneously.

Response time versus resolution time

These are not the same thing, and conflating them creates problems. Response time is how long it takes you to acknowledge the incident and begin investigation. Resolution time is how long it takes the problem to be fixed. You have much more control over the former than the latter — a site might be down because the hosting provider has a major infrastructure fault, and there is nothing you can do to resolve that faster. Your SLA should commit to a response time (which is genuinely within your control) and a target resolution time qualified by circumstances outside your control. A common formulation: “We will acknowledge Critical incidents within 1 hour and target resolution within 4 hours, except where the root cause lies with a third-party infrastructure provider.”

Realistic Numbers: What You Can and Cannot Commit To

The appeal of “99.9% uptime” is understandable — it sounds reassuring and it’s the figure hosting providers often quote. But hosting providers have redundant infrastructure, dedicated operations teams, and SLAs backed by service credits that come from a revenue base of thousands of customers. As a small agency, you are passing the hosting SLA through to your client and adding your own monitoring layer on top. You are not the hosting provider.

Here is what 99.9% uptime actually means in practice: 43 minutes of permitted downtime per month. A single incident that takes your team one hour to resolve — perhaps because it happened at 2am and nobody saw the alert until morning — already puts you in breach. For most agencies, committing to 99.9% monitored uptime with a meaningful financial remedy is simply not operationally feasible.

A more defensible position for a standard agency maintenance retainer is to commit to monitoring frequency and response time rather than uptime percentage. For example: “We monitor availability every 5 minutes and will acknowledge any Critical incident alert within 1 hour during standard support hours (9am–6pm Monday to Friday), or within 4 hours outside those hours.” That is a commitment you can honour with a properly configured alerting system, without needing an on-call rota.

If a client genuinely requires 99.9% uptime with out-of-hours response, that is a premium service tier that warrants a premium price — and appropriate hosting infrastructure (managed cloud with auto-scaling, not shared cPanel). Be explicit about the dependency: your uptime commitment is contingent on the site being hosted on infrastructure that you manage or that meets a defined standard.

Exclusions: What Must Not Be Your Responsibility

A maintenance SLA without a clear exclusions section is a liability. Clients often assume that a maintenance retainer covers everything that could go wrong with their website. Here are the exclusions you should always include, in plain English.

Third-party outages. If the hosting provider, CDN, DNS registrar, or any other third-party infrastructure experiences an outage, your uptime SLA does not apply. You should still alert the client and liaise with the third party, but resolution timelines are outside your control. Name the third parties explicitly if possible.

Client-initiated changes. If the client’s team (or another contractor they’ve engaged) makes a change to the site and it breaks, that is not covered by the maintenance SLA. This is especially important for WordPress sites where the client may have admin access. Your SLA covers the site as maintained by you; it does not cover recoveries from changes made by others without your knowledge.

Plugin and theme conflicts from client-requested updates. If a client asks you to update a specific plugin and that update breaks functionality, the rectification work sits outside the standard SLA and should be quoted as additional work. The monitoring covers availability, not application-level errors introduced by deliberate changes.

Force majeure. Cyberattacks, distributed denial-of-service incidents, and large-scale infrastructure events beyond your reasonable control should be excluded from SLA calculations. This is standard in enterprise contracts and should be standard in yours too.

Service Credits and Remedies: Keeping Them Proportionate

If you breach the SLA — a Critical incident goes unacknowledged for three hours, for instance — what is the client entitled to? Many agencies either have no answer to this question or have accidentally written themselves into unreasonable positions. The industry standard approach is service credits: a proportionate deduction from the next invoice rather than a right to terminate or seek consequential damages.

A workable credit structure for a £300/mo maintenance retainer might look like this: a 10% credit (£30) for each confirmed SLA breach in a calendar month, capped at 30% of monthly fees. This creates an incentive to perform without exposing the agency to claims that dwarf the value of the contract. The key clause to include is that service credits are the client’s sole remedy for SLA breaches — this explicitly excludes claims for consequential losses such as lost revenue, which would be uninsurable and potentially unlimited.

You should also define how credits are claimed and verified. Require the client to request a credit within 14 days of a breach, with the agency’s monitoring data as the authoritative source. This prevents disputes months later based on third-party monitoring tools that may have different measurement methodologies.

One thing to avoid entirely: promising “money-back guarantees” or “refund if the site goes down.” These sound reassuring in sales conversations but create enormous commercial risk. A site that experiences a 12-hour outage due to a hosting infrastructure failure is not grounds for a full retainer refund — and any contract that implies otherwise will eventually be tested.

The Monitoring Infrastructure Your SLA Requires

An SLA is only as good as the data behind it. If you’re committing to response times and monitoring frequencies, you need systems that actually deliver alerts reliably and create an auditable record of incidents. There are three things you need in place before you can confidently sign a maintenance contract with uptime commitments.

First, automated uptime monitoring with alerting. This means a system that checks the site at the agreed interval, sends an alert via multiple channels (email and SMS is a minimum; Slack or a shared inbox is better for team visibility), and logs every check with timestamps. Free tools like UptimeRobot can handle basic HTTP monitoring, but for agencies managing 10+ client sites, purpose-built monitoring built into your agency management platform is significantly more efficient — you see all clients in one place rather than managing separate dashboards.

Second, SSL certificate expiry monitoring. An expired certificate is one of the most avoidable incidents in web management, and yet it happens regularly because agencies rely on manual reminders. Automated monitoring should flag certificates expiring within 30 days, giving you time to renew before clients notice. If you’re hosting WordPress sites, plugin vulnerability monitoring is a meaningful addition here too — flagging plugins with known security issues before they become incidents.

Third, incident logging. Every alert, every acknowledgement, every resolution note should be recorded with timestamps. This is what you produce when a client challenges your SLA compliance. Without it, you’re arguing from memory against their interpretation. Marque CRM’s site monitoring integrates uptime checks, SSL alerts, and WordPress plugin monitoring directly into the client record, so your incident history is always attached to the right account — not buried in a separate tool.

Putting the Contract Together: A Practical Template Structure

A maintenance contract with a properly drafted uptime SLA does not need to be lengthy. Here is the section structure that covers everything without creating a document nobody will read.

  1. Scope of Services — what is monitored, what updates are included, what is explicitly excluded
  2. Monitoring Methodology — tool, check interval, outage confirmation criteria
  3. Uptime Definition — percentage calculation method, calendar month basis, exclusions from calculation
  4. Incident Severity Classification — Critical vs Standard definitions with examples
  5. Response and Resolution Targets — per severity tier, inside and outside business hours
  6. Exclusions — third-party outages, client-initiated changes, force majeure
  7. Service Credits — credit value, cap, claim process, authoritative data source
  8. Remedy Limitation — credits as sole remedy, exclusion of consequential losses
  9. Review Cadence — how and when SLA targets can be renegotiated (typically annually)

If your maintenance retainer is bundled into a broader agency contract rather than a standalone document, these sections should appear as a clearly labelled schedule or addendum. The key thing is that the client signs something that references these specific terms — a general services agreement with vague maintenance language does not constitute an SLA.

One final point on tone: the language of your SLA shapes the relationship. Contracts that read like they were written by a solicitor trying to minimise liability create distrust. The best maintenance contracts are written clearly enough that the client can read them without a lawyer, and confidently enough that they communicate you know what you’re doing. The specificity is the reassurance — it shows you’ve thought about what good service looks like rather than hiding behind vague promises.

“The best maintenance contracts are written clearly enough that the client can read them without a lawyer, and confidently enough that they communicate you know what you’re doing.”

If you are managing maintenance retainers across 10 or more clients and still relying on manual checks, separate monitoring dashboards, and ad-hoc incident records, that is the operational gap that makes SLA commitments genuinely risky. Standardising your monitoring infrastructure is not an optional upgrade — it is the prerequisite for offering maintenance contracts that are commercially sustainable. See how Marque CRM’s built-in site monitoring handles uptime, SSL, and WordPress plugin checks across your entire client base from a single dashboard, and read our guides on client reporting that justifies your retainer fees and building a billing system that doesn’t leak revenue.

Run the agency this describes

90 days, every feature unlocked, no card.

Start free trial