Ahosting Logo
Knowledge Base

How to Tell Clients About Maintenance and Outages

Clients tolerate downtime far better than they tolerate silencePlanned workUnplanned outageWhen to sendwell ahead, and again shortly beforeas soon as you know something, not whenyou know everythingWhat to saythe window in their time, and what isaffectedwhat is affected, what you are doing,when you will update againWhat to avoidserver time, and jargonsilence while you investigateAfterconfirm it is donesay what happened, brieflyThe first message matters most, and the commitment to update again is the part that keeps people from chasing you.

Clients tolerate downtime much better than they tolerate not knowing about it. That single observation determines almost everything worth doing here.

Planned work

Send notice well ahead, and include four things: when, in the client's time zone instead of the server's; how long; what will be unavailable; and whether they need to do anything.

"Between 02:00 and 04:00 on Sunday, websites will be unavailable for up to fifteen minutes during a reboot. Email will queue and deliver afterwards. No action needed" tells someone everything they need in one sentence.

Then repeat it shortly before. A notice sent three weeks earlier has been forgotten, and the reminder is what prevents the tickets.

Planning and running maintenance windows explains the technical side of choosing one.

The first message during an outage

Send it as soon as the problem is confirmed, before you know the cause.

This is the message people hesitate over, because it feels wrong to write without an answer. It is the most valuable one you will send: it stops every client contacting you separately, and it converts an unexplained failure into a handled one.

"We are aware that websites on this server are unavailable. We are investigating and will update within thirty minutes." Two sentences, no cause, entirely sufficient.

Update at a stated interval, even with nothing new

Say when the next update will come, then send it whether or not there is news.

"Still investigating, no cause identified yet, next update in thirty minutes" reads as work in progress. Silence for two hours reads as nobody doing anything, and that impression is what clients remember afterwards.

Keeping the interval you promised matters more than shortening it.

Do not speculate about causes

Early theories are usually wrong, and a client told it was a hardware fault will remember that even after it turns out to have been a configuration error.

Say what you know and what you are doing. "We have identified the cause and are applying a fix" is enough detail during an incident; the explanation belongs afterwards.

Somewhere other than the affected server

A status page hosted on the server that is down is not a status page. Neither is an email sent from the mail server that has failed.

Whatever channel you use for outage communication must be independent of the thing that fails. A separate host, or a service built for it. Setting up a simple status page goes over arranging that, and it is worth doing before you need it rather than during.

Afterwards

Send a short account once it is resolved: what happened, how long it lasted, and what changes as a result.

This is the part most often skipped, because the incident is over and everyone is tired. It is also the part that rebuilds confidence, and its absence leaves clients with an unexplained outage as their last impression.

Keep it factual and avoid blame, including of suppliers. "A storage fault on the host caused an eighty-minute outage. We have added monitoring that would have alerted us twenty minutes sooner" is a good message. A paragraph about the data centre is not.

Set expectations before anything happens

Onboarding is where clients learn what to expect: how they will be told, where to check, and what response times look like.

A client who was told at the start where the status page is will look there instead of writing to you, onboarding a new hosting client goes over what belongs in that first message, and reducing support tickets with documentation walks through linking it from the places people actually look.

The one to send even when nobody noticed

A brief outage at three in the morning that no client saw still deserves a line in the next update.

Clients who discover an unmentioned outage from their own monitoring conclude that you either did not know or did not say. Both are worse than the outage, and understanding uptime guarantees goes into why their monitoring and yours frequently disagree in the first place.

Write the messages before the incident

The first message of an outage is written badly because it is written under pressure, by somebody who would rather be fixing the problem.

Three short templates solve that: an acknowledgement, an interval update, and a resolution. Each with blanks for the time, the affected service and the next update.

Having them ready changes the behaviour instead of the wording. A message that takes thirty seconds to send gets sent, and one that has to be composed gets postponed until there is news, which is exactly the delay that damages confidence.

Keep them where they are reachable when the server is down, which is not on the server. Setting up a simple status page goes into it.

Say which service, not which system

Clients do not know what a mail server is. They know whether their email works.

"The mail server is being restarted" tells them nothing actionable. "Email will be unavailable for about ten minutes; messages sent to you during that time will arrive afterwards" tells them what to expect and what they need not worry about.

The second half of that is the part usually omitted, and it is the part that stops the tickets: people write in because they do not know whether something was lost, not because the service is briefly down.

One channel, said in advance

During an incident, clients look wherever they think to look. If that is three different places, two of them will be stale.

Pick one (a status page, or email) and say at onboarding that this is where it will be. Then use it every time, including for small things, so it is a habit in place of a place people are directed to during a crisis.

Consistency matters more than the channel. A status page nobody has been told about is a page nobody checks. Onboarding a new hosting client goes into saying it at the start.

What to do when it was your mistake

Outages caused by a change you made are the ones people handle worst, usually by describing them vaguely.

Say what happened plainly: a configuration change caused it, it lasted this long, here is what stops it recurring. Clients respond better to that than to an unexplained interruption, and considerably better than to a vague reference to maintenance.

What damages trust is not the mistake: it is the discovery later that the account was incomplete. Being straightforward at the time is also the cheaper option, because it does not require remembering what was said.