Microsoft 365 Outage Response: What Your MSP Should Do
By Tom Hermstad · HD Tech

What should your MSP do when Microsoft 365 goes down?
A proper Microsoft 365 outage response starts in the first few minutes — not after you call your MSP to ask what's happening. Your provider should immediately triage whether the fault is local or Microsoft-wide, activate backup communication channels outside of email, and issue you a plain-English status update — not a link to the Microsoft status page. If your MSP's response is "we're waiting for Microsoft to fix it," that's not managed IT. That's a spectator sport.
How Often Does Microsoft 365 Actually Go Down?
More often than most business owners realize. Even with Microsoft 365 platform-wide quarterly uptime figures that have ranged between 99.5% and 99.99% — with Q1 2026 coming in as low as 99.526% — outages do occur, and when they do, the impact can be severe. Documented incidents show that some outages have lasted many hours before full resolution.
For many SMBs, that kind of outage window — with a sales team unable to send email or pull a quote — translates directly to lost revenue. And the numbers back that up: Gartner estimates the average cost of IT downtime at $5,600 per minute, which compounds fast when a production floor goes idle or a bid deadline passes with no response sent.
Multiple incidents. Back to back. Affecting businesses across industries.
AI-powered monitoring tools are already changing how fast a prepared MSP detects these problems — flagging degraded authentication response times, unusual login failures, and latency spikes before the first user files a help-desk ticket. That early-warning layer matters when every minute of downtime carries a measurable price tag. This is the forward posture a modern MSP should hold — and it's exactly what we build into the Lifeguard Loop™ from day one.
For every client caught in that outage window, the real question wasn't "Is Microsoft fixing it?" The real question was: What was your MSP doing while you sat there unable to send email?
Whether you run a manufacturing floor, a medical practice, a law firm, or a financial services operation — downtime isn't just inconvenient. It's a liability. You have contractual obligations, compliance requirements, and clients expecting answers.
Your competitor's team submitted that bid. Yours didn't. That's how managed IT becomes your competitive edge — and it only happens when your MSP has a practiced response plan, not a reactive scramble.
Over my career protecting Orange County businesses, I've watched this exact silence cost clients contracts. That's not a hypothetical — it's a pattern I've seen play out more times than I care to count, and it's exactly why preparation beats reaction every single time.
What a Competent MSP Does — Minute by Minute
This is the Relentless Response Engine™ in action. Not a checklist that lives in a binder. A practiced, documented process your team can execute without scrambling.
How does your MSP determine where the fault actually lives?
The first job is scoping the problem — fast. According to Microsoft's own admin guidance, admins must actively monitor the Service Health dashboard in the Microsoft 365 Admin Center — not wait for users to report symptoms.
A competent MSP checks:
- Is this one user, one site, one tenant, or a regional Microsoft issue? The answer dictates everything.
- Is local DNS, internet routing, or authentication the culprit — Azure AD / Entra ID, meaning the Microsoft system that controls who can log in to what, or is it Microsoft's core infrastructure?
- What's affected and what still works? SharePoint down doesn't mean Teams is down. Exchange affected doesn't mean OneDrive is offline.
You should have a plain-English answer from your MSP promptly — not "we're looking into it." Specifics.
A seasoned MSP runs AI-powered monitoring at this layer — tools that detect outage signals like degraded authentication response times, unusual login failures, and latency spikes before a single user reports a problem. The nephew who set up your M365 account doesn't run that layer.
How does a real MSP communicate during an active outage?
Here's where the gap between a real MSP and a reactive one becomes obvious.
I've taken those early-morning calls from business owners whose MSP's entire response was a status-page link. No explanation. No workaround. No next update time. Just a URL. That's not incident management — that's passing the buck.
Industry guidance for MSPs is clear: if email is down, you cannot use email to communicate the outage. Your MSP should already have a non-email channel established — SMS trees, phone bridges, a status dashboard, or a direct-dial protocol.
Shortly after the outage is confirmed, your staff should know:
- What's down and what's still functional
- What workarounds to use right now (cached Teams access, mobile apps, phone)
- What they should NOT do (don't mass-restart endpoints, don't attempt password resets if authentication is the issue)
- When the next update is coming
That last point matters more than people realize. Committing to a specific next update time buys patience. Silence breeds panic and shadow IT.
What happens when an outage stretches beyond the first hour?
When an outage stretches on, your MSP's true competence — or lack of it — is completely exposed.
A pro MSP already has the playbook open and is shifting into continuity mode fast. The nephew who set up your M365 account isn't running a triage checklist early in the morning.
I had a client years ago — a mid-sized manufacturer in Orange County. Exchange Online went down one morning. Their previous MSP sent one email (to an email system that was offline), then went quiet for hours.
By the time I got brought in on an emergency basis, their sales team had missed a bid deadline and their controller couldn't send wire instructions to a vendor. That silence cost them a contract. That's not a horror story I'm inventing — that's what happens when there's no continuity plan and no communication protocol.
Honestly, this is the part of the job I love most — not because downtime is fun, but because having the right response ready before the crisis hits is where years of hard-won experience pays off in a real, measurable way. At HD Tech, this is what we mean when we say preparation over prevention.
Why inbound email continuity is non-negotiable
When Exchange Online goes offline, inbound email doesn't just queue up waiting patiently. Without a continuity service in place, messages may bounce back to senders rather than hold safely for delivery.
MSPs can implement inbound email continuity through secondary mail services or MX record failover — that's a backup routing system that redirects incoming email to a secondary server so messages hold safely and deliver once Exchange comes back online, instead of bouncing back to senders. If your MSP hasn't set this up, ask them why today.
Spambrella's Continuity Service, for example, includes an emergency inbox and mail spooling — meaning inbound messages are held safely and your team can continue sending and receiving through an emergency inbox for the duration of an outage.
Your MSP should also be:
- Maintaining a running incident record — a timestamped log of every action taken, every status change, and every communication sent, giving you documented proof that the response was handled properly
- Monitoring Microsoft's Service Health for updates and translating those updates into business language before forwarding them to you
- Documenting closure — not just "it's back up" but a post-mortem summary (a short written summary of what broke, how long it lasted, and what changes so it doesn't happen the same way again) with what happened, how long it lasted, and what changes (if any) reduce exposure next time
If your MSP's incident record is a Slack thread with a handful of messages, that's not an incident record. That's a conversation.
The Two Things Most MSPs Get Wrong
1. Treating the Microsoft Status Page as the Answer
Forwarding a link to the Microsoft Service Health dashboard is not incident communication. It's deflection. Your team doesn't know how to read it, and it tells them nothing about what to do next. Your MSP owns the translation — technical status into business impact, with clear instructions.
2. Declaring It "Microsoft's Problem" and Going Passive
Yes, Microsoft owns the infrastructure. No, that doesn't mean your MSP has nothing to do. Real incident response for MSPs includes owning the communication, activating continuity services, scoping workarounds, and documenting everything — independent of when Microsoft resolves the underlying fault.
It's not if an outage happens. It's when. The question is whether your team ever stops working because of it.
What does a Microsoft 365 outage mean for compliance-sensitive businesses?
Here's something I see trip up business owners more than almost anything else: they treat outage response and compliance as two separate problems. They're not.
When that outage happens, your incident record — your timestamped, documented proof that your team responded properly — isn't just good IT practice. It's evidence.
If you're in a regulated industry — healthcare, financial services, legal, or government contracting — an M365 outage isn't just a productivity problem. It's a compliance exposure.
Here's the question your auditor will eventually ask:
If I asked your MSP to produce the incident record from the last M365 outage today, what would they hand over?
A timestamped log with actions taken and communications sent? Or a Slack thread and a shrug?
What auditors actually look for after an outage
An extended outage with no documented incident response is a gap your next auditor will find.
For a manufacturer, that gap shows up immediately — a purchase order stuck in limbo, or a floor supervisor who couldn't pull a digital work order and had to halt a production run. According to Verizon's 2026 Data Breach Investigations Report Manufacturing Snapshot, System Intrusion, Social Engineering, and Basic Web Application Attacks represent 91% of breaches in manufacturing — and auditors in that space are increasingly scrutinizing outage documentation as part of vendor risk reviews.
Compliance frameworks don't care that Microsoft had a rough Tuesday.
What your MSP must document — every time
Your compliance posture lives or dies on three things your MSP must deliver:
- Incident record — timestamped proof that your team responded properly, start to finish
- Continuity activation — evidence that you had controls in place before the outage, not scrambling during it
- Post-mortem summary — documented proof of what happened, how long it lasted, and what changed to reduce exposure next time
What every healthcare CEO should know about HIPAA in 2026 — here's what that looks like when HIPAA meets audit season and availability requirements are on the table.
For any business with contract deadlines, document exchange obligations, or client SLAs tied to uptime — downtime has consequences that extend well beyond the hours the system is offline. You still needed a plan.
The MSP Accountability Checklist
Trust your MSP. Then verify. Every single time. That's not paranoia — that's years of experience talking. The Lifeguard Loop™ is built on exactly this principle: we earn trust by showing our work, and we expect you to hold us to it.
During a major SaaS outage, your MSP should:
- Confirm fault location (local vs. tenant vs. regional Microsoft issue) promptly
- Issue a plain-English business impact summary — not a status page link
- Activate non-email communication channels if email is affected
- Brief your staff on what's down, what still works, and what not to do
- Implement inbound email continuity if Exchange is affected for an extended period
- Maintain a timestamped incident record throughout
- Deliver a written post-mortem summary promptly after resolution
- Draw on years of local, seen-it-before experience — not a gut reaction to an unfamiliar situation
Walk through that checklist after every outage — not as an exercise in blame, but as a measure of competitive readiness. Every box your MSP can check is a box your competitor's vendor might not be able to.
That's how managed IT becomes your competitive edge. The businesses that can keep operating during an outage, document it cleanly, and hand an auditor a complete incident record are the ones that win contracts and keep clients.
The ones who can't are the ones who send a status-page link and go quiet.
If your MSP handled only some of these during the last outage you experienced, you have an accountability gap. That gap doesn't just cost you hours. It costs you client trust, compliance standing, and revenue you'll never fully account for.
Frequently Asked Questions
Your MSP should confirm within the first few minutes whether the issue is local or Microsoft-wide — and get you a plain-English summary fast: what's broken, what still works, and what your team should do right now. They check the Service Health dashboard directly rather than waiting for users to flood the help desk. Silence in the early stages is a failure — full stop.
Your MSP can't flip a switch at Microsoft's data center — but they can confirm scope, switch your team to backup communication channels, and activate email continuity so inbound messages don't bounce. They should walk your staff through workarounds and document everything in real time. Real MSP guidance treats every outage as a formal incident — triage, communicate, document, close. Sitting on their hands is not incident management.
An extended outage with no documented incident response is a gap your next auditor will find. HIPAA, financial services obligations, and contractual SLAs don't care that Microsoft had a rough Tuesday. Your MSP must maintain a timestamped incident record, activate continuity services, and deliver a written post-mortem after resolution. "We waited for Microsoft to fix it" is not a compliance answer.
Ask them to show you the document. A real process has a triage checklist, a communication protocol that doesn't depend on email, defined escalation paths, and a post-mortem template. If they look at you blankly or say "we handle it case by case," that's your answer. A prepared MSP running a real Relentless Response Engine™ already has the playbook open before your first user reports a problem. That's the standard — no exceptions.
The triage steps are the same regardless of industry. What changes is the consequence — and how fast it compounds. A manufacturer loses production hours. A healthcare practice risks HIPAA exposure. A law firm misses a filing deadline. Whatever your business does, the Relentless Response Engine™ delivers a documented, practiced process that holds up under pressure and hands your auditor a complete record. That's not a nice-to-have. That's the job.
The nephew who set up your M365 account won't have a triage checklist ready at 7 a.m. — a seasoned pro with many years of experience will.
Got a few minutes? Book a Discovery Call with Tom directly. Ask your toughest questions. Get honest, plain-English recommendations on where your security posture stands and what to shore up — from someone who's been protecting Orange County businesses for many years.
No pitch. No pressure. Just hard-won experience working for you.
Schedule your Discovery Call with Tom →
It's not if, it's when.

Tom Hermstad
President & CMO, HD Tech
Tom Hermstad has led HD Tech since 1995, building one of Southern California's most trusted managed IT and cybersecurity firms. He specializes in helping Orange County businesses eliminate IT headaches and stay ahead of evolving cyber threats — in plain English.
