>samit_hota
Back to security news
SN-2026-182HighResolved

Network Maintenance Automation Bug Triggers Widespread Microsoft 365 Outage

Samit Hota·
CVE ID
N/A
Affected Products / Orgs
Microsoft 365, Microsoft Azure, Exchange Online, Microsoft Teams
#news#general-incident#microsoft

Automated operational guardrails failed during a routine infrastructure update, resulting in a severe Microsoft 365 outage that left enterprise users worldwide unable to access Exchange Online, Microsoft Teams, and underlying Azure services. The incident highlights the inherent systemic risks of centralizing administrative power within automated network management software.

Technical Root Cause

The disruption originated within Microsoft’s automated network maintenance request system, a internal service designed to orchestrate routine router and switch updates across global data centers. During a scheduled maintenance job, a validation bug within the control software executed a series of unintended routing changes. Rather than modifying a isolated subnet, the routine inadvertently removed core BGP (Border Gateway Protocol) and internal IP routes across a significantly larger fleet of routing devices than authorized.

This unexpected route revocation effectively severed communication paths between core Azure management planes, authentication endpoints, and front-end service gateways. Because the automated system operated with elevated network orchestration privileges, the faulty instructions propagated rapidly through the backbone network before manual safety cutoffs could intervene.

Operational Impact and Recovery

The propagation of invalid routing tables caused immediate service degradation across multiple geographic regions. Enterprise tenants experienced lost telemetry, failed logins, and dropped real-time media streams in Teams, alongside mail routing delays in Exchange Online. Dependent third-party applications relying on Azure Active Directory (Microsoft Entra ID) authentication also faced widespread connection timeouts.

Engineers stabilized the environment by halting all automated network jobs, isolating the maintenance service, and manually restoring valid IP routing tables across affected edge and core routers. Microsoft subsequently confirmed that telemetry and connection queues cleared once valid topology routes re-established across internal backbones.

Mitigating Automation Control Risks

While this incident was triggered by internal maintenance tooling rather than external adversary activity, it underlines critical operational safety practices for enterprise infrastructure teams:

  • Stricter Canarification of Network Changes: Automation scripts must execute in isolated blast-radius zones with obligatory validation pauses, preventing changes from hitting global edge routers simultaneously.
  • Out-of-Band Control Planes: Administrative management channels should remain topologically distinct from production routing paths so engineers retain operational visibility even when core BGP tables collapse.
  • Automated Route Pre-Validation: Network orchestration engines should evaluate intended BGP updates against hard operational policy bounds before pushing route revocations live.

Found something similar in your stack?

Let's find out before it becomes an incident.

Book an advisory call