Skip to main content
DNS Checker(beta)
What Happens When One DNS Provider Goes Down: The Hidden Fragility of TLD Ecosystems
Updated 10 min read

What Happens When One DNS Provider Goes Down: The Hidden Fragility of TLD Ecosystems

Ishan Karunaratne

Ishan Karunaratne

Software Architect & Infrastructure Engineer

On October 21, 2016, the Mirai botnet launched a DDoS attack against Dyn, a managed DNS provider. Twitter went down. GitHub went down. Netflix, Reddit, Spotify, and dozens of other major services went dark. The attack didn't target any of these companies directly. It targeted the DNS provider they all shared.

This article uses historical provider-concentration figures to explore shared failure scenarios. The original analysis reported 240.3 million domains across 1,929 zone files and 112 zones where one provider served more than half the observed domains. The file inventory and provider-assignment method need reconciliation before these can be treated as verified TLD-wide market measurements.

For the current provider concentration data by TLD, see the provider concentration dashboard. For provider market share rankings, see DNS provider rankings. This article explores the failure scenarios: what actually happens when concentrated DNS infrastructure goes down.

Data qualification (September 2026): The reported Cloudflare count and share (28.9 million; 11.21%) imply a denominator of about 257.8 million, while the GoDaddy figures (47.7 million; 18.49%) imply about 258.0 million. Neither matches the stated 240.3 million corpus. The historical figures remain visible, but their snapshot/counting units have not been reconciled. The 1,929 files also need an inventory to distinguish delegated TLDs from other zones or repeated snapshots; compare scope with the IANA root-zone database. Do not use these figures as current market-share estimates.

The Anatomy of a DNS Provider Failure

When a DNS provider goes down, the failure cascades in ways that aren't obvious until you trace the dependency chain:

T+0 seconds: Provider goes offline. Name servers stop responding to queries. This could be a DDoS attack, a BGP misconfiguration, a software bug, or hardware failure.

T+0 to T+30 seconds: Resolvers detect the failure. Recursive resolvers (Google, Cloudflare, ISPs) send queries to the provider's name servers and get timeouts. They retry. More timeouts.

T+30 seconds to T+5 minutes: Cached records still work. Any resolver that has the domain's records in cache continues to serve them. Domains with high TTLs (3600s or more) are temporarily shielded. Domains with low TTLs (60-300s) start failing as caches expire.

T+5 to T+60 minutes: Cascading failures begin. As caches expire, more domains become unreachable. Each failed DNS lookup triggers downstream failures:

  • Websites return connection errors
  • Email bounces with temporary failures
  • API calls timeout, breaking dependent services
  • Authentication systems fail (OAuth callbacks, SSO redirects)
  • Payment processors can't reach merchant domains
  • CDN edge nodes can't resolve origin servers

T+1 hour+: The long tail. Even after the provider recovers, resolvers that cached the failure keep returning errors for a short period. Be precise about which failure this is: an unreachable authoritative server does not produce NXDOMAIN, which is an authoritative statement that a name does not exist. It produces timeouts or SERVFAIL. RFC 9520 requires resolvers to cache such resolution failures for at least one second and no longer than five minutes, so the tail here is bounded by that, not by the zone's negative-cache TTL.

This is a failure scenario, not a reconstruction of every cited provider incident. A routing outage can affect services differently from an authoritative DNS outage; establish the affected products and network paths before attributing a historical incident to DNS concentration.

Why TLD-Level Concentration Matters

The historical table reports GoDaddy at 18.49% and Cloudflare at 11.21%, subject to the unresolved denominator above. Global shares can conceal higher concentration within a particular zone; neither measure alone determines outage impact.

When I measured concentration at the TLD level, the picture is very different. Some TLDs have provider diversity. Others are essentially single-vendor ecosystems.

The numbers: 36 TLDs have a single provider above 90%. 62 above 70%. 112 above 50%. For those 112 TLDs, a single provider outage doesn't just affect "some" domains. It could affect a large share of domains that depend exclusively on that provider.

Consider .shop with over 1.1 million domains and 62% at a single provider. Or .top with over 4.2 million domains and nearly 59% at one provider. These are examples from the historical report, not a verified current ranking or an outage forecast.

Browse the full TLD-level concentration data, including HHI scores, on the provider concentration dashboard.

Scenario: Failure of a Widely Used Provider

The historical report assigns 28.9 million domains and an 11.21% share to Cloudflare. These values should not be read as current counts or used to forecast failure while their denominator remains unreconciled. A high zone-specific provider share is an exposure indicator, not proof that every dependent domain would stop resolving.

The June 2022 Cloudflare outage report illustrates the shape of the risk, though it is worth describing accurately. A routing configuration change took 19 data centers offline between 06:27 and 07:42 UTC, roughly 75 minutes, and affected about half of Cloudflare's HTTP requests. It was a geographically partial network outage rather than a global authoritative-DNS failure across every TLD Cloudflare serves. The concentration concern is what a comparable event would mean if it did hit the authoritative DNS path for a TLD where one provider holds the majority of delegations.

Now imagine a sustained attack. Not 90 minutes, but 90 hours. Persistent DDoS against Cloudflare's authoritative DNS infrastructure, targeting the anycast prefixes that serve their name servers. The exposure could be substantial, but the unreconciled historical count is not an outage forecast. Domains with independent authority or usable cached answers may continue resolving.

The Compounding Dependencies

DNS concentration risk is worse than it appears because DNS providers aren't just DNS providers. Modern infrastructure stacks create dependency chains:

Cloudflare offers DNS alongside CDN, WAF, and compute services. A domain can depend on several of those products, but an authoritative DNS failure does not automatically disable the others. Map each dependency and the effect of its failure before designing failover.

GoDaddy offers DNS alongside other services. A shared supplier can create correlated dependencies, but the reported 47.7 million domain figure does not establish how many customers use its full stack or how an outage would affect them.

This means multi-provider DNS, the standard recommendation for concentration risk, only partially mitigates the problem. If your Cloudflare DNS failover works but your Cloudflare CDN is also down, your domain resolves to an IP that serves errors.

The HHI Framework: Measuring Market Power in DNS

I use the Herfindahl-Hirschman Index (HHI) to quantify concentration beyond just the top provider's share. HHI sums the squared market shares of all providers in a TLD. The U.S. Department of Justice uses HHI to assess market concentration in antitrust cases, the same framework applies to DNS.

The 2023 DOJ/FTC Merger Guidelines return to these concentration thresholds (Guideline 1 and footnote 15):

  • HHI below 1,000: unconcentrated
  • HHI 1,000 to 1,800: moderately concentrated
  • HHI above 1,800: highly concentrated

(If you have seen 1,500 and 2,500 quoted instead, those are the superseded 2010 thresholds, which is what I originally used here.)

The TLDs I've flagged with the highest concentration have HHI values of 3,000 to 5,000+, which is far into the highly concentrated range. Two cautions on reading those numbers, though.

First, a high HHI does not mean a single provider. HHI is the sum of squared shares, so two providers splitting a TLD 50/50 produce exactly 5,000, and a 70/30 split produces 5,800. In the 50/50 case either provider failing takes out half the TLD's domains, not all of them. Only an HHI of 10,000 necessarily means one provider with 100% share. To estimate actual outage impact you have to look at the top provider's share directly, which is why I report that separately above.

Second, this is an analogy, not antitrust doctrine. The DOJ's structural presumption applies to a merger that raises HHI by more than 100 points in an already-concentrated market; a static HHI reading, however high, does not mean a market "would trigger antitrust review." I borrow HHI because it captures the shape of a distribution better than a top-provider percentage alone, not because DNS hosting is under antitrust scrutiny.

Why Concentration Persists

If concentration is so risky, why does it persist? Several structural factors:

Registrar bundling. Most domain owners register a domain and get DNS hosting bundled for free. They never think about DNS as a separate service because it comes with the registration. This means the largest registrars (GoDaddy, Namecheap, IONOS) automatically become the largest DNS providers.

Migration friction. Moving DNS to a different provider requires changing NS records at the registrar, recreating all zone records at the new provider, and waiting for propagation. This takes effort and creates a risk window. Most domain owners never bother.

Pricing. Premium DNS services (NS1, DNSimple, Route 53) charge monthly fees. The bundled DNS from registrars is free. For the millions of domains that cost $10/year to register, paying $5/month for DNS doesn't make economic sense.

Multi-provider complexity. Running DNS across two providers requires zone synchronization, monitoring, and failover logic. This is operationally complex and not well-supported by most registrar interfaces. Only enterprises with dedicated DNS teams typically maintain multi-provider setups.

RFC 8767 permits resolvers to serve stale answers during some failures. Likewise, failure of authoritative DNS does not automatically imply that the same provider’s CDN or WAF has failed; those dependencies need separate analysis.

Practical Resilience: What Actually Helps

For Critical Domains

If your domain is business-critical (revenue-generating, customer-facing, authentication infrastructure), the investment in multi-provider DNS is worth it:

  • Primary + secondary DNS. Configure your domain with NS records from two independent providers. Use zone transfer (AXFR/IXFR) or API sync to keep records consistent. If one provider goes down, the other continues serving.
  • DNS provider on different infrastructure. Ensure your two DNS providers use different networks, different anycast prefixes, and different hosting infrastructure. Two providers that share the same upstream transit don't provide real independence.

For All Domains

  • Set reasonable TTLs. During an outage, cached records are your lifeline. A TTL of 3600 (1 hour) gives you a buffer. A TTL of 60 seconds means you're exposed almost immediately.
  • Monitor NS responsiveness. Use external monitoring that queries your name servers from multiple locations. An alert on NS timeout lets you respond before users notice.
  • Know your TLD's concentration. Before registering a domain, check whether the TLD has healthy provider diversity. Be careful how you read that risk, though: it is a property of the TLD's domain population, not an extra dependency on your own domain. Your zone is served by the nameservers named in your own delegation, so if you use providers independent of the TLD's dominant one, you are not exposed to that provider's outages just because your neighbours are. What high concentration predicts is a large correlated failure across that TLD, which matters for the ecosystem and for anyone depending on many domains under it, and it is a reasonable signal about the registrar defaults you will be nudged toward. The same provider-bundling dynamic that drives concentration also shapes whether a TLD's domains are DNSSEC-signed, so a TLD dominated by one provider tends to inherit that provider's security defaults too.

For the Ecosystem

  • Registries should publish concentration metrics. If ICANN required registries to report HHI or top-provider share as part of their annual compliance, it would create transparency and pressure for improvement.
  • Registrars should make multi-provider DNS easier. A registrar that offers one-click setup for secondary DNS with a partner provider would differentiate itself on resilience.
  • The research community should model correlated failures. The interaction between DNS concentration, BGP routing concentration, and CDN dependency creates failure scenarios that are poorly understood and underresearched.

For live provider concentration data across all TLDs, see the provider concentration dashboard. For provider market share rankings, see DNS provider rankings. To explore individual providers, visit the provider directory.


Frequently Asked Questions

Sources

This article was researched and structured by the author with AI assistance for drafting and technical verification.

About the Author

Ishan Karunaratne
Ishan Karunaratne

Software Architect & Infrastructure Engineer

US Army veteran with a B.S. in Information Technology, CompTIA A+, Network+, and Security+ certified. 20+ years building and securing web infrastructure.

B.S. Information Technology, Online SystemsCompTIA A+ (2009)CompTIA Network+ (2009)CompTIA Security+ (2009)US Army Veteran, Operation Iraqi Freedom

Share this article

DNS terms in this guide

Plain-English definitions for the key terms referenced above.

Related Articles

Typo-Like Nameservers: Investigating 145,061 Historical Delegations

A historical pipeline flagged 145,061 delegations to a typo-like nameserver domain. Similar spelling alone does not establish malicious control or takeover.

How Expired Name Servers Become Domain Hijacking Vectors

Historical nameserver-delegation candidates illustrate why owners should verify authority, provider status and domain control. An expiry-looking hostname does not prove takeover risk.

Why DNSSEC Is Still Failing: Lessons from 240 Million Domains

Historical zone snapshots reported low parent-DS presence. Examine the measurement limits, provider incentives, and operational reasons DNSSEC deployment can remain incomplete.

Complete Guide to DNS Attacks and DNS Security (Prevention, Testing & Mitigation)

A comprehensive guide to DNS attack types including cache poisoning, amplification, tunneling, zone walking, and hijacking. Learn how attackers exploit DNS, how to test your own domains, and how to harden your infrastructure.