📚 Free Download: The Complete Documentary Sources & Data Guide

Every statistic in this documentary, with its original source, in one free PDF you can keep and cite.

Download Free PDF Book

Why “Documentary,” Not “Post”

You will hear people say Google removed millions of websites from the internet. You will hear that Google deletes developer accounts and apps by the hundreds of thousands. You will hear that Google crawls your website to train its AI, not to help people find it. You will hear that Google has simply taken over the internet.

Every one of those sentences contains something true. Almost none of them are stated precisely. This piece exists to close that gap — to take the raw grievance, much of it my own, and hold it up against the actual public record: court testimony, official transparency reports, independent trackers, and our own site’s real Search Console data.

This is not a rant. It is not a press release. It is a documentary, built the way a documentary should be built: claim by claim, source by source, with the exaggerations trimmed off and the genuinely damning parts left standing — because, as it turns out, the accurate version is damning enough on its own.

Chapter 1: “Google Removed Millions of Websites From the Internet” — What Actually Happened

Start with the biggest, angriest claim, because it’s also the most technically wrong — and the most revealing once you correct it.

Google did not delete millions of websites from the internet. A deindexed site is still online. You can still open it in a browser by typing the URL directly. Other search engines can still crawl it. What Google does is remove it from its own index — the database that decides what appears in Google Search results. For a site that depends on Google for 80–95% of its visitors, the practical difference between “deleted from the internet” and “erased from the only door people actually use to find you” is close to zero. But the words mean different things, and a documentary that wants to be trusted has to use the right ones.

So what is actually documented?

The scale of enforcement. Google’s March 2024 core update, combined with new spam policies introduced the same month — scaled content abuse, site reputation abuse, and expired domain abuse — was aimed at what Google itself described as “low-quality, unoriginal content.” Google said the update reduced this kind of content in search results by 45%, beating its own internal 40% target. Independent trackers recorded roughly 1,446 websites hit with “Pure Spam” manual actions in the early weeks of that rollout alone.

Named casualties, not vague statistics. This is where the story stops being abstract:

  • ZacJohnson.com, a site that had been publishing 325+ articles a day between September 2023 and March 2024 — over 60,000 articles in six months — was deindexed as what SEO analysts call a textbook case of “scaled content abuse.”
  • Site reputation abuse (“parasite SEO,” where a large, trusted domain hosts a low-quality affiliate section to borrow its authority) triggered manual actions against affiliate and coupon subdirectories on CNN Underscored, USA Today, LA Times, Forbes Advisor, WSJ Buyside, Time Stamped, AP Buyline, and Men’s Journal through 2024. Per Sistrix visibility data: Time Stamped lost roughly 97% of its visibility, WSJ Buyside about 77%, CNN Underscored about 63%, and Forbes Advisor about 43%.
  • AP Buyline was shut down entirely. The European Commission has since opened a formal investigation into whether Google’s site reputation abuse policy unfairly targets legitimate publisher monetization — that inquiry was still open as of late 2025.

Notice something important here: several of the biggest names hit by this specific policy were not small independent creators. CNN, USA Today, Forbes, and the Wall Street Journal are about as “big brand” as it gets, and their affiliate sections still got demoted. The “big brands are always safe” version of this story is too simple. What actually happened is more specific: Google went after a particular monetization pattern — trusted domains renting out their authority to third-party content farms — regardless of the parent brand’s size.

But independent creators absorbed the deepest wounds. The September 2023 Helpful Content Update is the other half of this story, and here the “small creators get crushed” narrative holds up under scrutiny. HouseFresh, an independent air-purifier review site, lost roughly 91% of its Google traffic — from about 4,000 daily visitors down to around 200. Retro Dodo lost about 85%. A Digitaloft study of 671 independent travel publishers found 78% of them lost traffic between August 2022 and March 2024. HouseFresh did eventually recover, in October 2025, after a further core update — but SEO analyst Glenn Gabe’s assessment is blunt: HCU classifications tend to be “sticky once applied,” and most sites that get hit never see that kind of recovery.

What Google itself has stopped telling us. Google used to publish detailed annual “how we fought spam” transparency reports. In 2019 it disclosed discovering roughly 25 billion spammy pages a day and sending over 90 million messages to webmasters. In 2022 it said its spam-detection system, SpamBrain, was catching five times more spam sites than the year before. Then, conspicuously, Google did not publish its usual detailed webspam report for 2023 — the exact year generative AI caused a documented flood of low-effort content across the web. Draw your own conclusion about the timing, but the gap in Google’s own disclosure is itself part of the record.

The honest verdict on Claim 1: “Millions of websites removed from the internet” is not an accurate description of anything that has happened. What is accurate, documented, and serious: Google has deindexed or demoted hundreds of named sites and vastly more individual pages through 2024’s spam-policy enforcement, the effect on a site’s business can be as final as being taken offline even though the site technically still exists, and independent publishers have suffered disproportionately compared to large brands — except in the specific case of parasite-SEO enforcement, where big names got hit too.

Chapter 2: The Play Store Purge — This One Is Real, and the Numbers Are Official

If Claim 1 needed correcting, Claim 2 barely needs any correction at all. This is the part of the story with hard, Google-published numbers behind it.

As part of a 2021 developer class-action settlement (the same litigation that produced a $90 million payout to roughly 48,000 US developers), Google agreed to publish an annual Play Store transparency report. The most recent figures, compiled from that report:

Year Apps removed Developer accounts terminated Accounts reinstated
2024 ~3.9 million (~11,000/day) ~155,000 ~1,500
2025 ~2 million ~126,000 ~12,500

The 2024 breakdown of why apps were pulled: 55% for data protection/privacy violations, 16% for ineligible content, 15% for data defects, and 9% for scams or fraud.

There is a separate, easily confused metric worth naming precisely, because press coverage regularly mixes the two up: Google’s Android Security team also publishes numbers on apps blocked before they ever reached the Play Store — 2.36 million in 2024, alongside 158,000 “bad” developer accounts banned pre-publication. That figure (158,000 banned) is not the same as the 155,000 terminated accounts from the Transparency Report, even though the two numbers are close enough that they get conflated constantly in secondary reporting. This documentary is treating them as what they are: two different reports, measuring two different things, that happen to land near each other.

The catalog shrank by nearly half. Independent app-analytics firm Appfigures found the live Play Store catalog fell from roughly 3.4 million apps in early 2024 to about 1.8 million by early 2025 — a 47% drop, driven largely by a July 2024 minimum-functionality and spam policy that took effect that August, plus tighter developer verification. In the same window, Apple’s App Store grew slightly (1.6 million to 1.64 million), which is the detail that rules out “the whole industry is shrinking” as an explanation. This was a Google-specific policy decision.

The part the transparency report doesn’t capture: wrongful terminations. A recurring pattern shows up across independent developer complaints — accounts terminated for “a pattern of high risk or abuse” or flagged as “associated” with a previously banned account, via automated systems that give no specifics and offer opaque, often unsuccessful appeals. Google does not explain what makes one account “related” to another. Developers have documented clear false positives — in one case, a security vendor (Ikarus) confirmed its own scanner had wrongly flagged a legitimate companion app as a trojan. In early 2026 Google formalized a 180-day appeal window, after which a termination becomes permanent regardless of merit.

The honest verdict on Claim 2: This claim is true and unusually well-documented, because Google is legally required to publish the numbers. Roughly 3.9 million apps and 155,000 developer accounts in 2024 alone, a catalog that shrank by nearly half in a year, and a real, credible pattern of automated wrongful terminations with weak recourse. If anything, this is the strongest single claim in the entire complaint.

📚 Get the Full Documentary Sources

Download the complete PDF with every citation, court quote, and dataset used in this investigation.

Download Free PDF Book

Chapter 3: “Google Crawls to Train Its AI, Not to Index You” — Partly True, and Confirmed Under Oath

This is the claim I hear most often from creators, usually stated with total confidence, and it turns out to be the most technically tangled one — and the one where the truth is stranger, and more useful to know, than the simplified version.

The technical reality first. Googlebot is the crawler that indexes your site for Search. Google-Extended is not a separate crawler at all — it’s a robots.txt control token, introduced in September 2023, that governs whether content Google has already crawled gets used to train and ground Gemini and Vertex AI. It has no distinct user-agent string of its own; Google crawls with its existing bot identities either way. Google’s own documentation, updated in 2024 and reaffirmed as recently as April 2025, states plainly: “Google-Extended does not impact a site’s inclusion or ranking in Google Search.” Blocking it costs you nothing in Search.

Here is the nuance almost everyone gets wrong, in both directions. Blocking Google-Extended does not remove your content from AI Overviews or AI Mode — because those features live inside Search itself, powered by ordinary Googlebot-indexed content, not by the Google-Extended training pipeline. The only tools that suppress AI Overviews are blunt instruments — nosnippet, data-nosnippet, max-snippet, or noindex — and every one of them also damages your normal search visibility. There is currently no surgical way to stay fully visible in Google Search while opting your content out of Google’s AI-generated answers. That’s not a rumor. That’s the documented mechanism.

And then there’s the sworn testimony. In the September 2025 US antitrust remedies trial, DOJ attorney Diana Aguilar asked Google DeepMind VP Eli Collins directly: “the search org has the ability to train on the data that publishers had opted out of training, correct?” Collins answered: “Correct — for use in search.” An internal Google document shown in that same trial revealed that Google removed 80 billion of 160 billion training tokens after honoring opt-outs for DeepMind model training specifically — but that opt-out never extended to how the Search organization uses that same content to power AI Overviews.

The honest verdict on Claim 3: The strict version — “Google secretly funnels all crawling into AI training under the guise of indexing” — is not accurate; Google-Extended genuinely separates AI training from indexing on paper. But the deeper accusation is validated by testimony given under oath in federal court: the Search organization does train on content that publishers explicitly opted out of for AI purposes, provided that use is “for search.” For a creator, the practical upshot is the one that matters: there is no clean door out of Google’s AI answers that doesn’t also cost you Google Search visibility.

Chapter 4: How Big Is “Taken Over”? The Numbers Behind the Phrase

Market share. Per StatCounter Global Stats, Google held 90.39% of worldwide all-device search in May 2026. It peaked at 92.9% in 2023, dipped to an unusually low 89.57% in July 2025 — its steepest annual decline in a decade — and partially recovered after the launch of AI Mode. On mobile specifically, Google’s share sits around 94.6%; on desktop, roughly 79–82%, its lowest in over 20 years. By country, India runs around 97% and the US around 85%. Chrome itself holds about 69.65% of the browser market, reinforcing Google as the default search gateway for most of the world before a user even opens a new tab.

Scale. Google has confirmed over 5 trillion searches per year as of 2025 — roughly 14 billion a day — up from the “2 trillion” figure it last cited back in 2016.

The click collapse. This is the number that should worry every publisher, big or small. Ahrefs, analyzing 300,000 keywords through Google Search Console data (published February 2026), found top-ranking pages now see a 58% lower average click-through rate than they did just eight months earlier. In their own words: “out of every 100 clicks that once went to top-ranking websites, Google now keeps 58.” Position-1 click-through rate on AI-Overview-triggering queries fell from 0.073% to 0.016% between December 2023 and December 2025 — a roughly 78% decline. Separately, Pew Research Center (July 2025, 900 US adults, 68,879 real searches from March 2025) found that when a Google user encountered an AI summary, they clicked through to a traditional result in only 8% of visits — versus 15% when no AI summary appeared. Only 1% clicked a link inside the AI summary itself.

The lawsuits this has already produced. Chegg sued Google in February 2025 after non-subscriber traffic fell 49% year-over-year (versus an 8% decline a year earlier); its stock is down more than 98% from its 2021 peak. Penske Media — publisher of Rolling Stone, Billboard, Variety, and The Hollywood Reporter — sued in September 2025, the first major US publisher suit specifically over AI Overviews, citing affiliate revenue down more than a third from its peak. Google’s public response, through spokesperson José Castañeda, is that both suits are “meritless” and that “AI Overviews send traffic to a greater diversity of sites.” In 2024, federal judge Amit Mehta ruled that Google illegally maintains a monopoly in search; in 2025 he ordered Google to share certain search data with competitors.

The honest verdict on Claim 4: “Taken over the internet” is a rhetorical flourish, not a technical claim — but for search and content discovery specifically, it is directionally accurate almost to the point of understatement. A company with roughly 90% of global search share is not one competitor among several. It is closer to the only door.

Chapter 5: Our Own Data — The Part of This Documentary I Didn’t Have to Look Up

Everything above came from public records, court transcripts, and third-party analytics firms. This part comes from our own Google Search Console account for this website, checked the same day this documentary was written.

Right now, this site has 1 indexed page against 292 pages sitting unindexed. Over the trailing 12 months, total organic clicks from Google Search: 14. Total impressions: 1,090. Average click-through rate: 1.3%. Average position: 4.1.

Read that last figure again. An average ranking position of 4.1 — genuinely competitive, first-page placement — produced 14 clicks across an entire year. That is not a contradiction; it is exactly what the Ahrefs and Pew data above predict when so much search real estate now goes to AI summaries and zero-click features instead of blue links.

There’s a second, stranger detail in our own Search Console data worth being honest about: some of the top search queries bringing impressions to this site have nothing to do with what this site is. Terms like “your shop,” “shopsy,” “is your shop open today,” and “where is your shop” appear in our top-50 queries with real impression counts, alongside entirely unrelated religious-history queries like “jewish god” and “judaism god.” We do not run a shop. We have never published content about that topic. Whatever Google’s index currently associates with this domain, it does not fully match what this site actually is or publishes — which is itself a small, honest data point about how imprecise even Google’s own understanding of a site’s content can be.

I’m including this not as a complaint but as evidence, in the same spirit as the rest of this documentary: this is what “292 unindexed, 1 indexed, 14 clicks a year” actually looks like from the inside, on a real, currently operating website, using free-tier services because that’s what’s available to an independent creator without Google’s blessing. It is the single most concrete illustration in this piece of what the statistics in Chapters 1 and 4 mean in practice for a site exactly like this one.

What the Evidence Does and Doesn’t Support

It would be easy to end this documentary with the flattened version — “big brands get trusted, small creators get called spam” — because it’s emotionally satisfying and partly true. But the record is more specific than that, and a documentary owes you the specific version:

  • What’s supported: Independent, individually-run sites (HouseFresh, Retro Dodo, the 78% of surveyed travel publishers) absorbed the deepest and most durable traffic losses from Google’s 2023–2024 quality updates. Recovery, when it happens at all, is rare and slow.
  • What’s more complicated: The single largest enforcement action against “trust abuse” specifically targeted large, trusted brand names — CNN, USA Today, Forbes, WSJ — for renting out their domain authority to third parties. Size didn’t protect them from that particular policy; it was the mechanism they exploited that got flagged.
  • What’s genuinely asymmetric: Ad spend is not a documented factor in Google’s ranking or indexing algorithms — Google states organic results and paid ads are separate systems, and no evidence in the public record contradicts that. What is asymmetric is resilience: large publishers can absorb a 60–90% traffic loss and survive on other revenue; a single-person site cannot, and that structural imbalance produces the same outcome — independent creators disappearing — without requiring any deliberate bias in the algorithm itself.

Recommendations, If You’re Building on This Land

  1. Use precise language. “Deindexed” is not “deleted from the internet.” “Terminated account” is not “your app was deleted from users’ phones.” The accurate version of your grievance is more credible, and it’s damning enough without inflation.
  2. Treat Google organic as a declining channel, not a foundation. With a documented 58% CTR decline on top rankings, if Google search sends you more than half your traffic, that concentration is now a structural risk, not just a business fact.
  3. Decide deliberately on Google-Extended, but know its limits. Blocking it opts you out of Gemini training with zero SEO cost. It will not remove you from AI Overviews. There is no setting that does both without hurting your Search visibility.
  4. If you publish apps, document everything. Compliance records, test logs, and account-separation evidence matter, because automated “association” flags are the leading cause of the wrongful terminations documented above. File any appeal well inside the 180-day window — there’s no upside to waiting.
  5. Build what a policy update can’t take. An email list, a direct audience, a product people pay for directly — none of these disappear because an algorithm changed its mind about what counts as “helpful.”

Frequently Asked Questions (FAQ)

Q: Has Google actually deleted millions of websites?

A: No — that specific framing isn’t accurate. Google has deindexed hundreds of named sites and demoted many more through 2024’s spam and quality enforcement, which can be commercially fatal, but the sites themselves remain online and reachable by direct URL. The “millions” figure is documented for apps removed from the Play Store, not websites removed from search.

Q: Does blocking Google-Extended stop my content from appearing in AI Overviews?

A: No. Google-Extended only controls Gemini/Vertex AI training data. AI Overviews are powered by ordinary Googlebot-indexed content, so the only way to suppress them is with tools like noindex or nosnippet — which also remove you from normal search results.

Q: How many apps and developer accounts did Google actually remove in 2024?

A: According to Google’s own Play Store transparency report, approximately 3.9 million apps were removed and about 155,000 developer accounts were terminated in 2024, with roughly 1,500 accounts reinstated on appeal.

Q: Is it true Google trains its AI on content publishers opted out of?

A: For AI model training specifically (Gemini/Vertex), no — Google says it honors those opt-outs. But under oath in a September 2025 antitrust trial, a Google DeepMind VP confirmed the Search organization does train on opted-out content when the use is “for search,” which includes AI Overviews.

Q: If my site is deindexed or losing traffic, is it because Google favors big brands?

A: Partly, but not simply. Independent publishers have suffered the deepest, most durable losses from quality updates. But the largest single “trust abuse” enforcement action targeted big-name brands specifically. The clearer pattern is that large publishers can survive a steep traffic loss on other revenue streams, while independent creators usually cannot — which produces the same visible outcome without requiring deliberate favoritism in the algorithm itself.

Conclusion: The Precise Version Is the Powerful Version

I started this piece wanting to say Google took over the internet, deleted millions of sites, and quietly trains its AI on work it never paid for. Having gone through the actual record — court testimony, transparency reports, independent trackers, and our own site’s real numbers — I’m not backing away from the underlying grievance. I’m sharpening it.

Google has not deleted the internet. It has built a system where a page can exist, be well-written, rank in position 4.1, and still earn 14 clicks in a year, because the doorway it depends on now keeps most of the value for itself. It has terminated hundreds of thousands of developer accounts through processes with almost no real recourse. It has, under oath, confirmed that opting out of AI training doesn’t mean what most publishers assumed it meant. And it holds a 90% share of the only map most people use to find anything at all.

None of that requires exaggeration. Stated precisely, it’s already the story. This is what building on Google’s land actually looks like right now, documented rather than dramatized — including on this exact website, with its exact numbers, on the exact day this was written.

📚 Free Download: The Complete Documentary Sources & Data Guide

Keep every statistic, court quote, and source from this documentary for your own reference.

Download Free PDF Book

Sources: Google Play Annual Transparency Report (via Surfshark, 2025); Appfigures/TechCrunch Play Store catalog analysis; StatCounter Global Stats (May 2026); Ahrefs click-through-rate study (Feb 2026, 300,000 keywords); Pew Research Center AI Overview study (July 2025); US v. Google antitrust remedies trial testimony (September 2025); Sistrix visibility tracking; Digitaloft travel-publisher study; Google Search Central documentation on Google-Extended; this site’s own Google Search Console data.

Disclaimer: This article synthesizes publicly reported statistics, court testimony, and independent research current as of August 2026, alongside our own site’s real analytics data. Some figures (site reputation abuse enforcement, ongoing EU investigation, Chegg/Penske litigation) describe matters still unresolved at time of writing and may change.