Google’s Biggest SEO Leak Changed Everything… But Almost Nobody Adapted
Guest: Mike King / iPullRank (NEW) · Host: Edward Sturm (NEW) | Published: 2026-08-05 | Source: transcript (64:01, 6.8K views)
Summary
Edward Sturm (episode 1,127) interviews Mike King, one of the first public analysts of Google’s May 2024 Content Warehouse API leak (NEW: google-api-leak / content-warehouse). King is explicit that the leak plus DOJ testimony, Mark Williams-Cook’s SERP JSON leak, and the earlier Yandex source leak “confirmed everything” SEOs suspected Google misleads about — and that the SEO-software industry still has not adapted. This page separates what he attributes to the leaks from what is his later interpretation or agency testing.
Leak-attributed (King’s reading of Content Warehouse / related dumps): Google does have a site-level authority concept despite public denials of “domain authority.” The index is stratified into four buckets — high, medium, low, plus fresh docs — and link equity is a sliding scale by bucket, not a uniform PageRank pipe. Traffic and first-page ranking are the practical proxy for which bucket a URL lives in; editorial/news sites sit in the high bucket (not merely because of crawl rate). DR/DA are “entertainment metrics.” Vector embeddings represent people, sites, entities, pages, and passages and are rolled up at multiple levels — which, for King, means the link graph is less central than the industry still pretends. Twiddlers (boosts/demotions after initial scoring; Panda started as one) appear in both Yandex code and Google’s systems. User-behavior signals in the leak data: a temporary site-quality score cloned from a similar known site, then dwell / bounce from search used to keep or demote it. Generative-AI farms typically fail that human-signal test, not an “AI detector.”
King’s interpretation (not a leaked field): high-bucket docs in RAM, mid on SSD, junk on spinning disk; mentions usually overpower the link graph because they create consistent entity context (Rand Fishkin learned this from King), but it is not a hard rule — some query spaces still let links win; “make friends or make news,” not guest-post spam.
Post-leak GEO practice (agency tests, not the 2024 dump): ChatGPT retrieval is real-time, not an index — it uses Bing, Google, Exa, SerpAPI (visible in the conversation network response). Optimize query fan-outs, not the prompt; an agentic critic drops passages at each stage. A client drowning in nginx 499s (client abandoned the request) gained ~300% AI visibility in three months after edge-caching for TTFB — the site did not feel slow to humans (SPA illusion). Training-data presence is audited via Common Crawl; they try to add crawl paths from pages already in CC (including Wikipedia). He has no 1-2-3 “get into the weights” recipe and is still figuring it out. Old tactics returning for AI: white-on-white text, cloaking, micro-sites — “no rules on these platforms.” Most-cited surfaces: Reddit, YouTube, LinkedIn; he tells clients to start their own subreddit rather than astroturf. llms.txt: he used to agree it was theater; Claude uses it heavily, and Lighthouse’s agentic review looks at it — Google not using it for search is not the same as agents ignoring it.
He rebrands the job relevance engineering (AI + IR + content strategy + UX + digital PR), insists AI search is not “just SEO” (his unpopular opinion), and would have every SEO learn to vibe-code. 90-day action: grep logs for 499s and fix speed.
Key Claims
- Leak-confirmed (King): site-level authority exists; index in four buckets (high/mid/low/fresh docs); link value depends on the source’s bucket; traffic/ranking is the proxy; embeddings everywhere; Twiddlers/boosts; search-session quality (dwell/bounce) feeds site quality. (leak-reading, not a primary dump on this page)
- DR/DA are entertainment metrics; SEO tools have not rebuilt around embeddings or bucketed link value. (King-opinion)
- Mentions usually beat links for entity understanding because of consistency across the web; exceptions exist. (interpretation; Rand attributed this to King)
- ChatGPT RAG is live-fetch, not indexed; 499 timeouts silently drop you. Edge-cache / TTFB fix → ~300% visibility / 3 months. (agency test)
- Optimize fan-out queries + passage relevance, not prompts. Fan-outs jitter run-to-run but the topic is stable. Comprehensive audience-driven content, not a million near-dupes.
- AI referral is a rounding error that converts; treat the channel as brand visibility (citation rate + accuracy) plus input metrics (bots, synthetic-query ranks, passage scores, entity salience).
- Common Crawl presence is his current training-data proxy; Wikipedia/in-CC links as crawl paths. No scientific implant method. Spam “millions of AI mentions” is the old play returning.
- Claude uses
llms.txt; Google search does not have to. Ahrefs’ “unsupported” claim is incomplete if agents are in scope. - AI-scaled content is not automatically spam; Google cannot reliably detect it, so it uses human signals. Watermarking is the detection bet.
Notable quotes
“They’d say things like, ‘Oh, there’s no domain authority’… and there’s literally something in there called site authority.”
Why citable: The leak’s most famous contradiction of Google’s public line.
“We’ve historically believed that whether or not a site gets traffic or if the site ranks or not doesn’t matter as far as link building, but we found that it’s definitively true that that matters.”
Why citable: Directly upgrades Brand & User Signals as Ranking Drivers / Dooley’s “traffic is #1” from creator-opinion toward leak-informed.
“When AI has to choose… ChatGPT will give up… they had a ton of 499s. And by fixing that one thing, we improved visibility in that 3 months by like 300%.”
Why citable: Concrete, log-based GEO tactic that matches Ahrefs’ “slow pages get dropped from RAG” without being an Ahrefs talking point.
Connections
Entities mentioned: Edward Sturm (NEW), Mike King / iPullRank (NEW), Google, ChatGPT, Claude, Reddit, Perplexity (implicit via GEO), Rand Fishkin, Mark Williams-Cook, Yandex, Bing, Brave, Profound, Qoria (King’s open-source fan-out tool) Concepts referenced: google-api-leak / content-warehouse (NEW), GEO / AEO (Getting Recommended by AI), Entity-Based SEO, Brand & User Signals as Ranking Drivers, Local Link Building & Authority, Data Poisoning & LLM Backdoors, 250 Authority Protocol, AI-Assisted Development, Schema Markup (Structured Data) (not a focus)
Contradictions / Tensions
- Leak vs. speculation (this page’s job): bucketed index, site authority, embeddings, Twiddlers, user-session quality → attributed to leaks. RAM/SSD/HDD storage map, “mentions > links” as a rule, 499→300%, Common Crawl implant, white-on-white working, own-your-subreddit → King interpretation or tests. Do not launder the second list into “the leak proved.”
- vs. AI SEO Course for Beginners: Complete AEO Tutorial (Ahrefs) on
llms.txt: Ahrefs = no major provider officially supports it. King = Claude uses it; Lighthouse agentic review reads it. Record both; scope is search vs. agents. - vs. AI SEO Course for Beginners: Complete AEO Tutorial (Ahrefs) on mechanism: both say mentions + retrieval + speed; King adds leak-grounded index buckets and user-signal scoring. Ahrefs has the 75k-brand correlations King lacks here.
- vs. 250 Authority Protocol: King will not claim a document-count implant. His training-data move is get into Common Crawl via real crawl paths. Closer to Anthropic’s “access, not count” caveat.
- vs. Caleb Ulku on age/signals: leak-informed user-session quality supports Dooley’s traffic thesis more than Caleb’s “models have no sense of age.”
- vs. Joy/Caleb on Reddit: they treat Reddit as scarce human content; King treats it as hostile and recommends a brand-owned subreddit plus controlled replies — a different operationalization of the same “Reddit is cited” fact.
- Commercial bias: King is an agency owner who withholds “magic metrics” because secrets don’t help an agency the way they helped Moz. Treat unpublished “if we improve X, visibility moves” as unaudited. Host sells Compact Keywords (ignore).
Notes
Use this as the wiki’s leak-interpreter source, not as a substitute for the primary Content Warehouse docs. Highest-value transfers: site authority exists; links are bucket-weighted; embeddings/entities are first-class; user signals are in the machinery; 499/TTFB is a GEO bug. Flag NEW: edward-sturm, mike-king, google-api-leak / content-warehouse.