TRAIN Any AI To Recommend You (This Study Proves How)

Author: Caleb Ulku | Published: 2026-06-12 | Source: transcript (17:52, 9.9K views at ingest)


Summary

This is the corpus’s core AEO (getting recommended by AI models) source. Caleb’s premise: when someone asks ChatGPT, Claude, Gemini, or Grok to recommend a local business, the model has no internal database — it either searches the web in real time and builds an answer from whatever it finds (Reddit, forums, Medium, review sites), or, for well-known entities, answers from training memory. Critically, models have no sense of content age — a 2019 complaint weighs the same as a fresh five-star review. Most local businesses are invisible to AI or misrepresented by content they don’t control.

He anchors the strategy in an Anthropic study (with the UK AI Security Institute and the Alan Turing Institute) showing that ~250 documents (≈0.00016% of training data) were enough to establish a behavioral pattern in models from 600M to 13B parameters — and, surprisingly, the count didn’t need to scale with model size (“the defenses don’t scale”). He honestly caveats that the study tested a poisoning trigger, not brand influence, and that quality filters/deduplication are real barriers — but argues the principle (low threshold to establish a pattern) transfers.

Two deliverables: (1) a “Model Training Data Risk Auditor” prompt that forces an AI to surface what Reddit/forums/reviews currently say about a brand (defense); and (2) the 250 Authority Protocol (offense) — 250 diverse pieces of content across four buckets (own site/case studies; professional platforms like LinkedIn; community content like Reddit/Quora; third-party validation like press, chambers, sponsorships). The mechanism is consensus: AI trusts information that recurs across multiple trusted sources, formats, and perspectives, and discounts echo chambers (250 identical posts on one domain = one source). Production uses a multi-step Claude pipeline (topical map → outline grounded in real questions → long writing prompt → human editor to catch hallucinations), ~30–40 pieces in a few hours, full protocol in 1–2 months. Framing: stop selling “SEO,” start selling “AI presence” while the window is open.


Key Claims

  • AI models building local recommendations search the web in real time unless the entity is well-known enough to be in training memory — the goal is to move from “looked up” to “recommended from memory.”
  • Models pull heavily from Reddit, forums, Medium, review sites and have no concept of time; stale negative content persists.
  • Anthropic / UK AISI / Alan Turing Institute study: ~250 documents (~0.00016% of data) established a pattern across 600M–13B-parameter models; the threshold did not scale with model size.
  • Honest caveat (his own): the study tested a specific trigger mechanism, not brand marketing; quality filters and deduplication are real barriers between publishing and training influence.
  • 250 Authority Protocol: 250 pieces across 4 buckets (own content, professional platforms, community content, third-party validation); AI rewards consensus + diversity, discounts single-source echo chambers.
  • Claude preferred over ChatGPT for content (claimed higher helpfulness, more natural, passes AI detection better) — his agency’s stated default.
  • Human review is mandatory — AI hallucinates numbers/regulations, dangerous for high-ticket/legal/medical.
  • Positioning shift: sell “AI presence,” not just SEO; “AI companies are 3 years into evaluating content quality; the window won’t stay open.”

Notable quotes

“When an AI is deciding what to believe about a topic, it’s not looking at one page, it’s looking for consensus. Does the same information show up across multiple trusted sources in multiple formats from multiple perspectives?”

Why citable: Names the core GEO mechanism — LLM recommendation as consensus-detection across sources — which is what distinguishes GEO/AEO strategy from page-level Google SEO.

“You’re not just going to be selling SEO anymore. You’re going to be selling AI presence.”

Why citable: The strategic thesis of the whole AI-search thread; useful content-fuel framing for a positioning piece.


Connections

Entities mentioned: Caleb Ulku, Anthropic, Claude, ChatGPT, Reddit, Google · Gemini, Grok, UK AI Security Institute, Alan Turing Institute, LinkedIn, Quora, Medium (minor) Concepts referenced: GEO / AEO (Getting Recommended by AI), 250 Authority Protocol, Local SEO, Entity-Based SEO


Contradictions / Tensions

  • Overreach risk (named): extrapolating a data-poisoning trigger study into a brand-marketing playbook is a leap; Caleb partially acknowledges this, but the “250 pieces → into ChatGPT’s training data” implication is stronger than the evidence supports. Tag 250 Authority Protocol as creator-opinion.
  • Consistent internally with the rest of the corpus (consensus, Reddit emphasis, entity thinking) but this is the most speculative/emerging source — consensus: emerging on GEO / AEO (Getting Recommended by AI).
  • The Anthropic study itself is a credible primary result; the marketing application is the contested part. Keep the two separate.

Notes

Backfill candidate for a direct read of the underlying Anthropic + UK AISI + Alan Turing Institute paper to verify the “250 documents” figure and its scope before treating it as established.