Skip to content
Implementa.

SEO for ChatGPT (GEO) · Guide 35 of 35

What sources does AI cite: the map of domains that decides whether it names you

When ChatGPT recommends a tool or names a vendor, it isn’t reading your website: it’s synthesizing from a very small set of sources it weights above the rest. And that set is shorter than you’d think. In an analysis of 1.2 million ChatGPT responses, 67% of the citations on a topic came from just 30 domains; and in a 5W Public Relations study of U.S. citations, Wikipedia (≈13%) and Reddit (≈12%) alone account for more than one in four. Translation: there’s a club of sources that decides who exists for the model, and your website is rarely a member. This is the map of that club —which domains AI weights when it talks about your category, how to spot them, and what to do to get inside—.

There’s an idea that’s hard to let go of: that if your website is good, AI will find you. It doesn’t work that way. When ChatGPT recommends a tool or names a vendor, it doesn’t crawl the internet reading sites one by one: it retrieves and synthesizes from a very small set of sources it weights above the rest. And that set is genuinely short. In an analysis of 1.2 million ChatGPT responses, 67% of the citations on a topic came from just 30 domains. Existing for the model isn’t showing up on Google: it’s being inside that club of sources. This is the map of who runs it and how to get in.

AI doesn’t read you: it retrieves from a club of sources

A generative engine doesn’t rank pages by links the way classic search does. When you ask it, it breaks the question down, retrieves fragments from the sources it trusts, and writes an answer citing a few of them. The effect is brutal concentration: in the 5W Public Relations study of ChatGPT citations in the U.S., Wikipedia (around 13%) and Reddit (close to 12%) alone add up to more than one in four citations, and outside those two almost no domain passes 3%. The list of who decides is short, and your domain is almost never on it.

This changes where your visibility is decided. According to industry analyses, most of the citations AI uses to build a recommendation don’t come from the brand’s domain but from third parties —directories, reviews, media, forums—. Your site is still necessary (AI retrieves it to confirm what you do and who you are), but it rarely gets you onto the list. The list is written by the sources the model already weights. That’s why "optimizing for GEO" isn’t just touching your site: it’s understanding which sources AI cites in your category and working to be in them.

The domains that almost always carry weight (and why)

No two categories are identical, but the pattern repeats. Five source types show up again and again when AI talks about products, services or vendors:

Source typeExamplesWhy AI weights it
EncyclopedicWikipediaVerifiable, neutral entity. It’s the most-cited source by ChatGPT: it turns you into "something that exists" in the model’s eyes.
CommunityReddit, niche forumsOpinion with no brand filter. AI reads it as real user experience, not marketing.
Reviews / directoriesG2, Capterra, your sector’s directoriesIndependent verification of what you say about yourself. In software, G2 is among the most cited precisely because you don’t control it.
Authority mediaForbes, Business Insider, trade pressCredible third-party signal. A signed mention weighs more than ten posts on your own blog.
MultimodalYouTubeDominates in Google AI Overviews and adds up in the rest. Well-tagged video gets in where text can’t.

The common thread is uncomfortable but clear: AI trusts what others say about you more than what you say about yourself. Independent verification beats self-reported information. That’s why a real review on G2 or a mention in a media outlet moves your visibility needle far more than rewriting your homepage for the tenth time.

Each engine reads a different internet

The rookie mistake is treating "AI" as one thing. It isn’t. Domain overlap between ChatGPT and Perplexity is only around 11%: they’re two almost separate libraries. Perplexity leans heavily on Reddit and live sources; Google AI Overviews favor YouTube and multimodal content; ChatGPT concentrates on Wikipedia, Reddit and established media. The same brand can be queen in one engine and not exist in another. We unpack it in ChatGPT vs Perplexity: same brand, two visibilities and in how to appear in Perplexity.

The practical consequence: a source audit is done per engine. Measuring your presence in a single "AI" aggregate gives you a number that averages opposite realities and leaves you blind exactly where you have the gap. If your buyer asks in Perplexity, your source map is Perplexity’s, not ChatGPT’s.

How to find out which sources it cites in YOUR category

The good news: you can audit this without a paid tool to start. You just need method and consistency. Four steps:

  1. Write 15-25 category questions, without your brand. How a buyer who doesn’t know you yet actually asks ("best tool for X", "how to choose a provider of Y"). It’s the same set you use to audit your visibility with prompts.
  2. Fire them at the engines that show their sources. Perplexity, Google AI Overviews and ChatGPT with search on list the domains they cite. That’s where you read the map; note every source that shows up.
  3. Count repetitions, not impressions. After two or three rounds you’ll see the same ones repeat: a directory or review platform in your sector, one or two media outlets, a Reddit thread, one or two competitor sites. That short list is the club that decides your category.
  4. Cross it with where you are. Mark which of those sources you already appear in and which you don’t. The gap between "sources AI cites" and "sources you exist in" is, literally, your work plan.

How to get into the set of sources that cites you

Once you have the map, the job is getting into those sources, not inventing others. In order of return:

  1. Claim and complete your listing on the review platforms and directories that already show up. G2, Capterra or your sector’s vertical, with correct data and real, recent reviews. It’s the fastest lever in software categories.
  2. Consolidate your entity. Same description, same category and same data across your site, LinkedIn and directories, so AI can "square" who you are. The structured data that does matter helps with that disambiguation, but it’s plumbing, not the engine.
  3. Earn third-party mentions. A signed article in a trade outlet, a podcast appearance, a published case. AI weights what others say about you; give it verifiable material to cite.
  4. Participate where the community lives. Reddit and niche forums carry a lot of weight. Don’t spam: contribute useful answers AI can retrieve as real experience.
  5. Keep your site clean and legible. Direct answer up top, verifiable data, clear entity. It won’t get you on the list by itself, but it’s what AI retrieves to confirm you once you’re there.

And measure it over time, not once: today’s snapshot doesn’t tell you if you’re improving. Combine the visibility audit for the initial map, tracking mentions week by week to see the movement, and the framework for how to measure your AI visibility to score each answer. If you don’t want to build and maintain this by hand, AI visibility monitoring keeps the source map live and watched per engine; and when the reading calls for action, GEO optimization works on your presence in those sources, not on a report.

Frequently asked questions

A very concentrated handful. In 5W’s analysis of ChatGPT citations in the U.S., Wikipedia (around 13%) and Reddit (close to 12%) add up to more than 25% of all citations, and outside those two almost no domain passes 3%. After them, authority media (Forbes, Business Insider), YouTube and —in software categories— review platforms like G2 show up repeatedly. The lesson isn’t "publish on Wikipedia": it’s that ChatGPT builds the answer from third-party sources it trusts, not from your domain. That’s why visibility work happens as much off your site as on it.

No, and the gap is big. Domain overlap between ChatGPT and Perplexity is only around 11%: each engine reads a different internet. Perplexity leans heavily on Reddit and live, fresh sources; Google AI Overviews favor YouTube and multimodal content; ChatGPT concentrates on Wikipedia, Reddit and established media. The same brand can dominate one and vanish in another. That’s why a source audit is done per engine, not in a single aggregate that reassures you and lies about where you actually stand.

By asking it yourself, methodically. Fire 15-25 category questions (without your brand) at each engine and, in the ones that show their sources —Perplexity, Google AI Overviews, ChatGPT with search— note which domains show up again and again. In two or three rounds you’ll see the pattern: it’s almost always a directory or review platform in your sector, one or two median outlets, a Reddit thread and one or two competitor sites. That short list is your map: those are the pages you need to exist on for AI to retrieve you. You don’t need a paid tool to start; you need to ask the same thing every time and take notes.

Yes, but with the right role. Your site is necessary —AI retrieves it to confirm what you do, your data and your entity— but it rarely gets you onto the recommended list: that’s decided by the third-party sources the model weights. So the work has two legs at once: making your site clean and legible (direct answer, verifiable data, consistent entity) and, above all, earning cited presence on the domains that already show up in your category —real reviews on G2/Capterra or your sector’s directory, signed mentions in media, useful participation in the communities AI reads—. One leg without the other leaves half the job undone.

Free AI Impact Plan

The guide is generic. Your plan isn't.

Tell us about your company and we'll ship back a diagnosis with priorities, numbers and what to implement first. No sales call, no charge.

What sources does AI cite: the map of domains that decides whether it names you · Implementa