Credible Roots / Field NotesWikipedia Page CreationKnowledge PanelWikidata ProfileAI Search VisibilityPersonal Branding: The GuideDistributionAI VideoDistributionPersonal Brand WebsiteWhat It CostsNow or LaterCase StudyAboutEditorial StandardsCorrectionsAuthorContactPrivacy
Field note
How to get cited by ChatGPT, Perplexity and Gemini
Answer engines cite what they can retrieve, verify and quote. That makes citation less a marketing problem than a question of whether clear, corroborated statements about you exist in public.
AI answer engines mostly compose answers from pages they retrieve at the moment someone asks, which makes citation a question of being findable, quotable and corroborated rather than a ranking exercise. Three things decide it. First, access: if your robots.txt blocks GPTBot, ClaudeBot, PerplexityBot or Google-Extended, those systems cannot retrieve your pages to cite them at all. Second, the shape of your writing: a passage that states a complete, specific fact near the top of a page can be lifted and attributed, whereas positioning language gives an extractor nothing to quote. Names, dates, numbers and defined subjects survive extraction; adjectives do not. Third, corroboration: when a claim about you appears only on your own site, a careful system treats it as a claim, but when independent sources repeat the same fact, the system will assert it.
- Access first. If your robots.txt blocks a system, it cannot cite you.
- Write in liftable passages. Lead with the answer, then explain.
- Corroboration decides ties. Systems trust facts that appear in several independent places.
- Be specific about something. Broad positioning gives an engine no reason to name you.
How answers actually get sourced
When one of these tools answers a question with links, it is usually not reciting memorised text. It is searching, retrieving a set of pages, and composing an answer grounded in what it just read.
That mechanism explains most of what follows. Being cited depends on being retrievable at the moment someone asks, and on being easy to quote once retrieved. It is closer to being a good source for a journalist on deadline than to traditional ranking.
It also explains why obscure pages sometimes get cited over famous ones. If your page answers the specific question asked, more directly than a larger site does, it can win the citation without ranking first for anything.
How a citation actually happens
Blocked crawlers drop you out at step two. Vague positioning copy drops you out at step three.
Make the passage liftable
The single highest leverage change is structural. Put the answer first, in a complete sentence that makes sense with nothing around it.
Compare two versions of the same fact. "With over two decades of experience across multiple sectors, our approach has always been client focused" gives an engine nothing to quote. "Priya Raman has advised manufacturers on supply chain resilience since 2004, and led the response study cited by the industry association in 2023" can be lifted verbatim and attributed.
The pattern that works: a direct claim near the top of the page, short paragraphs, subheadings phrased as the questions people ask, and specifics rather than adjectives. Names, dates, numbers and defined subjects survive extraction. Positioning language does not.
Let the right crawlers in
This is mechanical and often overlooked. Different systems use different crawlers, and your robots.txt decides which ones may fetch your pages. If you want to be citable, the relevant agents need to be allowed.
It is worth checking rather than assuming, because plenty of sites block these agents by default through a host setting or a security product, then wonder why they never appear in AI answers. There is a legitimate opposite choice too: some publishers block training crawlers deliberately to protect licensing. Just make it a decision rather than an accident.
The same applies to rendering. If your content only exists after JavaScript runs, some retrieval systems will see an empty page. Server rendered text is safer.
Corroboration beats optimization
Everything above is necessary and none of it is sufficient, because these systems weigh agreement across sources.
If the only place a claim about you appears is your own site, a careful system treats it as a claim. If the same fact appears in a trade publication, a conference program and a university page, it becomes something the system will assert. This is the same logic behind Wikipedia's sourcing rules and Google's entity confidence, arriving from a different direction.
Which means the work is not really technical. Publish clear material on a defined subject, be the person others reference when discussing it, and let independent sources repeat the facts. The formatting makes you quotable. The corroboration makes you trusted.
Questions people actually ask
How do AI tools decide who to cite?
Answer engines retrieve pages relevant to the question, then quote the ones that state something clearly, factually and in a self-contained way. Pages that make a direct claim near the top, and that are corroborated elsewhere, are far easier to cite than pages that bury the point.
Does blocking AI crawlers hurt your visibility?
If you block a crawler in robots.txt, that system cannot retrieve your pages to cite them. Some organizations block deliberately for licensing reasons. If being cited is the goal, blocking works directly against it.
Is llms.txt required to be cited by AI?
No. It is a proposed convention, not an adopted standard, and no major answer engine requires it. It is cheap to add and may help, but crawler access, clear factual writing and independent corroboration matter far more.
Why does ChatGPT mention competitors and not you?
Usually because there is more retrievable, quotable material about them, spread across sources the system trusts. The fix is not a trick. It is producing clear material on a specific subject and being referenced by others discussing it.
What the citation studies actually find
This area is full of confident numbers that disagree with each other, so it is worth separating what is consistent from what is contested.
Published research
| Finding | Figure | Source |
|---|---|---|
| AI citations that come from earned media rather than a brand's own site | about 84% | Muck Rack, May 2026 |
| A brand's own website's share of what AI engines reference | 5 to 10% | Muck Rack, 2026 |
| Brands more likely to be cited via third-party sources than their own domain | 6.5x | AirOps, October 2025 |
| LLM citations drawn from the first 30% of an article | 44.2% | Kevin Indig, Growth Memo, 2026 |
Contested: the overlap between AI citations and Google's top 10 is reported anywhere from about 12% to 38% depending on the study and the engine. The studies disagree on the number and agree on the conclusion, which is that ranking well is not sufficient to be cited. One honest caveat: AI referral traffic is still roughly 1% of total web traffic (Search Engine Land analysis of 3.3 billion sessions), so this is currently high-quality and low-volume.
References
- OpenAI. Overview of OpenAI crawlers — GPTBot, OAI-SearchBot and ChatGPT-User.
- Anthropic. Crawler documentation — ClaudeBot, Claude-User and Claude-SearchBot.
- Google Search Central. Crawlers and user agents, including the Google-Extended control.
- RFC 9309. Robots Exclusion Protocol.
- Muck Rack, 2026, on the share of AI citations sourced from earned media. A vendor study, treated here as indicative rather than definitive.
- AirOps, 2025, on third-party versus owned-domain citation likelihood. A vendor study, treated here as indicative.
- Kevin Indig, Growth Memo, 2026, on where in an article citations are drawn from. A practitioner analysis, not controlled research.
How this was made: researched against primary sources, drafted with AI assistance, then reviewed and approved by the named author before publication. Our editorial standards.
Related reading
Not ready for a call? Being recognized as an entity comes first. The five steps are here, and none of them cost anything.
Start with your entity →Want to see what AI says about you now?
We will run the questions your buyers ask, show you who gets named instead of you, and map what it would take to change the answer.
Apply for a call