Credible Roots / Field NotesWikipedia Page CreationKnowledge PanelWikidata ProfileAI Search VisibilityPersonal Branding: The GuideDistributionAI VideoDistributionPersonal Brand WebsiteWhat It CostsNow or LaterCase StudyAboutEditorial StandardsCorrectionsAuthorContactPrivacy

Field note

How to get cited by ChatGPT, Perplexity and Gemini

Answer engines cite what they can retrieve, verify and quote. That makes citation less a marketing problem than a question of whether clear, corroborated statements about you exist in public.

Atharv Sankpal

By Atharv Sankpal

More than five years in personal branding. Runs atharvsankpal.com and works with founders and businesses in the US on their digital presence. Previously produced AI video for brands.

Published 2026-08-13 · 7 min read

Researched against primary sources, drafted with AI assistance, then reviewed and approved by the named author. How we write · About the author

The short answer

AI answer engines mostly compose answers from pages they retrieve at the moment someone asks, which makes citation a question of being findable, quotable and corroborated rather than a ranking exercise. Three things decide it. First, access: if your robots.txt blocks GPTBot, ClaudeBot, PerplexityBot or Google-Extended, those systems cannot retrieve your pages to cite them at all. Second, the shape of your writing: a passage that states a complete, specific fact near the top of a page can be lifted and attributed, whereas positioning language gives an extractor nothing to quote. Names, dates, numbers and defined subjects survive extraction; adjectives do not. Third, corroboration: when a claim about you appears only on your own site, a careful system treats it as a claim, but when independent sources repeat the same fact, the system will assert it.

How answers actually get sourced

When one of these tools answers a question with links, it is usually not reciting memorised text. It is searching, retrieving a set of pages, and composing an answer grounded in what it just read.

That mechanism explains most of what follows. Being cited depends on being retrievable at the moment someone asks, and on being easy to quote once retrieved. It is closer to being a good source for a journalist on deadline than to traditional ranking.

It also explains why obscure pages sometimes get cited over famous ones. If your page answers the specific question asked, more directly than a larger site does, it can win the citation without ranking first for anything.

How a citation actually happens

01
Someone asks
“Who is the expert on X?”
02
The engine retrieves
Pages it is allowed to crawl.
03
It looks for a passage
One that stands on its own.
04
It quotes and cites
The clearest source wins.

Blocked crawlers drop you out at step two. Vague positioning copy drops you out at step three.

Make the passage liftable

The single highest leverage change is structural. Put the answer first, in a complete sentence that makes sense with nothing around it.

Compare two versions of the same fact. "With over two decades of experience across multiple sectors, our approach has always been client focused" gives an engine nothing to quote. "Priya Raman has advised manufacturers on supply chain resilience since 2004, and led the response study cited by the industry association in 2023" can be lifted verbatim and attributed.

The pattern that works: a direct claim near the top of the page, short paragraphs, subheadings phrased as the questions people ask, and specifics rather than adjectives. Names, dates, numbers and defined subjects survive extraction. Positioning language does not.

Let the right crawlers in

This is mechanical and often overlooked. Different systems use different crawlers, and your robots.txt decides which ones may fetch your pages. If you want to be citable, the relevant agents need to be allowed.

It is worth checking rather than assuming, because plenty of sites block these agents by default through a host setting or a security product, then wonder why they never appear in AI answers. There is a legitimate opposite choice too: some publishers block training crawlers deliberately to protect licensing. Just make it a decision rather than an accident.

The same applies to rendering. If your content only exists after JavaScript runs, some retrieval systems will see an empty page. Server rendered text is safer.

Corroboration beats optimization

Everything above is necessary and none of it is sufficient, because these systems weigh agreement across sources.

If the only place a claim about you appears is your own site, a careful system treats it as a claim. If the same fact appears in a trade publication, a conference program and a university page, it becomes something the system will assert. This is the same logic behind Wikipedia's sourcing rules and Google's entity confidence, arriving from a different direction.

Which means the work is not really technical. Publish clear material on a defined subject, be the person others reference when discussing it, and let independent sources repeat the facts. The formatting makes you quotable. The corroboration makes you trusted.

Questions people actually ask

How do AI tools decide who to cite?

Answer engines retrieve pages relevant to the question, then quote the ones that state something clearly, factually and in a self-contained way. Pages that make a direct claim near the top, and that are corroborated elsewhere, are far easier to cite than pages that bury the point.

Does blocking AI crawlers hurt your visibility?

If you block a crawler in robots.txt, that system cannot retrieve your pages to cite them. Some organizations block deliberately for licensing reasons. If being cited is the goal, blocking works directly against it.

Is llms.txt required to be cited by AI?

No. It is a proposed convention, not an adopted standard, and no major answer engine requires it. It is cheap to add and may help, but crawler access, clear factual writing and independent corroboration matter far more.

Why does ChatGPT mention competitors and not you?

Usually because there is more retrievable, quotable material about them, spread across sources the system trusts. The fix is not a trick. It is producing clear material on a specific subject and being referenced by others discussing it.

What the citation studies actually find

This area is full of confident numbers that disagree with each other, so it is worth separating what is consistent from what is contested.

Published research

FindingFigureSource
AI citations that come from earned media rather than a brand's own siteabout 84%Muck Rack, May 2026
A brand's own website's share of what AI engines reference5 to 10%Muck Rack, 2026
Brands more likely to be cited via third-party sources than their own domain6.5xAirOps, October 2025
LLM citations drawn from the first 30% of an article44.2%Kevin Indig, Growth Memo, 2026

Contested: the overlap between AI citations and Google's top 10 is reported anywhere from about 12% to 38% depending on the study and the engine. The studies disagree on the number and agree on the conclusion, which is that ranking well is not sufficient to be cited. One honest caveat: AI referral traffic is still roughly 1% of total web traffic (Search Engine Land analysis of 3.3 billion sessions), so this is currently high-quality and low-volume.

References

How this was made: researched against primary sources, drafted with AI assistance, then reviewed and approved by the named author before publication. Our editorial standards.

Not ready for a call? Being recognized as an entity comes first. The five steps are here, and none of them cost anything.

Start with your entity →

Want to see what AI says about you now?

We will run the questions your buyers ask, show you who gets named instead of you, and map what it would take to change the answer.

Apply for a call