Skip to content
← Blog
Practice8 min

How to get cited by ChatGPT

ChatGPT cites the sources it retrieved while composing an answer, and those are frequently not the pages ranking for the original query. Earning a citation means being present in the material it retrieves: third-party comparisons and reviews, documentation it can quote cleanly, and pages structured so a single passage answers a sub-question completely.

Alex ChenPractice, Volexi
Published 16 July 2026

How does ChatGPT decide what to cite?

It expands your question into sub-queries, retrieves candidate sources for each, then writes one answer from the sources that survived. Citations come from that retrieved set, not from a ranking.

This is the single most useful thing to understand, because it explains an outcome that otherwise looks arbitrary: a page can hold position one and not be cited in an answer to the query it ranks for, while a forum thread nobody optimised gets quoted.

It also explains why the fix is often not on your website. If the retrieved set for your category is dominated by review sites and comparison roundups, publishing another page on your own domain does not enter the competition.

Why do third-party pages beat my own pages?

Because a model asked to recommend a tool treats an independent comparison as better evidence than a vendor claiming to be the best. Your own page is a source about you, not a source about the category.

The asymmetry is worth accepting rather than fighting. Vendor pages do get cited, but usually for factual sub-questions: pricing, specifications, supported integrations, compliance. Recommendation questions go to third parties almost by default.

That splits the work in two. Own the factual questions with clear pages on your domain, and earn the recommendation questions by changing what independent sources say, which is outreach, reviews and community presence rather than publishing.

What earns a citation

Ordered by how often each one changes an outcome, based on what shows up in the citation sets we capture.

01
Presence in independent comparisons

Roundups and alternatives pages carry recommendation answers. If the five roundups ranking for your category do not list you, you are absent from the retrieved set before writing begins.

02
One page that answers the whole fan-out

A question expands into sub-questions about price, compliance, migration, integrations and fit. A page covering all of them in one document outperforms five pages each covering one, because a single retrieval satisfies more of the answer.

03
Extractable structure

Feature tables, definition lists and short question-shaped headings can be lifted whole. Long narrative paragraphs have to be summarised, and summarising loses the attribution.

04
Specific, checkable facts

Numbers, dates, versions and named entities give a model something to quote with confidence. Vague claims get paraphrased into someone else's sentence.

05
Documentation quality

Public docs are heavily retrieved for factual sub-questions and are usually the most neglected surface in a marketing programme.

Does adding llms.txt get me cited?

On its own, no. Support is not guaranteed by any major engine, so treat it as cheap insurance rather than a lever.

The reason to publish one anyway is that writing it forces a decision about which twenty pages actually represent you. That decision is usually worth more than the file.

The measurement problem underneath all of this

Every recommendation above is testable, and almost nobody tests them, because a single check tells you nothing. Answers vary run to run, so the same prompt asked twice in an afternoon can name different brands in a different order.

Anything you change has to be measured as a rate across repeated runs and multiple engines, before and after. Otherwise you are reading noise and attributing it to the work you just did.

Common questions

Does ChatGPT cite the same sources as Google AI Overviews?

Frequently not. They retrieve from different pools, which is why a brand can be well represented in one and absent from the other. Checking a single surface gives a misleading read.

Will blocking GPTBot protect my content?

It will also remove you from the retrieved set, so you cannot be cited. That is a legitimate choice for a publisher selling subscriptions and a poor one for a company that wants to be recommended.

How many prompts do I need to track to get a stable read?

Enough to cover the sub-questions in your category rather than a headline term or two, and repeated on a schedule. A stable read comes from repetition over time, not from a larger single sample.

Alex Chen
Practice, Volexi

Covers the work that changes whether an engine names you.

volexi
MONITOR · ACT · MEASURE
The AI search console for marketing teams and the agencies that serve them.
© 2026 JOLTCLICK LIMITED T/A VOLEXIGDPR-READY · EU DATA RESIDENCY