Logo
SEO & Pricing

How ChatGPT Chooses Which Websites to Cite

How ChatGPT chooses which websites to cite
Published on: 14/07/2026
Modified: 21/08/2026

ChatGPT does not use a single public ranking where websites keep fixed positions. When a question would benefit from current or additional information from the web, ChatGPT Search may rewrite it as one or more targeted queries, retrieve results from third-party search providers and partner content, and build an answer that links to selected pages.

OpenAI says search results are ranked using multiple factors intended to surface relevant, reliable information, but it does not publish the full ranking model, exact weights, or a formula that determines which URL will be cited. No one can reliably promise a citation for a specific result. What you can do is confirm that the site is eligible for Search, that one clear page is the best match for the topic, and that the content gives a direct, current, and verifiable answer.

The key is to separate four processes that are often treated as if they were the same:

  1. information the model already learned during training;
  2. whether ChatGPT runs a web search;
  3. which pages are discovered and considered as potential sources;
  4. which sources appear as citations in the final answer.

These are different stages. A crawler may be able to access a page without that page being selected for a particular query. ChatGPT may mention a brand without linking to it. A link in the Sources panel also does not mean that the page supports every sentence in the response.

The Short Answer: How a Citation Is Generated

Based on OpenAI’s current documentation, the process can be summarized this way:

  1. ChatGPT decides whether to search the web. It may do this automatically when the question would benefit from current information, or the user may select Search manually.
  2. The original question may be rewritten. ChatGPT Search can send one or more targeted queries and then run narrower follow-up queries after reviewing the first results.
  3. Potential sources are retrieved. OpenAI states that ChatGPT Search uses third-party search providers as well as content supplied directly by partners.
  4. The answer is generated. ChatGPT combines the retrieved information with the context of the conversation.
  5. Sources are displayed. Responses that use Search may include inline citations, while the Sources panel may include both cited sources and other relevant links.

In other words, ChatGPT may not search the exact words the user typed. A page written around one rigid phrase may fail to match a more specific query generated during the search process.

Does ChatGPT Always Cite Sources?

No. Citations are tied to responses that use Search, and not every conversation triggers a web search.

ChatGPT may search automatically when current web information would improve the answer, and a user can also select Search manually. If Search was not used, the absence of a citation does not show that the website has a technical problem.

When Search is used, the answer may include:

  • citations next to specific claims;
  • a Sources button with a source panel;
  • images with separate source information;
  • direct navigational links to a website or brand.

OpenAI notes that the Sources panel may include cited sources and other relevant links. When evaluating visibility, check whether a specific URL is attached to a claim, not simply whether the domain appears somewhere in the panel.

What Is the Difference Between Discovery, Selection, and Citation?

Discovery

Discovery means that a page can enter the search process. OpenAI uses OAI-SearchBot for the automatic crawling that supports ChatGPT’s search features.

If a site blocks this crawler, OpenAI says its pages will not appear in ChatGPT Search answers, although the domain may still appear as a navigational link in some situations.

Selection as a Potential Source

After Search is triggered, the system works with generated queries and retrieved results. Official OpenAI materials confirm the use of third-party search providers and partner content, but they do not publish the exact formula used to rank every candidate source.

No official source establishes that one universal factor, such as word count, Schema.org markup, the number of FAQ entries, or a Google ranking, determines selection on its own.

Citation

A citation is the visible connection between a source and the generated answer. A URL may be considered during search without appearing as a citation. Conversely, the Sources panel may show a relevant link that is not attached to a specific sentence.

This distinction matters when you measure visibility. A missing citation is not proof that the system failed to discover the page.

The Three OpenAI User Agents That Are Most Commonly Confused

User agentPrimary roleDoes it determine participation in ChatGPT Search?
OAI-SearchBotAutomatic crawling for ChatGPT search featuresYes. This is the crawler used to manage opt-out from Search and automatic crawling.
GPTBotCrawling content that may be used to train generative AI foundation modelsNo. Allowing it is not a requirement for participation in ChatGPT Search.
ChatGPT-UserVisiting a page following a specific action or request by a userNo. OpenAI states that it is not used to determine whether content may appear in Search.

The settings for OAI-SearchBot and GPTBot are independent. A website may allow the Search crawler while blocking the use of its content for training through GPTBot.

An example robots.txt configuration for such a policy is:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

This is only an example. The site’s actual robots.txt file must be reviewed alongside its other directives, firewall and CDN rules, and the IP ranges published by OpenAI.

OpenAI also says its search systems may need about 24 hours to adjust after a robots.txt change.

What Is Officially Confirmed as Necessary?

OAI-SearchBot Must Not Be Blocked

For a site to be eligible to appear in ChatGPT Search through automatic crawling, it must allow OAI-SearchBot and accept requests from OpenAI’s published searchbot IP ranges.

Robots.txt is only the first check. The crawler may be allowed there and still be blocked by:

  • Cloudflare or another CDN security system;
  • a web application firewall;
  • an anti-bot plugin;
  • hosting protection;
  • rate limiting;
  • rules that return different statuses depending on the user agent;
  • login, cookie, or JavaScript requirements.

The Page Must Be Accessible as a Web Source

ChatGPT Search relies on information it can access on the web and link back to. If the useful answer exists only behind a login, inside a closed interface, or in an element that cannot be extracted reliably, the page is less useful as a source.

The Topic Must Match the Search Query Used

OpenAI’s documentation shows that ChatGPT Search may rewrite a question as one or more targeted queries. The page therefore needs to solve the user’s actual problem, not repeat a single keyword phrase.

If a user asks, “Which SEO agency offers ChatGPT optimization in Bulgaria?”, the system may search around providers, location, and service. If the question is “How do I allow the ChatGPT crawler in robots.txt?”, the query and the most appropriate source page will be different.

What May Improve the Chances, Based on Practical Interpretation?

The recommendations below are not an official OpenAI list of ranking factors. They are practical guidance based on how search, retrieval, and claim-level citations work.

One Clear Primary URL for Each Question

When several pages on the same site give overlapping but slightly different answers, there may be no obvious primary source for the topic.

A useful model is:

  • one page serves as the primary source for the topic;
  • related pages cover it only where the context requires it;
  • relevant internal links point to the primary page;
  • the title, H1, and opening answer reflect the page’s real purpose.

A Direct Answer That Can Be Connected to Evidence

The page should quickly show:

  • what the answer is;
  • under what conditions it is valid;
  • what the exceptions are;
  • where the information comes from;
  • when it was checked.

Generic claims such as “high-quality content is important” do not provide a specific, verifiable answer that can be meaningfully cited.

A Clear Distinction Between Fact, Experience, and Assumption

When official documentation does not disclose a mechanism, this should be stated clearly.

Suitable wording includes:

  • “OpenAI officially states...”;
  • “The documentation does not publish...”;
  • “Our practical test indicates...”;
  • “This is an interpretation, not a confirmed ranking factor.”

This helps readers separate verifiable facts from professional judgment.

Current and Specific Company Information

For queries about services, companies, specialists, or products, the main information should not conflict across different pages:

  • organization name;
  • services;
  • markets served;
  • address and contact details;
  • authors and expert roles;
  • prices, terms, and dates when publicly available;
  • links to primary documents and official profiles.

Consistency does not guarantee a citation, but it reduces the chance that the system will encounter conflicting claims about the business.

External Verifiability

A claim that appears only on a company’s own site is self-published. Editorial coverage, official registries, professional profiles, primary data, and credible independent sources can make that claim easier to verify.

This is not a reason to mass-create profiles or buy unrelated links. An external source should genuinely confirm the information and help the reader.

What Does Not Guarantee a Citation in ChatGPT?

Allowing GPTBot

GPTBot is associated with crawling content that may be used to train OpenAI’s generative AI foundation models. It is not the control for ChatGPT Search eligibility. That role belongs to OAI-SearchBot.

A Single llms.txt File

OpenAI’s current crawler documentation covers OAI-SearchBot, GPTBot, ChatGPT-User, robots.txt, and published IP ranges. It does not require an llms.txt file for participation in ChatGPT Search.

An llms.txt file cannot replace accessible content, a clear primary URL, and an allowed Search crawler.

Schema.org Markup by Itself

OpenAI has not published a rule stating that any schema type guarantees a citation. Structured data should accurately describe the visible page, but it is not a pass for automatic inclusion.

Ranking First for One Keyword

ChatGPT Search may rewrite a question as different, more specific queries. Ranking first for one tracked phrase does not prove that the same URL will be selected for every conversational version of that question.

Longer Content or More FAQ Questions

Adding words does not fix a missing or unclear answer. An FAQ section helps only when the questions are real, are not answered better in the main content, and do not duplicate another primary page.

One Successful Screenshot

One appearance shows that the page was used at one moment and in one context. It does not prove consistent visibility across related queries.

Practical Checklist: Why a Website Is Not Being Cited

1. Check Whether Search Was Used at All

Before diagnosing the site, confirm that the response contains inline citations or a Sources button. If ChatGPT did not use web search, the absence of your domain is not evidence of a technical problem.

2. Check OAI-SearchBot

Determine:

  • whether there is a separate rule for OAI-SearchBot;
  • whether the crawler falls under a broader Disallow rule;
  • whether the published IP ranges are allowed;
  • what HTTP status it receives;
  • whether it receives the same content as a normal user;
  • whether it is blocked by the CDN or firewall.

3. Check the Correct Primary URL

For every test query, determine which page should be cited.

If you cannot identify one URL with confidence, the problem is probably the site’s content architecture, not ChatGPT alone. Do not create a new article for every wording variation. Improve or clarify the existing primary page first.

4. Check Whether the Page Provides a Self-Contained Answer

A reader should be able to understand the main answer without piecing it together from several pages.

Check whether the URL contains:

  • a direct answer at the beginning;
  • definitions only where necessary;
  • conditions and exceptions;
  • practical steps;
  • an update date;
  • visible sources;
  • clear authorship or editorial responsibility.

5. Check the Evidence

For every important claim, ask:

  • Is it an official source?
  • Is it a primary source?
  • Is it reliable data?
  • Is it based on your own practical experience?
  • Is it a hypothesis, and is it labeled as such?

A page that mixes facts with unlabeled assumptions is harder to verify and less useful, whether or not it receives a citation.

6. Check the Internal Links

The primary URL should receive links from relevant parent and related pages. Anchor text should describe the topic clearly instead of sending competing signals to several URLs about the same problem.

7. Check Whether the Query Is Looking for Your Type of Page

An informational guide, service page, product category, and company profile each serve a different purpose.

For “How does OAI-SearchBot work?”, a technical guide is likely to be the right source. For “agency for ChatGPT optimization,” the system may look for a service page or a page that helps users choose a provider. One URL should not try to satisfy both purposes completely.

How to Test Citations Without Misleading Yourself

OpenAI does not provide a report showing every question for which a domain was considered as a potential source. Practical monitoring therefore needs a consistent method.

Create a limited set of queries based on real business tasks:

  • general informational questions;
  • specific problems;
  • comparisons;
  • provider queries;
  • local queries;
  • questions where freshness matters.

For every test, record:

FieldWhat to Record
PromptThe exact wording without later editing
Date and contextDate, country, language, and whether the conversation contains previous context
SearchWhether web search was triggered
Cited domainWhich website was shown
Cited URLThe exact page, not only the domain
Supported claimWhich part of the answer is connected to the citation
AccuracyWhether the source genuinely supports the statement
RepeatabilityWhether the result appears in subsequent tests

Do not change your strategy after one result. Compare groups of closely related queries and look for recurring issues:

  • the wrong primary URL;
  • missing or outdated information;
  • a blocked crawler;
  • a competing page with a more direct answer;
  • conflicting company information;
  • insufficient evidence.

The Most Common Reasons a Page Is Not Cited

  1. ChatGPT did not trigger Search. The answer was generated without visible web sources.
  2. OAI-SearchBot is blocked. The cause may be robots.txt, a firewall, CDN, or anti-bot system.
  3. The page does not match the reformulated query. It may be too broad or solve a different intent.
  4. There is no clear primary URL. Several pages compete with similar content.
  5. The answer is unclear or difficult to extract. Important facts are hidden, scattered, or presented only visually.
  6. The claims are not sufficiently verifiable. Primary sources, dates, or clear editorial responsibility are missing.
  7. Other sources solve the specific task better. Technical eligibility does not mean automatic selection.
  8. You are checking only the domain rather than the specific citation. The Sources panel may contain other relevant links.

How This Relates to GEO and AEO

GEO and AEO are industry terms for visibility in generated and direct answers. Their definitions, overlap, and differences from SEO are covered in our separate guide, What Are GEO and AEO?

This article has a narrower purpose: it explains ChatGPT Search, OpenAI’s crawlers, and visible citations. It is not a complete definition of GEO or AEO, and it is not the service page.

Frequently Asked Questions

Does OpenAI Publish the Exact Factors Used to Select Sources?

No. Official OpenAI materials explain when Search may run, how queries may be rewritten, the use of third-party search providers and partner content, crawler controls, and how citations appear. They do not publish a complete ranking model, signal weights, or a guaranteed URL-selection formula.

Do I Need to Allow GPTBot to Be Cited?

No. OAI-SearchBot is the crawler used for ChatGPT’s search features. GPTBot is associated with crawling content that may be used to train generative AI foundation models. The settings are independent.

Can a Blocked Website Still Be Mentioned?

OpenAI states that websites that deny access to OAI-SearchBot will not be displayed in ChatGPT Search answers, but they may still appear as navigational links. This should not be treated as a normal citation-visibility strategy.

Is Every Link in the Sources Panel a Real Citation?

Not necessarily. OpenAI’s Help Center says the panel may contain cited sources and other relevant links. Check whether the URL is attached to a specific claim.

How Long Should I Wait After Changing Robots.txt?

OpenAI says its search systems may need about 24 hours to adjust to a robots.txt change. After that period, also review server logs, CDN rules, and firewall behavior.

Can Citation Be Guaranteed?

No. You can improve technical eligibility, content structure, evidence, and alignment with the user’s need. The final source selected for a particular answer remains outside the control of the site or agency.

Checking Visibility in ChatGPT

If your site is not appearing as a source, the cause may be technical, related to the content itself, or rooted in the site’s information architecture.

As part of AI search engine optimization, we review:

  • access for OAI-SearchBot and the applicable IP ranges;
  • robots.txt, firewall, and CDN restrictions;
  • the correct primary URL for every important topic;
  • the quality and verifiability of the main claims;
  • internal links and the boundaries between pages;
  • current citations and the accuracy of the brand’s presentation;
  • what can realistically be measured and improved.

We do not promise guaranteed citations. Our goal is to remove verifiable barriers and build a site that offers clear, accessible, and reliable source pages for the questions your customers actually ask.

Sources

Sources checked on August 21, 2026.

Image source: ChatGPT

Гласувай
crosschevron-down