Promptwatch Logo

GPTBot

Crawls web content to improve OpenAI's generative AI models and ChatGPT; respects 'robots.txt' directives to exclude sites from training data.
OpenAIGPTBot
AI CrawlerAI Training

What is GPTBot?

GPTBot is OpenAI's crawler for web content that may be used to train its generative AI foundation models and improve products such as ChatGPT. Cloudflare links the bot to OpenAI's GPTBot documentation and classifies it as an AI crawler.

A request from GPTBot shows that OpenAI's training crawler fetched a page. It does not confirm that the page was selected for a dataset, retained for training, or reflected in a particular model output. OpenAI says a GPTBot disallow signals that the site's content should not be used to train its generative foundation models.

GPTBot is not OpenAI's search crawler. The current crawler overview assigns search discovery to OAI-SearchBot and certain user-triggered actions to ChatGPT-User. A publisher can allow OAI-SearchBot for ChatGPT search while blocking GPTBot for training, because the robots.txt settings are independent.

OpenAI notes that when both GPTBot and OAI-SearchBot are allowed, it may use the result of one crawl for both purposes to avoid duplicate requests. That optimization does not merge the controls. Blocking GPTBot remains the documented signal against training use, while OAI-SearchBot governs search access.

OpenAI's example user agent contains GPTBot/1.4, but the version can change. It may add a robots.txt marker while fetching the policy file. The stable match is GPTBot, and OpenAI publishes current crawler addresses at https://openai.com/gptbot.json.

This entry is IP verified, tracked by Promptwatch, and recorded as respecting robots.txt. Cloudflare also marks robots compliance as true. The user-agent header can still be spoofed, so identity-sensitive rules should compare the source with OpenAI's current published ranges rather than trusting the string alone.

Relevant for AI searchChatGPTGPT models

Is GPTBot relevant for AI search?

Yes. GPTBot feeds ChatGPT, GPT models, so the pages it can reach shape what those AI products say about you.

GPTBot gathers public web content that can end up in the training data for large language models. Once your pages are in that set, they influence how the operator's models talk about you for that model generation. Allowing it lets your own writing carry weight in those answers; blocking it means the models learn about you from third parties instead.

Track GPTBot with Promptwatch

Promptwatch classifies GPTBot (GPTBot) in real time from your server and CDN logs. See exactly when it visits, which pages it requests, the status codes it gets, and how those crawls map to AI citations. Requests are validated against the operator's published IP ranges to filter out spoofed user agents.

How to handle GPTBot

To tell OpenAI not to crawl the site for possible foundation-model training, publish:

User-agent: GPTBot
Disallow: /

Keep the search decision separate. If the site should remain available to ChatGPT search, do not apply the GPTBot rule to OAI-SearchBot; give that crawler its own group. A GPTBot block is not a request to remove a page from ordinary search engines or from OpenAI's search crawl.

GPTBot is documented as following robots.txt. Check later access logs to confirm the policy is served correctly, and compare claimed traffic with https://openai.com/gptbot.json when origin identity matters. Use authentication rather than robots.txt for nonpublic content.

Examples

  • A newspaper disallows GPTBot across its archive but leaves OAI-SearchBot open so public stories can still be discovered in ChatGPT search.
  • A documentation site permits GPTBot on product manuals and excludes a path containing material licensed from another publisher.
  • A network team checks a burst labeled `GPTBot/1.4` against OpenAI's current IP file before treating it as authentic crawler traffic.

Frequently asked questions about GPTBot

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

OpenAI operates GPTBot as a crawler for content that may be used to train its generative AI foundation models.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard