Promptwatch Logo

Robots.txt

Robots.txt is a root file that tells search engine and AI crawlers which pages to crawl or avoid—critical for managing GPTBot, PerplexityBot, and ClaudeBot.
Updated September 6, 2026
SEO

Definition

Robots.txt is a text file placed in a website's root directory that provides crawling instructions to web robots (bots and crawlers) about which pages or sections of the site should or should not be crawled. It follows the Robots Exclusion Standard and serves as the first communication between your website and any crawler—including the ai-crawlers that now dominate many sites' traffic.

In 2026, robots.txt management has become a strategic AI visibility decision. AI crawlers—GPTBot (OpenAI), PerplexityBot, ClaudeBot (Anthropic), Google-Extended, and others—now account for over 95% of crawler traffic on many websites. Your robots.txt configuration directly determines whether these AI systems can access and potentially cite your content. Blocking AI crawlers means your content won't appear in ChatGPT, Perplexity, or Claude responses.

Key robots.txt directives include User-agent (which crawler the rules apply to), Disallow (paths to block), Allow (exceptions within blocked paths), Sitemap (location of xml-sitemaps), and Crawl-delay (request pacing). You can set rules for specific AI crawlers independently—for example, allowing GPTBot while restricting other bots, or granting all AI crawlers access to your blog but blocking them from gated content.

Robots.txt works alongside the newer llms-txt standard, which serves as an AI-specific complement. While robots.txt tells crawlers what not to access, llms.txt proactively guides AI systems to your most valuable, citation-worthy content.

Important limitations: robots.txt only controls crawling, not indexing. Pages blocked by robots.txt can still appear in search results if linked from other sites (use noindex meta tags to prevent indexing). Robots.txt is a voluntary standard—malicious bots may ignore it. Never use robots.txt to hide sensitive data; use proper authentication instead.

Best practices: keep rules simple and readable, don't block CSS/JavaScript needed for rendering, reference your XML sitemaps, test changes before deployment using Google's robots.txt tester, and regularly review your AI crawler rules as new bots emerge and your content strategy evolves. The free robots.txt generator for AI is the fastest way to draft AI-aware rules, and our guide to seeing which AI bots are crawling your site shows how to confirm them in your logs.

Examples of Robots.txt

  • A publisher allows GPTBot and PerplexityBot access to their articles but blocks them from paywalled premium content, balancing AI visibility with content monetization
  • An e-commerce site blocks AI crawlers from checkout, account, and filtered product pages while allowing access to product pages and buying guides—directing AI citation toward valuable content
  • A media company reviews their robots.txt and discovers they accidentally blocked ClaudeBot, explaining why their content never appears in Claude responses—fixing it restores AI visibility within weeks
  • A SaaS company creates user-agent-specific rules allowing all AI crawlers to access their documentation and blog while blocking admin and staging directories
  • An SEO team reviews robots.txt alongside AI Overview citations, Bing/Copilot visibility, sitemap freshness, structured data validation, and AI crawler access before updating priority pages.

Terms related to Robots.txt

AI Web Crawlers

AI crawlers are bots from AI companies fetching web content for training and retrieval—95%+ of crawler traffic, central to AI search and GEO.

AI

LLMs.txt

LLMs.txt is a proposed specification for controlling how AI crawlers and language models access website content, a robots.txt equivalent for LLM interactions.

GEO

Crawl Budget

Crawl budget is the number of pages search and AI crawlers fetch in a timeframe; for large sites it decides what ChatGPT can cite.

SEO

XML Sitemaps

XML sitemaps are structured files listing website URLs with metadata to guide search engine and AI crawler discovery, crawl priority, and freshness.

SEO

Crawling and Indexing

Crawling and indexing are how search engines and AI crawlers discover and store web content, including GPTBot, ClaudeBot, and llms.txt.

SEO

AI Indexing

How AI systems discover, process, and store web content for generating responses—distinct from traditional search indexing and critical for GEO.

AI

OpenAI Crawlers

OpenAI crawlers such as GPTBot, OAI-SearchBot, and ChatGPT-User have different purposes for training, ChatGPT search, and user-triggered browsing.

AI

IndexNow

IndexNow is a protocol that notifies participating search engines of URL changes, improving freshness for AI search and LLM answers.

SEO

TDM Rights Reservation

TDM rights reservation reserves rights around text and data mining by AI systems—shaping AI search and GEO source control.

AI

AI Crawler Logs

AI crawler logs are server log records showing how AI bots, retrieval agents, and user-triggered AI browsers access a site for AI search and GEO visibility.

Analytics

AGENTS.md

AGENTS.md is a machine-readable instructions file that tells AI agents how to understand, navigate, and act on a site for AI search and agentic workflows.

GEO

Content Signals

Content Signals is a robots.txt extension stating how crawled content may be used—search, AI input, or AI training—beyond just fetch permission.

SEO

Really Simple Licensing (RSL)

Really Simple Licensing (RSL) is an open XML standard for publishers to declare licensing terms for AI crawlers—shaping AI search and GEO visibility.

AI

Pay Per Crawl

Pay per crawl is an access model where a site charges AI crawlers per request via HTTP 402, instead of serving content free or blocking it.

SEO

AI Crawler Verification

AI crawler verification confirms crawler requests are genuine AI bots by checking user-agent against provider IP ranges, filtering spoofed traffic.

SEO

Frequently Asked Questions about Robots.txt

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

For most businesses seeking AI visibility, allow AI crawlers (GPTBot, PerplexityBot, ClaudeBot, Google-Extended) access to your public, citation-worthy content. Block them from private areas, gated content, and low-value pages. Blocking all AI crawlers means your content won't appear in AI responses—a significant visibility loss as AI search captures 12–15% market share.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard