Promptwatch Logo

Content Signals

Content Signals is a robots.txt extension stating how crawled content may be used—search, AI input, or AI training—beyond just fetch permission.
Updated September 6, 2026
SEO

Definition

Content Signals is an extension to robots.txt that separates permission to crawl from permission to use. A publisher adds a Content-Signal line declaring intent for three distinct purposes: search (including the page in a search index and linking to it), ai-input (using the page as grounding material for a generated answer), and ai-train (using it to train or fine-tune a model). Each is set to yes or no, so a site can welcome search and AI citation while declining training.

The distinction matters because the classic robots.txt vocabulary cannot express it. Disallowing a crawler to prevent training also removes the content from retrieval, which removes it from AI answers. Many publishers spent 2024 and 2025 discovering that a blanket block protected their content and erased their AI search visibility at the same time.

Cloudflare popularized the format in 2025 and began adding default signals for sites on its network, which pushed the syntax onto a large share of the web quickly; see our Cloudflare integration write-up for how that rollout affects crawler behavior. Content Signals is a preference expressed in plain text, not an enforcement mechanism: it carries the weight of a stated license term, and compliance is voluntary in the same way robots.txt compliance is. Publishers wanting stronger footing generally pair it with RSL for formal licensing and pay per crawl or bot rules for enforcement.

For GEO teams, Content Signals is the lowest-effort way to make an accurate declaration: one line of text that says "cite me, don't train on me" instead of an all-or-nothing block that costs citations.

Examples of Content Signals

  • A site adds `Content-Signal: ai-train=no, search=yes, ai-input=yes` to robots.txt so it stays citable in AI answers while opting out of model training.
  • A publisher that previously blocked GPTBot entirely replaces the block with content signals and recovers citations in ChatGPT and Perplexity.
  • A CDN applies a default content signal policy across customer domains, instantly changing declared usage terms for a large portion of the web.
  • A legal team documents its AI usage stance in a robots.txt content signal as a machine-readable companion to the site's terms of service.

Terms related to Content Signals

Robots.txt

Robots.txt is a root file that tells search engine and AI crawlers which pages to crawl or avoid—critical for managing GPTBot, PerplexityBot, and ClaudeBot.

SEO

AI Web Crawlers

AI crawlers are bots from AI companies fetching web content for training and retrieval—95%+ of crawler traffic, central to AI search and GEO.

AI

Really Simple Licensing (RSL)

Really Simple Licensing (RSL) is an open XML standard for publishers to declare licensing terms for AI crawlers—shaping AI search and GEO visibility.

AI

Pay Per Crawl

Pay per crawl is an access model where a site charges AI crawlers per request via HTTP 402, instead of serving content free or blocking it.

SEO

TDM Rights Reservation

TDM rights reservation reserves rights around text and data mining by AI systems—shaping AI search and GEO source control.

AI

AI Training Data

AI training data is the text, images, and code used to train LLMs like GPT and Claude—shaping the baseline knowledge models use in AI search and GEO.

AI

LLMs.txt

LLMs.txt is a proposed specification for controlling how AI crawlers and language models access website content, a robots.txt equivalent for LLM interactions.

GEO

AI Indexing

How AI systems discover, process, and store web content for generating responses—distinct from traditional search indexing and critical for GEO.

AI

Publisher Licensing

Publisher licensing governs how AI companies access, train on, retrieve, and cite professional content—shaping which sources appear in AI search and GEO.

AI

AI Search Visibility

How frequently and prominently brands appear in AI-generated responses—measured through Share of Model across ChatGPT, Perplexity, and AI Overviews.

GEO

AI Crawler Verification

AI crawler verification confirms crawler requests are genuine AI bots by checking user-agent against provider IP ranges, filtering spoofed traffic.

SEO

Frequently Asked Questions about Content Signals

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

`search` covers indexing the page and linking to it in results. `ai-input` covers using the page as grounding material when generating an answer, which is what produces AI citations. `ai-train` covers using the content to train or fine-tune a model. Each can be set independently to yes or no.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard