We raised €6M. Learn more →
Promptwatch Logo

Content Signals

Content Signals is a robots.txt extension that lets publishers state how crawled content may be used—for search, AI input, or AI training—rather than only whether it may be fetched.
Updated August 1, 2026
SEO

Definition

Content Signals is an extension to robots.txt that separates permission to crawl from permission to use. A publisher adds a Content-Signal line declaring intent for three distinct purposes: search (including the page in a search index and linking to it), ai-input (using the page as grounding material for a generated answer), and ai-train (using it to train or fine-tune a model). Each is set to yes or no, so a site can welcome search and AI citation while declining training.

The distinction matters because the classic robots.txt vocabulary cannot express it. Disallowing a crawler to prevent training also removes the content from retrieval, which removes it from AI answers. Many publishers spent 2024 and 2025 discovering that a blanket block protected their content and erased their AI search visibility at the same time.

Cloudflare popularized the format in 2025 and began adding default signals for sites on its network, which pushed the syntax onto a large share of the web quickly. Content Signals is a preference expressed in plain text, not an enforcement mechanism: it carries the weight of a stated license term, and compliance is voluntary in the same way robots.txt compliance is. Publishers wanting stronger footing generally pair it with RSL for formal licensing and pay per crawl or bot rules for enforcement.

For GEO teams, Content Signals is the lowest-effort way to make an accurate declaration: one line of text that says "cite me, don't train on me" instead of an all-or-nothing block that costs citations.

Examples of Content Signals

  • A site adds `Content-Signal: ai-train=no, search=yes, ai-input=yes` to robots.txt so it stays citable in AI answers while opting out of model training.
  • A publisher that previously blocked GPTBot entirely replaces the block with content signals and recovers citations in ChatGPT and Perplexity.
  • A CDN applies a default content signal policy across customer domains, instantly changing declared usage terms for a large portion of the web.
  • A legal team documents its AI usage stance in a robots.txt content signal as a machine-readable companion to the site's terms of service.

Terms related to Content Signals

Robots.txt

Root directory file instructing search engine and AI crawlers which pages to crawl or avoid—now critical for managing GPTBot, PerplexityBot, and ClaudeBot.

SEO

AI Web Crawlers

Bots deployed by AI companies to fetch web content for training and retrieval—comprising 95%+ of tracked crawler traffic, led by GPTBot and PerplexityBot.

AI

Really Simple Licensing (RSL)

Really Simple Licensing (RSL) is an open XML standard that lets publishers declare machine-readable usage, licensing, and compensation terms for AI crawlers and agents.

AI

Pay Per Crawl

Pay per crawl is an arrangement where a site charges AI crawlers for each request using HTTP 402 Payment Required, instead of serving content for free or blocking it outright.

SEO

TDM Rights Reservation

TDM rights reservation is the use of legal and technical notices to reserve rights around text and data mining by AI systems.

AI

AI Training Data

The text, images, code, and multimedia content used to train large language models like current GPT models, Claude, and Gemini for AI applications.

AI

LLMs.txt

LLMs.txt is a proposed specification for controlling how AI crawlers and language models access website content, functioning as a robots.txt equivalent specifically designed for LLM interactions.

GEO

AI Indexing

How AI systems discover, process, and store web content for generating responses—distinct from traditional search indexing and critical for GEO.

AI

Publisher Licensing

Publisher licensing governs how AI companies access, train on, retrieve, display, or cite professional content.

AI

AI Search Visibility

How frequently and prominently brands appear in AI-generated responses—measured through Share of Model across ChatGPT, Perplexity, and AI Overviews.

GEO

Frequently Asked Questions about Content Signals

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

`search` covers indexing the page and linking to it in results. `ai-input` covers using the page as grounding material when generating an answer, which is what produces AI citations. `ai-train` covers using the content to train or fine-tune a model. Each can be set independently to yes or no.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard