Promptwatch Logo

AI Search External

Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.
CloudflareCloudflare-AI-Search-External
AI SearchAI Crawler

What is AI Search External?

AI Search External is the identity Cloudflare AI Search uses when a crawl leaves the customer's own Cloudflare account and reaches another domain. It is not a separate global search engine. A Cloudflare customer starts the indexing job, and Cloudflare performs the external request on that customer's behalf.

External traversal is off by default. In the AI Search discovery settings, a customer must enable external links or subdomains before the crawler follows them beyond the source host. On domains inside the customer's Cloudflare account the crawler identifies as Cloudflare-AI-Search; on a domain outside that account it switches to Cloudflare-AI-Search-External.

Cloudflare's AI Search documentation says a website data source is fetched, converted to Markdown, split into chunks, and added to the customer's index. That index can support retrieval or generated answers in the application the customer builds. The customer also decides whether to expose search, chat completion, or MCP endpoints.

An external crawl can therefore make a page searchable inside one customer's AI Search instance. It does not place the page in a universal Cloudflare index, and it does not promise a public citation. Cloudflare does not document this crawler as a source of foundation-model training data, so a visit should not be described as model training.

The complete recorded user agent is Cloudflare-AI-Search-External (https://developers.cloudflare.com/ai-search; [email protected]). The stable token is Cloudflare-AI-Search-External. It states the operator, documentation location, and contact address without relying on a changing browser version.

Robots behavior needs a qualified answer. This page's structured bot status is unknown, while the supplied Cloudflare directory record marks the external variant as not following robots.txt. Current AI Search product documentation says pages disallowed by robots.txt are recorded as blocked_by_robots_txt instead of being indexed. Because those records disagree, publish a specific directive but use an actual deny rule when exclusion must be certain.

Cloudflare signs these requests with Web Bot Auth. The public key directory for the external identity is https://search.ai.cloudflare.com/external/.well-known/http-message-signatures-directory. Validating the HTTP message signature can prove that Cloudflare sent the request; it does not identify which Cloudflare customer configured the crawl or what that customer will do with indexed results.

Relevant for AI searchAI search

Is AI Search External relevant for AI search?

Yes. AI Search External feeds AI search, so the pages it can reach shape what those AI products say about you.

AI Search External crawls and indexes pages so the AI search or assistant behind it can retrieve them at answer time. A page it has never fetched cannot be quoted, summarized, or linked in that product's answers, so most sites keep it allowed to stay citable. Blocking it removes your pages from that AI surface and hands those citations to competitors.

How to handle AI Search External

Allow this crawler only if you accept another organization's Cloudflare AI Search instance indexing the page. Blocking it does not affect the internal Cloudflare-AI-Search identity, so write separate policy for that token when needed.

Record a site-wide preference with:

User-agent: Cloudflare-AI-Search-External
Disallow: /

The available robots records conflict, so do not rely on that file for a hard security boundary. If access must stop, deny the signed external agent at the edge or protect the content with authentication. Validate Web Bot Auth against https://search.ai.cloudflare.com/external/.well-known/http-message-signatures-directory before granting any exception based on identity.

Examples

  • A Cloudflare customer enables external link discovery for a documentation index, causing linked vendor pages to receive the external user agent.
  • A publisher permits normal Cloudflare AI Search jobs for its own account but blocks `Cloudflare-AI-Search-External` so other customers cannot add articles to their indexes.
  • An edge team verifies the request signature against the external key directory before classifying a suspicious user agent as genuine Cloudflare traffic.

Frequently asked questions about AI Search External

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

A Cloudflare AI Search customer has enabled crawling beyond its source host or account. External link and subdomain traversal are disabled by default.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard