Promptwatch Logo

meta-webindexer

Crawls web content to provide search results for Meta AI users.
meta-webindexer
AI CrawlerAI Assistant

What is meta-webindexer?

meta-webindexer crawls web pages for Meta AI search. Meta says it analyzes online content to improve the relevance and accuracy of those results, making this a search indexer rather than a generic browser agent.

The connection to AI search is explicit. Meta's documentation says that allowing Meta-WebIndexer helps Meta AI cite and link to a site's content in responses. Access makes a page eligible for that workflow, but it does not promise a citation, a position, or a particular description of the source.

Meta does not describe this crawler as its foundation-model training collector. That role belongs to the separate meta-externalagent identity. Blocking meta-webindexer is therefore a choice about Meta AI search retrieval, not a documented opt-out from every Meta training use.

Expected requests contain either meta-webindexer/1.1 or meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler). Both variants share the stable meta-webindexer token used in robots.txt.

Robots guidance is contradictory across published records. Meta tells site owners to control this crawler with robots.txt and says it can cache the file for up to 24 hours. Cloudflare's directory marks followsRobotsTxt as false. Treat compliance as partial until your logs show that the crawler stopped requesting an excluded path.

Cloudflare attributes Meta-WebIndexer to Meta and links to https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/. No Web Bot Auth key directory or validated IP method is attached to the entry, so a user-agent match is not proof by itself. Pair the AI crawler label with request history and access controls.

Relevant for AI search

Is meta-webindexer relevant for AI search?

Yes. meta-webindexer collects pages for an AI product, so what it can crawl influences how AI systems describe your brand.

meta-webindexer crawls and indexes pages so the AI search or assistant behind it can retrieve them at answer time. A page it has never fetched cannot be quoted, summarized, or linked in that product's answers, so most sites keep it allowed to stay citable. Blocking it removes your pages from that AI surface and hands those citations to competitors.

How to handle meta-webindexer

Keep meta-webindexer open if you want pages available for citation and linking in Meta AI search. To request exclusion from the entire site, use:

User-agent: meta-webindexer
Disallow: /

Meta says robots.txt changes can take up to 24 hours because the crawler caches the file. Cloudflare's record says the bot does not follow robots.txt, so check requests after that window. If it still reaches excluded URLs and the policy must be enforced, add an application or edge rule for those paths.

Examples

  • A public help center allows meta-webindexer because it wants Meta AI answers to link users to its current troubleshooting pages.
  • A subscription publisher excludes its article archive from Meta AI search while leaving public author and contact pages crawlable.
  • A site owner waits through Meta's 24-hour cache period, then checks whether meta-webindexer still requests a newly disallowed directory.

Frequently asked questions about meta-webindexer

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

Meta operates it. The official crawler documentation is at https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard