What is meta-webindexer?
meta-webindexer crawls web pages for Meta AI search. Meta says it analyzes online content to improve the relevance and accuracy of those results, making this a search indexer rather than a generic browser agent.
The connection to AI search is explicit. Meta's documentation says that allowing Meta-WebIndexer helps Meta AI cite and link to a site's content in responses. Access makes a page eligible for that workflow, but it does not promise a citation, a position, or a particular description of the source.
Meta does not describe this crawler as its foundation-model training collector. That role belongs to the separate meta-externalagent identity. Blocking meta-webindexer is therefore a choice about Meta AI search retrieval, not a documented opt-out from every Meta training use.
Expected requests contain either meta-webindexer/1.1 or meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler). Both variants share the stable meta-webindexer token used in robots.txt.
Robots guidance is contradictory across published records. Meta tells site owners to control this crawler with robots.txt and says it can cache the file for up to 24 hours. Cloudflare's directory marks followsRobotsTxt as false. Treat compliance as partial until your logs show that the crawler stopped requesting an excluded path.
Cloudflare attributes Meta-WebIndexer to Meta and links to https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/. No Web Bot Auth key directory or validated IP method is attached to the entry, so a user-agent match is not proof by itself. Pair the AI crawler label with request history and access controls.
