What is meta-externalagent?
meta-externalagent is Meta's crawler for direct web indexing. Meta says its uses include training foundation AI models and improving products, so this is the Meta identity to consider when setting a policy for model-training collection.
It runs as a crawler rather than a one-off fetch initiated by a person. That separates it from Meta-ExternalFetcher, which Meta documents as a user-requested fetcher that may bypass robots.txt. A rule aimed at meta-externalagent should not be assumed to govern every Meta bot.
Meta documents two forms that may appear in logs: meta-externalagent/1.1 and meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler). Both contain the same stable token and should match a robots.txt group named meta-externalagent.
Allowing the crawler makes public content available for the uses Meta describes, including foundation-model training. It does not prove that Meta selected a page for a dataset or that a future model will reproduce it. Blocking new crawls is also separate from Meta AI search visibility, which Meta assigns to the meta-webindexer crawler.
Meta explicitly supports robots.txt for meta-externalagent and says it may cache the file for up to 24 hours. A change can therefore take a day to show up in request behavior. The robots.txt guide explains why this is a crawl instruction rather than a way to protect confidential content.
Cloudflare identifies the operator as Meta and links to https://developers.facebook.com/docs/sharing/webmasters/crawler. Its entry has no Web Bot Auth directory, and this bot does not have validated IP tracking in the record. The documented user agent helps label traffic, but it is not cryptographic proof of origin.
