Promptwatch Logo

meta-externalagent

The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly.
meta-externalagent
AI CrawlerAI Training

What is meta-externalagent?

meta-externalagent is Meta's crawler for direct web indexing. Meta says its uses include training foundation AI models and improving products, so this is the Meta identity to consider when setting a policy for model-training collection.

It runs as a crawler rather than a one-off fetch initiated by a person. That separates it from Meta-ExternalFetcher, which Meta documents as a user-requested fetcher that may bypass robots.txt. A rule aimed at meta-externalagent should not be assumed to govern every Meta bot.

Meta documents two forms that may appear in logs: meta-externalagent/1.1 and meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler). Both contain the same stable token and should match a robots.txt group named meta-externalagent.

Allowing the crawler makes public content available for the uses Meta describes, including foundation-model training. It does not prove that Meta selected a page for a dataset or that a future model will reproduce it. Blocking new crawls is also separate from Meta AI search visibility, which Meta assigns to the meta-webindexer crawler.

Meta explicitly supports robots.txt for meta-externalagent and says it may cache the file for up to 24 hours. A change can therefore take a day to show up in request behavior. The robots.txt guide explains why this is a crawl instruction rather than a way to protect confidential content.

Cloudflare identifies the operator as Meta and links to https://developers.facebook.com/docs/sharing/webmasters/crawler. Its entry has no Web Bot Auth directory, and this bot does not have validated IP tracking in the record. The documented user agent helps label traffic, but it is not cryptographic proof of origin.

Relevant for AI searchMeta AI

Is meta-externalagent relevant for AI search?

Yes. meta-externalagent feeds Meta AI, so the pages it can reach shape what those AI products say about you.

meta-externalagent gathers public web content that can end up in the training data for large language models. Once your pages are in that set, they influence how the operator's models talk about you for that model generation. Allowing it lets your own writing carry weight in those answers; blocking it means the models learn about you from third parties instead.

How to handle meta-externalagent

Leave meta-externalagent allowed if your policy permits Meta to retrieve public content for foundation-model training and product indexing. To opt the whole hostname out of future crawls, publish:

User-agent: meta-externalagent
Disallow: /

Meta says the crawler follows robots.txt and may cache it for up to 24 hours, so wait that long before judging a new rule. Use narrower Disallow paths if only part of the site is excluded. Authentication or an edge block is still required when access must be technically prevented.

Examples

  • A documentation publisher allows ordinary search crawlers but adds a dedicated meta-externalagent group because its policy excludes foundation-model training.
  • A media site permits the crawler on press releases while disallowing an archive containing licensed articles.
  • An administrator changes the rule on Monday and checks logs after Meta's documented 24-hour robots.txt cache window instead of expecting an immediate stop.

Frequently asked questions about meta-externalagent

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

Meta operates it. Cloudflare's bot entry points to Meta's crawler documentation at https://developers.facebook.com/docs/sharing/webmasters/crawler.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard