What is Internet Archive Bot?
Internet Archive Bot crawls public web pages for the Internet Archive's Wayback Machine. The operator also calls it archive.org_bot, and its crawler record is linked from the user-agent strings listed by Cloudflare.
The crawler's purpose is preservation. It fetches pages so the Wayback Machine can retain snapshots from different points in time. A later visitor may be able to compare an older copy with the current page, even after the live site has changed.
That persistence is the main publishing consequence. Correcting or deleting a live page does not by itself rewrite a snapshot that the archive already holds. Sites with legal or records-management requirements should treat future crawl policy and requests about existing captures as separate tasks.
Internet Archive Bot is not an AI search or model-training crawler. Its own visit does not indicate that content entered a foundation-model dataset, and allowing it does not directly improve placement in an AI answer. Other people or systems may consult an archived page later, but that is separate from this crawler's archival job.
Cloudflare's bot directory lists two request forms, one containing special_archiver/3.1.1 and another containing archive.org_bot. Both point to the Internet Archive crawler page. The stable robots token for this bot is archive.org_bot.
This bot and Cloudflare's directory are recorded as following robots.txt. Cloudflare does not list a Web Bot Auth signature directory, however, so a matching user agent is not cryptographic proof. A robots directive controls crawl access; it should not be treated as a deletion command for snapshots already stored.
