What is Bibliothèque nationale de France Crawler?
A visit from bnf.fr_bot is part of the Bibliothèque nationale de France's internet legal deposit work. The BnF uses the crawler to preserve online publications as national documentary heritage. It is an archival collector, not a search engine bot.
The crawler runs on Heritrix. It starts from selected pages, follows links, and fetches files needed to replay a site, including images and style sheets. Heritrix can misread JavaScript and request URLs that were never real, so some 404 responses during a capture do not point to broken links on the site.
The BnF identifies the traffic with bnf.fr_bot and says it leaves a long delay between requests to reduce server load. Its own policy also says it may ignore robots.txt. French legal deposit rules allow the library to retrieve excluded files when they are needed to reconstruct the published site.
The collected pages go into a preservation collection for later consultation and research. The BnF does not assign this crawler an AI search or general model training role. Seeing it in a log therefore says nothing about whether an AI assistant will index, cite, or train on the page.
