Promptwatch Logo

Bibliothèque nationale de France Crawler

Bibliothèque nationale de France's mission is to collect, catalog, preserve, enrich and communicate the national documentary heritage.
Bibliothèque nationale de Francebnf.fr_bot
Academic Research

What is Bibliothèque nationale de France Crawler?

A visit from bnf.fr_bot is part of the Bibliothèque nationale de France's internet legal deposit work. The BnF uses the crawler to preserve online publications as national documentary heritage. It is an archival collector, not a search engine bot.

The crawler runs on Heritrix. It starts from selected pages, follows links, and fetches files needed to replay a site, including images and style sheets. Heritrix can misread JavaScript and request URLs that were never real, so some 404 responses during a capture do not point to broken links on the site.

The BnF identifies the traffic with bnf.fr_bot and says it leaves a long delay between requests to reduce server load. Its own policy also says it may ignore robots.txt. French legal deposit rules allow the library to retrieve excluded files when they are needed to reconstruct the published site.

The collected pages go into a preservation collection for later consultation and research. The BnF does not assign this crawler an AI search or general model training role. Seeing it in a log therefore says nothing about whether an AI assistant will index, cite, or train on the page.

Indirectly relevant

Is Bibliothèque nationale de France Crawler relevant for AI search?

Indirectly. Bibliothèque nationale de France Crawler has no AI product of its own, but its output can end up in the systems that AI answers draw on.

Bibliothèque nationale de France Crawler gathers pages for research corpora and web-measurement studies. Several widely used AI training datasets began as academic crawls, so pages collected for a study can later teach commercial models what your brand is.

How to handle Bibliothèque nationale de France Crawler

Allow the crawler to reach public pages and their assets if you want the BnF's legal deposit copy to replay correctly. Put private files, administration routes, and unpublished material behind authentication. The crawler may ignore an exclusion file, so robots.txt is not an access control for sensitive content.

The stable token still lets you record an exclusion preference:

User-agent: bnf.fr_bot
Disallow: /

Do not assume that rule will stop the capture. The BnF explicitly documents circumstances in which it bypasses robots.txt. If the crawl causes load problems or generates large numbers of bad JavaScript URLs, send timestamps and relevant log entries to [email protected]. The library publishes that address for site owners who need a collection adjusted.

Examples

  • A regional newspaper sees `bnf.fr_bot` fetch an article, its photographs, and the linked CSS. The publisher leaves those public assets available so the archived page keeps its layout.
  • Heritrix derives hundreds of nonexistent paths from a cultural site's JavaScript bundle. The site owner sends the timestamps and sample 404 requests to `[email protected]` so the BnF can adjust the capture.

Frequently asked questions about Bibliothèque nationale de France Crawler

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

The BnF captures online publications under France's internet legal deposit rules. Your site or a page on it was included in one of those preservation collections.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard