What is Library Of Congress Web Archiving?
Subject experts at the Library of Congress select websites for the Library's web archive. Each selection begins with one or more seed URLs, which may point to a single document, a section, or an entire domain. The Library normally notifies site owners before collection, except for government sites, and some projects ask for permission to archive or provide off-site access.
Heritrix is the program's main crawler. It follows links from a seed and downloads page resources needed for preservation. A project may capture a changing site weekly, monthly, or quarterly, while a less active selection can be visited on a longer interval.
The Library generally bypasses robots.txt because blocked images, CSS, and JavaScript can make an archived page incomplete. It prefers not to collect administration sections or other areas that were not intended for general access, and asks site owners to raise those scope concerns directly.
These captures support preservation and later research. The Library does not state that this crawler feeds AI search or trains general-purpose models. An archival request should not be counted as evidence that a page will appear in an AI response.
