What is kb.dk_bot?
Denmark's Royal Danish Library uses kb.dk_bot to build Netarkivet, the national web archive. Its legal deposit scope covers material in Danish, material published by Danes, and sites aimed at a Danish audience. The library has collected this part of the internet since 2005.
Netarkivet does not crawl every source on one uniform schedule. Broad collections can snapshot Danish domains up to four times a year. Selected news sites may be visited as often as 12 times a day, while event and research-driven collections create their own periods of activity.
The current Heritrix identity includes https://www.kb.dk/netarkivindsamling in the user agent. The library says its harvesters ignore robots.txt and can override robot meta tags. Those exclusions often hide files that are necessary to replay a page, which conflicts with the archive's preservation purpose.
Access to the collected material is restricted to research because the archive can contain sensitive personal information. Netarkivet does not describe kb.dk_bot as an AI search fetcher or a general model training crawler. Its appearance in a log records a preservation visit, not AI visibility or a likely citation.
