Promptwatch Logo

kb.dk_bot

Royal Danish Library collects the Danish Internet according to the Danish Legal Deposit Act for research purposes.
Netarkivetnetarkivindsamling
Academic Research

What is kb.dk_bot?

Denmark's Royal Danish Library uses kb.dk_bot to build Netarkivet, the national web archive. Its legal deposit scope covers material in Danish, material published by Danes, and sites aimed at a Danish audience. The library has collected this part of the internet since 2005.

Netarkivet does not crawl every source on one uniform schedule. Broad collections can snapshot Danish domains up to four times a year. Selected news sites may be visited as often as 12 times a day, while event and research-driven collections create their own periods of activity.

The current Heritrix identity includes https://www.kb.dk/netarkivindsamling in the user agent. The library says its harvesters ignore robots.txt and can override robot meta tags. Those exclusions often hide files that are necessary to replay a page, which conflicts with the archive's preservation purpose.

Access to the collected material is restricted to research because the archive can contain sensitive personal information. Netarkivet does not describe kb.dk_bot as an AI search fetcher or a general model training crawler. Its appearance in a log records a preservation visit, not AI visibility or a likely citation.

Indirectly relevant

Is kb.dk_bot relevant for AI search?

Indirectly. kb.dk_bot has no AI product of its own, but its output can end up in the systems that AI answers draw on.

kb.dk_bot gathers pages for research corpora and web-measurement studies. Several widely used AI training datasets began as academic crawls, so pages collected for a study can later teach commercial models what your brand is.

How to handle kb.dk_bot

If the site is public and falls within the Danish collection scope, allowing the requests gives Netarkivet a more complete copy. Use login protection for intranets, administration tools, and other private areas. The library says it collects publicly available material, so access controls should reflect what is genuinely public.

A rule for the recorded token can document your preference:

User-agent: netarkivindsamling
Disallow: /

Netarkivet openly states that it ignores robots.txt, so this snippet is not an enforcement measure. Contact the Royal Danish Library with timestamps and affected URLs if a collection causes heavy load or enters an unsuitable public route. A network block will stop requests, but it can leave the legal deposit capture incomplete.

Examples

  • A Danish newspaper receives `netarkivindsamling` requests throughout the day. Its operations team compares the pattern with Netarkivet's selective news collection, which can run much more often than a national snapshot.
  • Before closing a public campaign site, a company leaves its pages, images, and PDFs online for an archival capture. Its customer dashboard stays behind authentication and is never exposed for the sake of the archive.

Frequently asked questions about kb.dk_bot

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

It collects the Danish part of the internet for Netarkivet under the Danish Legal Deposit Act. The scope includes Danish-language material and sites published by or directed at Danes.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard