What is Arquivo Web Crawler?
Arquivo.pt takes periodic copies of the Portuguese web so old pages can be consulted later. Its broad crawls start from seed lists that include previously collected .pt home pages, submitted suggestions, and the .pt DNS listing. The crawler downloads public pages and follows their embedded links to discover more material.
Collection happens on more than one timetable. Arquivo.pt reports three to four broad crawls per year and a daily crawl of selected online publications. About 90 percent of a broad collection is gathered within seven days, although slower or larger sites can take longer.
Requests identify with Arquivo-web-crawler and may name Heritrix, Brozzler, or Browsertrix in the rest of the string. The operator documents a ten-second courtesy pause between HTTP requests to the same site. It also states that the crawler follows the Robots Exclusion Protocol.
A visit is about historical preservation, not AI search or model training. Collected pages do not appear in Arquivo.pt immediately; the service says they become available for consultation after one year. That archived availability still does not show that an AI provider used the page.
