Promptwatch Logo

Arquivo Web Crawler

Web crawler archives the Portuguese web.
ArquivoArquivo-web-crawler
Archiver

What is Arquivo Web Crawler?

Arquivo.pt takes periodic copies of the Portuguese web so old pages can be consulted later. Its broad crawls start from seed lists that include previously collected .pt home pages, submitted suggestions, and the .pt DNS listing. The crawler downloads public pages and follows their embedded links to discover more material.

Collection happens on more than one timetable. Arquivo.pt reports three to four broad crawls per year and a daily crawl of selected online publications. About 90 percent of a broad collection is gathered within seven days, although slower or larger sites can take longer.

Requests identify with Arquivo-web-crawler and may name Heritrix, Brozzler, or Browsertrix in the rest of the string. The operator documents a ten-second courtesy pause between HTTP requests to the same site. It also states that the crawler follows the Robots Exclusion Protocol.

A visit is about historical preservation, not AI search or model training. Collected pages do not appear in Arquivo.pt immediately; the service says they become available for consultation after one year. That archived availability still does not show that an AI provider used the page.

Indirectly relevant

Is Arquivo Web Crawler relevant for AI search?

Indirectly. Arquivo Web Crawler has no AI product of its own, but its output can end up in the systems that AI answers draw on.

Arquivo Web Crawler captures snapshots of pages and stores them, usually for good. Public archives are a common ingredient in AI training datasets, and AI tools sometimes cite an archived version when the live page is gone. What it captures today can still be describing your brand years from now.

How to handle Arquivo Web Crawler

Use a narrow exclusion for crawler traps or material that should not enter future collections. Keep stylesheets, scripts, and images accessible when you want the archived page to replay properly, since Arquivo.pt specifically advises authors to permit embedded files.

The operator documents this full robots.txt block:

User-agent: Arquivo-web-crawler
Disallow: /

Robots compliance is documented, not inferred from the user agent. Replace / with a path when only one section should be skipped. An empty Disallow explicitly permits the whole site. These rules govern later crawling and do not constitute a request to remove copies that Arquivo.pt already holds.

Examples

  • During a broad collection, a `.pt` publisher sees `Arquivo-web-crawler` return at roughly ten-second intervals while it follows links through public pages.
  • A museum disallows an infinite calendar path but leaves exhibit images and CSS open so the useful pages can still be preserved and replayed.

Frequently asked questions about Arquivo Web Crawler

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

It preserves public web content with a focus on the Portuguese web. Broad crawl seeds include `.pt` sites and submitted suggestions.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard