Promptwatch Logo

MirrorWebCrawler

MirrorWebCrawler makes archived copies of websites for MirrorWeb Ltd.
MirrorWeb Ltdmirrorweb
Aggregator

What is MirrorWebCrawler?

MirrorWebCrawler makes archived copies of websites for MirrorWeb Ltd. The service is sold to financial and public-sector organizations that need retained website records, rather than to people searching the live web.

These requests produce a website archive for the customer that commissioned it. MirrorWeb does not document this crawler as a source for generative search or language-model training. A capture can matter for compliance records, but it has no stated effect on public AI search visibility.

A complete capture may include more than HTML. MirrorWeb says its archive can preserve dynamic and interactive content as users saw it, then make completed captures available for later search and replay. Scripts, styles, images, and other page resources may therefore appear alongside document requests during a crawl.

The crawler looks like a Chrome browser in access logs, with https://www.mirrorweb.com appended to the user agent. MirrorWeb says customer-specific variations can exist and advises customers to allow the crawler when site security blocks a complete capture. It does not publish a robots.txt commitment, so compliance is unknown.

Indirectly relevant

Is MirrorWebCrawler relevant for AI search?

Indirectly. MirrorWebCrawler has no AI product of its own, but its output can end up in the systems that AI answers draw on.

MirrorWebCrawler collects content from many sites and redistributes it through its own product. Aggregated copies can end up in datasets that AI systems later learn from, and some aggregators are themselves sources that AI answer engines draw on.

How to handle MirrorWebCrawler

Organizations that use MirrorWeb should coordinate crawler restrictions with the team responsible for the archive. Blocking pages or supporting assets can leave a capture incomplete. MirrorWeb's own guidance recommends allowing its user agent when a firewall or bot filter interferes.

The published identifier can also be expressed as a robots rule:

User-agent: mirrorweb
Disallow: /

There is no documented promise that MirrorWebCrawler follows robots.txt, and a customer may have a tailored user agent. If a crawl is not authorized, confirm it against the organization's MirrorWeb setup and use origin, firewall, or CDN controls for a reliable block.

Examples

  • A financial firm finds missing assets in an archived page and allows the MirrorWeb user agent through its bot filter before the next scheduled capture.
  • A public body spots a Chrome-like request ending in `https://www.mirrorweb.com` and confirms that its records team commissioned the crawl.

Frequently asked questions about MirrorWebCrawler

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

MirrorWeb archives how a site appeared, including dynamic and interactive content. Supporting resources may be needed to capture and replay that page.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard