What is MirrorWebCrawler?
MirrorWebCrawler makes archived copies of websites for MirrorWeb Ltd. The service is sold to financial and public-sector organizations that need retained website records, rather than to people searching the live web.
These requests produce a website archive for the customer that commissioned it. MirrorWeb does not document this crawler as a source for generative search or language-model training. A capture can matter for compliance records, but it has no stated effect on public AI search visibility.
A complete capture may include more than HTML. MirrorWeb says its archive can preserve dynamic and interactive content as users saw it, then make completed captures available for later search and replay. Scripts, styles, images, and other page resources may therefore appear alongside document requests during a crawl.
The crawler looks like a Chrome browser in access logs, with https://www.mirrorweb.com appended to the user agent. MirrorWeb says customer-specific variations can exist and advises customers to allow the crawler when site security blocks a complete capture. It does not publish a robots.txt commitment, so compliance is unknown.
