Promptwatch Logo

Siteimprove Crawl

Siteimprove content suite (i.e. Quality Assurance, Accessibility, Policy, and SEO). Crawls run on ports are 80 for HTTP and 443 for HTTPS.
SiteimproveSiteCheck-sitecrawl
SEO

What is Siteimprove Crawl?

Siteimprove Crawl scans sites connected to Siteimprove's content suite. It requests pages over the standard web ports, 80 for HTTP and 443 for HTTPS, and supplies crawl data for quality assurance, accessibility, policy, and SEO checks.

A Siteimprove account can run a full scan on a schedule or recheck selected pages. Siteimprove separates fetching from later analysis, so Crawler Management can show a crawl as finished while link checks and other processing are still pending. Exclusions and deduplication can also make the final product totals smaller than the raw page and link counts.

This is recurring audit traffic, not a search engine visit. Siteimprove says its crawler normally pauses between requests and can slow down when a server struggles. Customers can also exclude sections or change scan timing. Those account controls are more precise than blocking every request after it reaches the site.

The available sources document Siteimprove Crawl as a site-audit client, not a public AI answer engine, and do not connect it to model training. Fixes prompted by a Siteimprove report may improve the site itself, but allowing this crawler has no documented direct effect on AI citations or rankings.

Indirectly relevant

Is Siteimprove Crawl relevant for AI search?

Indirectly. Siteimprove Crawl has no AI product of its own, but its output can end up in the systems that AI answers draw on.

Siteimprove Crawl is an SEO auditing crawler. It has no direct effect on search rankings or AI answers, but the reports it feeds are used to fix crawlability and content issues that do affect how search and AI systems see your site.

How to handle Siteimprove Crawl

Keep Siteimprove Crawl reachable when your organization or an agency uses Siteimprove. If scans are too heavy, first change the crawl schedule, request delay, or content exclusions in Siteimprove. An unexpected visit is also worth checking against your Siteimprove account before you block it.

You can state a crawl preference for the recorded token in robots.txt:

User-agent: SiteCheck-sitecrawl
Disallow: /

The source record does not establish that this client honors robots.txt, so the stanza is not a guaranteed block. Confirm the result in access logs and use a web server or edge rule if you need enforcement.

Examples

  • A university schedules a weekly Siteimprove scan. Requests bearing `SiteCheck-sitecrawl` cover the public site first, and accessibility results appear only after Siteimprove finishes processing the crawl.
  • A retailer sees the crawler burdening a fragile archive. The team excludes that section in Siteimprove, then checks its logs after the next scheduled scan to confirm the request volume fell.

Frequently asked questions about Siteimprove Crawl

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

It supplies page and link data to Siteimprove checks for quality assurance, accessibility, policy, and SEO. Siteimprove performs some checks after the fetch stage, so a completed crawl does not always mean the reports are already updated.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard