Promptwatch Logo

ICC Crawler

ICC-Crawler automatically crawls the Internet and collects web pages.
NICTICC-Crawler
AI Crawler

What is ICC Crawler?

ICC-Crawler is a web collection program run by the Universal Communication Research Institute at Japan's National Institute of Information and Communications Technology, or NICT. Its official crawler page says it automatically travels the Internet and collects web pages.

NICT's current policy says the crawler reads the robots.txt file on each target host and follows its access restrictions. When a site sets Crawl-Delay, ICC-Crawler uses either that interval or its own minimum interval, whichever is longer. This gives operators both scope and pacing controls.

An archived NICT description ties the collection to research on web search and data mining. The available material does not say that pages train a foundation model or feed a public AI answer service. An ICC-Crawler request should therefore be treated as research collection, not as proof of model training, AI search eligibility, or a future citation.

Cloudflare's bot directory records the user agent as ICC-Crawler/3.0 (Mozilla-compatible; ; https://ucri.nict.go.jp/en/icccrawler.html). The stable robots.txt token is ICC-Crawler. Versioned log matching should use the token rather than requiring the complete string.

The bot is verified in the directory, but this entry has no verified IP ranges or HTTP message-signature directory. The user-agent header still can be copied by another client. If an exception would bypass a security control, origin identity needs evidence beyond the string.

Sites that support NICT's web research can allow public pages while excluding expensive endpoints, personal-information routes, or file classes. NICT also invites operators to contact the institute if collection continues after a robots.txt restriction, giving sites a second route when the published rule does not produce the expected result.

Relevant for AI search

Is ICC Crawler relevant for AI search?

Yes. ICC Crawler collects pages for an AI product, so what it can crawl influences how AI systems describe your brand.

ICC Crawler fetches, indexes, or retrieves pages for an AI product, so the pages it can reach shape how that system describes your brand and products.

How to handle ICC Crawler

To stop collection across the host, use the token documented by NICT:

User-agent: ICC-Crawler
Disallow: /

NICT also documents path-specific Disallow rules, file-pattern exclusions, Allow exceptions, and Crawl-Delay. A research repository might permit final publications while excluding temporary exports rather than applying a site-wide block.

This crawler is recorded as respecting robots.txt. Review access logs after a change. If ICC-Crawler continues to fetch blocked paths, NICT's crawler page asks site owners to contact the institute so it can stop collection from the host. Use ordinary edge controls as well when a denial must be immediate or resistant to spoofed user agents.

Examples

  • A public research library allows ICC-Crawler to collect published reports but excludes its temporary document-conversion directory.
  • A media archive sets a crawl interval for ICC-Crawler and confirms from timestamps that the bot uses the slower rate.
  • A webmaster contacts NICT after requests continue on a path that is explicitly disallowed for `ICC-Crawler`.

Frequently asked questions about ICC Crawler

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

The Universal Communication Research Institute at Japan's National Institute of Information and Communications Technology operates ICC-Crawler.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard