Promptwatch Logo

Cotoyogi

Cotoyogi collects Japanese-language data resources for the Center for Research and Development on Data Lake within Japan's Research Organization of Informati.
Unverifiablethe CenterCotoyogi
AI CrawlerAI Training

What is Cotoyogi?

Cotoyogi collects Japanese-language data resources for the Center for Research and Development on Data Lake within Japan's Research Organization of Information and Systems. The center publishes a dedicated Cotoyogi crawler page with its identity and access controls.

The published user agent is Mozilla/5.0 (compatible; Cotoyogi/4.0; +https://ds.rois.ac.jp/center8/crawler/). Requests using that string are ordinary crawler fetches rather than visits made by a person. The documentation does not state which kinds of Japanese pages receive priority or how a URL enters the collection queue.

Cotoyogi follows site-level robots rules. Its operator documents full-site and path-prefix exclusions, * wildcards, end-of-path $ matching, and Crawl-delay values expressed in seconds. It also recognizes a robots meta nofollow directive as an instruction not to follow links from that HTML document.

The directory places Cotoyogi in the AI Training category, and Cloudflare describes the broader project as infrastructure for data use and AI research and development. The crawler page itself only states that Cotoyogi collects Japanese-language data resources. It does not identify a particular training dataset, model, or public AI search service that will use a given page.

For site owners, the likely consequence of allowing the crawler is inclusion in that research collection process, not a direct referral or a guaranteed citation. A block keeps Cotoyogi from fetching the covered paths when it follows the published policy. It has no effect on other crawlers or on copies of the material acquired through another route.

The operator lists source addresses from 157.1.136.4 through 157.1.136.11, which offers a useful second check alongside the user agent. Cloudflare also records Cotoyogi as robots compliant, but provides no HTTP signature directory. This entry remains marked unverifiable and does not mark logged traffic as IP verified, so avoid treating either signal as automatic authorization.

Relevant for AI search

Is Cotoyogi relevant for AI search?

Yes. Cotoyogi collects pages for an AI product, so what it can crawl influences how AI systems describe your brand.

Cotoyogi gathers public web content that can end up in the training data for large language models. Once your pages are in that set, they influence how the operator's models talk about you for that model generation. Allowing it lets your own writing carry weight in those answers; blocking it means the models learn about you from third parties instead.

How to handle Cotoyogi

Use Cotoyogi's documented token to exclude the whole site in robots.txt:

User-agent: Cotoyogi
Disallow: /

The operator also supports narrower path rules and crawl pacing. For example, a site can disallow /images/, exclude *.gif$, or add Crawl-delay: 30.0 under the same user-agent group. A page-level nofollow meta directive only stops link following from that document; it is not a substitute for disallowing the document itself.

If unexpected traffic continues, compare the full user agent and source address with the details on the operator's crawler page. The published address range is 157.1.136.4 to 157.1.136.11. Questions can be sent to the crawler contact listed by ROIS.

Examples

  • A Japanese dictionary project allows Cotoyogi to fetch its openly licensed entries but excludes a directory containing licensed image scans.
  • A university repository keeps its public Japanese abstracts crawlable and adds `Crawl-delay: 30.0` after Cotoyogi requests begin competing with normal traffic.
  • An administrator checks a claimed Cotoyogi request against both the documented version 4.0 user agent and the `157.1.136.4` to `157.1.136.11` address range.

Frequently asked questions about Cotoyogi

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

The Center for Research and Development on Data Lake, part of the Research Organization of Information and Systems in Japan, operates Cotoyogi.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard