What is Cotoyogi?
Cotoyogi collects Japanese-language data resources for the Center for Research and Development on Data Lake within Japan's Research Organization of Information and Systems. The center publishes a dedicated Cotoyogi crawler page with its identity and access controls.
The published user agent is Mozilla/5.0 (compatible; Cotoyogi/4.0; +https://ds.rois.ac.jp/center8/crawler/). Requests using that string are ordinary crawler fetches rather than visits made by a person. The documentation does not state which kinds of Japanese pages receive priority or how a URL enters the collection queue.
Cotoyogi follows site-level robots rules. Its operator documents full-site and path-prefix exclusions, * wildcards, end-of-path $ matching, and Crawl-delay values expressed in seconds. It also recognizes a robots meta nofollow directive as an instruction not to follow links from that HTML document.
The directory places Cotoyogi in the AI Training category, and Cloudflare describes the broader project as infrastructure for data use and AI research and development. The crawler page itself only states that Cotoyogi collects Japanese-language data resources. It does not identify a particular training dataset, model, or public AI search service that will use a given page.
For site owners, the likely consequence of allowing the crawler is inclusion in that research collection process, not a direct referral or a guaranteed citation. A block keeps Cotoyogi from fetching the covered paths when it follows the published policy. It has no effect on other crawlers or on copies of the material acquired through another route.
The operator lists source addresses from 157.1.136.4 through 157.1.136.11, which offers a useful second check alongside the user agent. Cloudflare also records Cotoyogi as robots compliant, but provides no HTTP signature directory. This entry remains marked unverifiable and does not mark logged traffic as IP verified, so avoid treating either signal as automatic authorization.
