What is CMU-Cylab?
CMU-Cylab is a Carnegie Mellon University crawler built for research into the online trade in counterfeit documents. It queries search engines for material related to that topic, visits the result pages, and follows links from them. The operator says the crawl remains tied to the research subject rather than indexing the open web for a general search service.
Its published scope includes pages exposed to search engines and pages reachable from those results. The crawler skips non-HTML extensions and pages with downloads, and it is not intended to probe administration panels or other sensitive endpoints. A listed rate limit caps traffic at one request every three seconds per domain.
The project keeps crawled HTML and metadata for the duration of the research. Carnegie Mellon says classification models may be trained on this material for the academic study. That is a limited research use of machine learning, not a general-purpose model training program, and CMU-Cylab has no documented connection to AI search answers or citations.
Requests carry CMU-Cylab/1.0 inside a browser-like user agent. The operator page publishes source IP addresses, which provide a stronger check than the name alone. It does not state a robots.txt policy, while the crawler record reports that exclusions are not followed, so compliance should be treated as unconfirmed.
