Promptwatch Logo

CMU-Cylab

An academic research bot, conducting research on patterns in counterfeit document sales over the internet.
Carnegie Mellon UniversityCMU-Cylab
Academic Research

What is CMU-Cylab?

CMU-Cylab is a Carnegie Mellon University crawler built for research into the online trade in counterfeit documents. It queries search engines for material related to that topic, visits the result pages, and follows links from them. The operator says the crawl remains tied to the research subject rather than indexing the open web for a general search service.

Its published scope includes pages exposed to search engines and pages reachable from those results. The crawler skips non-HTML extensions and pages with downloads, and it is not intended to probe administration panels or other sensitive endpoints. A listed rate limit caps traffic at one request every three seconds per domain.

The project keeps crawled HTML and metadata for the duration of the research. Carnegie Mellon says classification models may be trained on this material for the academic study. That is a limited research use of machine learning, not a general-purpose model training program, and CMU-Cylab has no documented connection to AI search answers or citations.

Requests carry CMU-Cylab/1.0 inside a browser-like user agent. The operator page publishes source IP addresses, which provide a stronger check than the name alone. It does not state a robots.txt policy, while the crawler record reports that exclusions are not followed, so compliance should be treated as unconfirmed.

Indirectly relevant

Is CMU-Cylab relevant for AI search?

Indirectly. CMU-Cylab has no AI product of its own, but its output can end up in the systems that AI answers draw on.

CMU-Cylab gathers pages for research corpora and web-measurement studies. Several widely used AI training datasets began as academic crawls, so pages collected for a study can later teach commercial models what your brand is.

How to handle CMU-Cylab

Decide whether public material on the site may be used in the counterfeit-document study. If you allow the project, keep account pages and private discussions protected by authentication. CMU-Cylab says it does not intentionally seek sensitive endpoints, but that statement is not a substitute for access control.

The documented token can be addressed directly:

User-agent: CMU-Cylab
Disallow: /

Carnegie Mellon's crawler page does not promise to obey robots.txt, and the accompanying crawler record marks it as noncompliant. Check later requests before relying on the rule. For a firm block or allowlist, compare the source address with the current list on the official CMU-Cylab page instead of trusting CMU-Cylab/1.0 on its own.

Examples

  • A marketplace sees `CMU-Cylab/1.0` move from a search result to several public listings about forged credentials. The security team checks the source IP against CMU's list and leaves customer account routes behind login.
  • An archive does not want its pages retained for the study. It adds a specific robots rule, observes that requests continue, and enforces the decision at the edge after confirming the traffic is genuine.

Frequently asked questions about CMU-Cylab

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

Carnegie Mellon University operates it to study patterns in the online trade in counterfeit documents.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard