What is SemanticScholarBot?
SemanticScholarBot looks for academic PDFs on selected domains for Semantic Scholar. According to the official crawler page, the service makes those PDFs available through semanticscholar.org so researchers can find and understand published work. This is a scholarly document crawl, not a general sweep of every page on a site.
The published user agent is Mozilla/5.0 (compatible) SemanticScholarBot (+https://www.semanticscholar.org/crawler). Requests carrying that string are expected when the crawler is looking for documents. A repository or publisher can compare the requested URLs with its public paper and PDF routes to see whether the traffic fits that stated purpose.
Semantic Scholar is an AI-powered academic search product. It sources papers through publisher relationships, data providers, and web crawls, then uses its own models to process and classify the collection. A successful crawl can therefore affect whether a paper is available through Semantic Scholar's research tools.
Semantic Scholar also publishes scholarly data resources used for natural language processing and text mining. Its crawler documentation does not say that every PDF it retrieves enters one of those datasets, however. A SemanticScholarBot visit is not proof that the document trained a general-purpose language model.
The bot record says SemanticScholarBot respects robots.txt, and the official page says its user agent can be filtered or rejected. Site owners can allow public papers while excluding draft, embargoed, or high-cost download paths. Server or edge rules are still appropriate when a restriction must be enforced, since robots.txt is advisory.
Identity evidence is limited to the documented user-agent string in the supplied facts. The record leaves the operator field blank, marks verification as unverifiable, and provides neither published IP ranges nor a Web Bot Auth signature directory. Treat the token as a useful log classification, not as proof that a request is genuine.
