What is TurnitinBot?
TurnitinBot collects online material for Turnitin's plagiarism prevention service. Student papers are compared with internet sources for matching passages, and Turnitin cannot know every possible source before a paper is submitted. Its general crawler therefore gathers public content broadly and filters out pages or links that are not useful to that comparison.
Some requests come through publisher and content-partner ingestion rather than the general web crawl. Crossref members using Similarity Check can provide full-text URLs for Turnitin to index. Other visits may reach public news sites, blogs, academic pages, or open-access repositories without a separate publishing arrangement.
Logs can show TurnitinBot/ContentIngest or the shorter Turnitin identity. Turnitin treats TurnitinBot and Turnitin as equivalent names for exclusion rules. Bad inbound links or an occasional parsing error can send the crawler to a URL that does not exist, which explains some 404 requests.
Turnitin's own crawler page says it follows robots.txt, while the accompanying crawler record marks it as noncompliant. That conflict makes the bot's metadata appropriately uncertain. Its published purpose is similarity checking, not AI search or general-purpose model training.
