Promptwatch Logo

New York Times Newsgathering

New York Times Newsgathering is a shared identity for scripts written inside the Times newsroom.
The New York Times
Aggregator

What is New York Times Newsgathering?

New York Times Newsgathering is a shared identity for scripts written inside the Times newsroom. The Times says those programs collect public, non-copyright data from government and commercial websites. Individual projects range from archival work to public-service reporting such as election pages and Covid-19 trackers.

This is not one crawler working through a general queue. A newsroom data-acquisition project requests the sources needed for a particular piece of reporting or data product. The Times says its teams control request volume with throttling and concurrency settings, and it publishes a contact address that reaches the leads for those projects.

A legitimate request has several identifying details. The browser-like user agent ends in nyt_scraping/[email protected], while the headers include X-SCRAPED-BY: The New York Times and X-CONTACT: [email protected]. The official JSON declaration also provides a reverse-DNS pattern ending in .bot.newsdev.nytimes.com and a list of static source addresses.

The Times describes a reporting workflow, not a language-model training crawler or an AI search index. A request may support a newsroom article, archive, or public data page, but it is not evidence that the source will train a model or appear in an AI-generated answer.

Indirectly relevant

Is New York Times Newsgathering relevant for AI search?

Indirectly. New York Times Newsgathering has no AI product of its own, but its output can end up in the systems that AI answers draw on.

New York Times Newsgathering collects content from many sites and redistributes it through its own product. Aggregated copies can end up in datasets that AI systems later learn from, and some aggregators are themselves sources that AI answer engines draw on.

How to handle New York Times Newsgathering

Before granting an exception, compare the traffic with the live Times declaration. Check the user-agent suffix and both custom headers, then confirm the source against the published address list or reverse DNS. An email address in a header can be copied, so it should not be the only test.

The declaration provides no stable robots token and makes no robots.txt promise. Limit any allowlist to the public data endpoints the newsroom project needs. Send timestamps, request IDs, and unexpected paths to [email protected] if the volume or scope looks wrong. Authentication and edge policy should continue to protect non-public data.

Examples

  • An elections office checks the Times headers and source details before allowing a newsroom script to poll its public results feed.
  • A commercial data site finds the scraper on an unintended route and sends the published contact a request ID, path, and timestamp so the responsible team can correct the job.

Frequently asked questions about New York Times Newsgathering

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

No. The identity covers scripts and scrapers written for separate newsroom data-acquisition projects.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard