"Why scrape the ChatGPT interface when there's an API?"
It's the question we field most often, usually from someone weighing Promptwatch against a tool that does the opposite. So we tested it. Using Promptwatch, we put the same commercial prompts through both routes on the same day. One through ChatGPT's interface, which is what a real person searching ChatGPT actually sees. One through OpenAI's API. Same provider. Same questions. Same locale. 562 citations between them.
They came back describing two different markets. That gap is why Promptwatch tracks the interface, not the API.
ChatGPT UI vs API: the main differences
| ChatGPT UI | ChatGPT API | |
|---|---|---|
| What you're measuring | What your customers actually see | What the model can generate |
| Prompts that returned sources | 84% | 26% |
| Total citations | 392 | 170 |
| Unique domains cited | 149 | 67 |
| Editorial pages found (listicles, comparisons, reviews) | 48 | 3 |
| Community sources found (Reddit, forums) | 6 | 0 |
| Query fan-out visible | Yes | Never. No search runs, so no sub-queries exist |
| Rich results (images, shopping, local, ads) | Visible | Structurally impossible |
The two methods overlapped on just 25.6% of their sources. Seventy percent of what the interface found never appeared in the API results at all.
Why this matters for your visibility tracking
When you monitor brand mentions and citations in AI search, whether you're doing it manually or with a tool, where the data comes from changes everything. Most monitoring vendors rely on the OpenAI API or similar server-side integrations. We chose differently: Promptwatch collects data directly from the live ChatGPT user interface, the same interface your customers actually see when they ask a question.
This distinction compounds over time. It's the difference between measuring what's theoretically possible and what's actually happening in the real world. When you track real prompts and measure real citations, you see what's actually moving your visibility. When you rely on API data, you're measuring noise.
Here's what we found.

It made up a company
We asked: "platform that explains local real estate rules for expats buying homes in Europe."
The API recommended HomeRule Europe, complete with country guides, cost calculators, a vetted-expert directory, even a revenue model. Confident. Detailed. Well-structured.
HomeRule Europe does not exist.
That's the risk of API monitoring: a hallucination becomes a data point. You'd log it in your brand tracker with visibility 95, position 1. A fictional competitor lands at the top of your competitive set. No warning. No way to know unless you click through and realize the link is dead.
The interface, same prompt, same minute, returned four real platforms with working links. Two were emerging competitors the client hadn't been tracking. Those are real companies doing real work to rank in that query. That's where your attention should go.
This is exactly why Promptwatch monitors the live interface: we capture what users actually see, not what an API says it would have shown. When you track real prompts and real citations, hallucinations get filtered out immediately because there's nowhere to click to, no domain to return citations from.
Most of the category came back empty
| ChatGPT UI | ChatGPT API | |
|---|---|---|
| Prompts that returned sources | 84% | 26% |
| Total citations | 392 | 170 |
| Unique domains | 149 | 67 |
Three out of four API answers arrived with no sources at all. No links, nothing to click, nothing to verify.
If you tracked this category through an API, your report would show almost no citations, and you could reasonably conclude there was little AI visibility worth competing for. There is. Run the same prompts through the interface and 84% come back with sources. The visibility is there. The API simply isn't showing it to you.
Its "sources" training data, not the web
Web search was off. No retrieval happened. It still produced 170 citations, because the model typed URLs from memory:
- **[Idealista](https://www.idealista.com/)** – Especially strong in Spain...
- **[Kyero](https://www.kyero.com/)** – Focused on overseas homes...
87% of them point at a homepage. Sixty are a bare domain and nothing else. Eleven add only /en. Just eleven point to an actual page.
That's what model hallucination looks like: the model remembers Idealista exists. It doesn't know which page answers the question, because it never looked. It's generating plausible URLs, not retrieving them.
Why does the source matter?
When Promptwatch scrapes the live ChatGPT interface, every citation it logs is a link ChatGPT actually puts in front of a user. You can open it, read the page, and see exactly what content the answer was built on.
Three signals get muddled here, and they measure different things. Your server and CDN logs show when OpenAI's crawler fetched a page. Citation tracking shows when that page was cited in an answer. Referral traffic tagged utm_source=chatgpt.com shows when a reader clicked through to you. Crawled, cited, clicked. A page can be crawled and never cited, or cited and never clicked, and each gap tells you something different about your content. API monitoring gives you none of the three.
If you're optimizing based on API citations, you're optimizing against training data. If you're optimizing based on real citations from the interface, you're optimizing against what actually gets your pages read.
No fan-out. No rich results.
Two things an API cannot give you, no matter how it's configured.
Query fan-out. When someone asks ChatGPT a question, it doesn't run one search. It breaks the question into multiple sub-queries and searches each one separately: "expat home search europe multilingual property portals", "affordable second homes europe property portals".
Those sub-queries are the actual demand signal behind the prompt. They show you which angles of your content people care about. They're the single richest content signal in AI search.
An API returns a single string. It has no way to show you the sub-queries ChatGPT ran. A model with search disabled generates no fan-outs at all. There is nothing to capture, and nothing to optimize for.
Promptwatch captures these because we monitor the live interface. When ChatGPT breaks a question apart and searches multiple angles, we see each one. When you track real prompts, you get visibility into the semantic breakdown of what your customers are actually asking for.
Rich results. Image carousels, shopping modules, local business panels, ads. An API returns a string of text. It cannot contain any of them. If your brand shows up inside one of those elements, or a competitor does, API monitoring will never surface it.

The competitors you don't know about yet
API monitoring shows you the incumbents you could have named yourself. Interface monitoring shows you the emerging competitors, the Reddit threads gaining traction, the listicles outranking you in a query you thought you owned.
There is a structural reason for that. A model's weights are frozen at its training cutoff. A competitor that launched after that date does not exist inside them, and no amount of prompting will conjure one. Retrieval is the only route a new entrant has into an answer. In our test, the interface named two emerging platforms the client had never tracked. Asked the same question, the API named one that had never existed at all.
When you track real prompts in real time, you see the competitive landscape as it actually exists right now, not as it existed last quarter. You catch new entrants before they've saturated the market. You see which editorial tactics (listicles, Q&A formats, Reddit compilations) are actually getting cited, so you can match them. You see which content angles are driving visibility so you can double down on them.
This is exactly what Promptwatch does: monitor your customers' actual prompts, surface the actual competitors showing up in their actual answers, track position by position, and update weekly as the competitive landscape shifts.
How Promptwatch collects data differently
Promptwatch monitors ChatGPT, Gemini, Claude, and Perplexity by scraping the live user interface the same way a customer would use it. This gives us access to:
- Real citations. Not model memory, but the actual pages ChatGPT served to users as clickable links. We can show you which exact pages are being cited and how often, not hallucinated URLs pointing nowhere. Cross-referenced against your crawler logs, you can see the full path from OpenAI fetching a page to that page being cited in an answer.
- Query fan-outs. The sub-queries each prompt breaks into, showing you the semantic demand behind what customers actually search for. When you track prompts, you see not just "someone asked about European real estate" but "they asked about visa requirements, affordable markets, expat-friendly portals, and local rules". Each sub-query is a distinct content opportunity you might be missing.
- Competitive positions. Exactly where each competitor lands in each answer, position by position. Not buried in an API response. Real market positions that change week to week as the landscape shifts. This is how you catch emerging competitors before your SEO tools even register them.
- Emerging tactics. New sites and content formats (Reddit threads, listicles, UGC) that rank high enough to be cited. API monitoring would miss these entirely. Interface monitoring surfaces them immediately so you can decide whether to match them, outrank them, or contribute to the conversation.

When you track real prompts, you measure visibility using citation rate (percentage of tracked prompts where you're cited), share of voice (how often you show up vs competitors on the same prompts), and visibility scores (a composite metric you can trend over time). Same methodology as SEO measurement, but grounded in what actually gets your pages read by AI systems.
