Markdown Doesn't Matter in AI Search (Yet)
Citation Share by Content Format
Based on 1,665,674 citations (last 7 days) across ChatGPT, Claude, Perplexity, and Google AI Overviews. These results aren't inconclusive or misleading: markdown simply does not influence citations in AI search. There are plenty of reasons to use .md files, but optimizing for AI search shouldn't be the strategy.
HTML
99.94%
Markdown (.md)
0.050%
Images
0.015%
What this means for you
Across 1,665,674 citations in a single week, HTML pages took 99.94% while .md files took 0.05%: a two-thousand-to-one ratio. AI search engines retrieve and cite rendered web pages, so the popular advice to publish markdown versions of your content "for the LLMs" has essentially zero payoff in search citations. Where you compete is in how parseable and useful your existing HTML is, not in offering an alternative format.
How to act on it
- Skip building .md mirrors of your marketing and blog pages for AI search visibility. Put that effort into the HTML pages that account for 99.94% of citations: clear heading hierarchy, semantic markup, and content that answers a question directly.
- If your pages rely heavily on client-side JavaScript to render, fix that before anything else. A page a retrieval system can't parse loses citations because it's unreadable, not because it isn't markdown.
- Reserve .md endpoints for developer documentation, where coding agents fetch them directly. That's a different channel with different economics than consumer AI search.
Markdown is not for AI search. It's for AI agents
At just 0.05% of citations, .md files are virtually absent from AI search results. ChatGPT Search and AI Overviews cite HTML web pages, news articles, forums, and documentation sites, not raw markdown files. The confusion between AI search and AI agents has led to misplaced optimization efforts.
Where markdown does matter is in the rapidly growing world of AI coding agents. We're seeing a huge uptick in documentation and site information being crawled by Claude's bot (powering Claude Code) and OpenAI's bot (powering Codex). These agents don't search the web the way ChatGPT Search does. They fetch specific docs, READMEs, and API references to complete coding tasks. Markdown is simpler, cleaner, costs less parsing time, and keeps the context window clean. This shift is accelerating with new agent-native search engines like Exa, already integrated into tools like OpenCode, that let agents search the web and retrieve URLs on their own. As more AI agents adopt these purpose-built search layers, the gap between consumer-facing AI search and developer-facing AI agent retrieval will only widen.
| Request Path | Requests | Crawlers |
|---|---|---|
| /docs/****/************.md | 3,241 | anthropic-claudebotopenai-searchbot |
| /docs/****/*********.md | 2,870 | anthropic-claudebotopenai-searchbot |
| /docs/******/************.md | 1,934 | anthropic-claudebotopenai-searchbot |
| /docs/****/**********.md | 1,512 | anthropic-claudebotopenai-searchbot |
| /docs/******/****************.md | 1,207 | anthropic-claudebotopenai-searchbot |
| /docs/********/*************.md | 986 | anthropic-claudebotopenai-searchbot |
| /docs/****/****************.md | 743 | anthropic-claudebotopenai-searchbot |
| /docs/******/********.md | 651 | anthropic-claudebotopenai-searchbot |
| /docs/********/**********.md | 489 | anthropic-claudebotopenai-searchbot |
| /docs/****/*******************.md | 372 | anthropic-claudebot |
| /docs/********/****************.md | 318 | openai-searchbot |
| /docs/*****/**********.md | 245 | openai-searchbot |
| /docs/***********/**********.md | 189 | openai-searchbot |
| ... | ... | ... |
How we collect this data
We collect millions of prompt responses, citations, and click data from the actual user interfaces of major AI platforms: over 26 billion data points and growing. This gives us one of the largest datasets on how AI search engines cite sources and recommend brands.
Real UI monitoring
Data straight from the interfaces of ChatGPT, Gemini, Perplexity, Claude, AI Overviews, and more.
26B+ data points
Over 26 billion analyzed citations, prompts, and responses, one of the largest AI search datasets available.
Continuously updated
Refreshed constantly so the trends you see reflect the latest behavior of AI search engines.
Aggregated & public
Published freely for the GEO community, based on aggregated, non-identifiable trends.
Want to start tracking your own AI search data? Get started with Promptwatch
Track What AI Search Actually Cites
Understand which content formats AI search engines actually cite. Track your visibility across ChatGPT Search and Google AI Overviews with real citation data.
