Across the five AI engines, just 0.4% of responses cite PDF documents — the overwhelming majority of citations point to ordinary written web pages.
Only 80 of 19,279 responses (0.4%) across the five tracked AI engines cite PDF documents. Written, HTML web pages account for essentially all of the rest — so the answer to whether AI models prefer PDF content over articles is a clear no, at least for this prompt set.
PDF documents citation rate by AI engine
Google AI Mode cites pdf documents most (0.8%) and Perplexity least (0.1%) — a 12× spread, but even the leader stays in low single digits. Pdf content is a niche citation source, not a mainstream one, in every engine tracked. The engine that leans on it most (Google AI Mode) is worth noting if PDF is central to your content, but no engine treats it as a primary format.
The underlying numbers
| AI engine | Responses | Citing pdf documents | Rate |
|---|---|---|---|
| Google AI Mode | 3,975 | 31 | 0.8% |
| ChatGPT | 5,119 | 35 | 0.7% |
| Gemini | 5,123 | 10 | 0.2% |
| Google AI Overviews | 2,010 | 2 | 0.1% |
| Perplexity | 3,052 | 2 | 0.1% |
What actually gets cited instead
The PDFs that do get cited skew toward official and institutional documents (government bodies, standards organisations, large consultancies) that exist primarily in PDF form — not marketing collateral. If content can live as an HTML page, that version is far more likely to be surfaced and quoted than a PDF of the same material.
PDF documents citations over time
Tracked daily in Google AI Mode (its strongest engine), the PDF citation rate stays low and flat across the window — this is a stable structural preference for written pages, not a temporary dip.
What this means for content strategy
If your goal is AI-search citations, written HTML pages remain the workhorse format — PDF documents are cited far too rarely to build a strategy around. Publish in HTML first. Reserve PDFs for formal documents where the format itself adds credibility (reports, specs, filings); for everything else, an HTML page is the citable asset. If you must ship a PDF, publish an HTML companion page too.
Why PDFs are almost invisible to AI search
The 0.4% PDF citation rate is even lower than the video citation rate (2.6%). This is striking because PDFs are, in principle, easier for AI engines to process than video — they contain extractable text, they are static documents, and they are widely used for authoritative content like research papers, government reports, and technical specifications.
So why are PDFs cited so rarely? There are several likely explanations:
PDFs are often gated or hard to discover. Many PDFs live behind forms, paywalls, or in deep directory structures that search crawlers may not fully explore. An HTML page is typically linked from the site’s navigation; a PDF is often an isolated file.
PDFs are harder to parse consistently. While PDFs contain text, the extraction process can be unreliable — multi-column layouts, embedded images, and complex formatting can produce garbled or incomplete text. HTML is a structured format designed for text extraction.
PDFs lack metadata and context. An HTML page typically has a title tag, meta description, headings, and schema markup. A PDF has none of these — or has them in a less standardized form. AI engines rely on this metadata to understand what a page is about.
PDFs are not the web’s primary format. The web is overwhelmingly HTML. AI engines are optimized for HTML. PDFs are a secondary format that requires special handling.
The 0.4% citation rate does not mean PDFs are worthless. It means they are a niche format for AI citations. If your content can live as an HTML page, that is the version that will earn citations.
When PDFs do get cited
The PDFs that do make it into the cited set have a clear pattern: they are official, authoritative documents that exist primarily or exclusively in PDF form. Examples from our data include:
- Government white papers and regulatory filings
- Academic research papers (especially from .edu domains)
- Technical standards documents (ISO, W3C, IETF)
- Large consultancy reports (McKinsey, Deloitte, BCG)
These documents share two traits: they are authoritative by nature of their source, and they are not available in an equivalent HTML version. The AI engine cites the PDF because it is the best available source for that information — not because PDF is a preferred format.
For most organizations, the content you produce (blog posts, guides, product documentation, case studies) does not have these traits. It can — and should — live as HTML.
The HTML companion strategy
If you have content that must be published as a PDF (a formal report, a white paper, a specification), the most effective approach is to publish an HTML companion page. This is a web page that:
- Summarizes the key findings or specifications from the PDF
- Provides context and background that the PDF may lack
- Links to the PDF for readers who want the full document
- Includes proper heading structure, metadata, and schema markup
The HTML companion page becomes the citeable asset. The PDF remains available for download, but the AI engine cites the HTML page — which can then direct human readers to the PDF. This approach preserves the authority and formality of the PDF format while ensuring AI discoverability.
Practical recommendations
Default to HTML. Unless there is a specific reason to use PDF, publish content as HTML pages. This is the format AI engines cite.
Create HTML companions for existing PDFs. If you have PDFs that represent valuable, citeable content, create HTML summary pages that capture the key points and link to the full PDF.
Don’t gate your PDFs unnecessarily. If a PDF requires a form fill or login to access, AI crawlers will not see it. Make the PDF directly accessible if you want it to be cited.
Treat PDFs as a supplementary format. For formal documents (reports, white papers, specifications), offer both a PDF download and an HTML version. The HTML version earns the citation.
Methodology
Computed from AmICited’s 1,905 tracked prompts (June 24, 2026 – July 23, 2026, 2026). A response counts as citing a PDF when a cited source URL path ends in .pdf; everything else is treated as an HTML/written page. Rates are citing-responses over total responses per engine; the format label is derived entirely from the cited URLs in the data — no page fetching. Copilot is tracked but returned no citation data. No external links appear in this report.
