How do answer engines choose between two similar pages?
When two pages cover the same ground, the tiebreakers are extractability, specificity, corroboration and attribution — in roughly that order.
Whichever page makes it easier to give a correct, attributable answer. In practice: the one with a self-contained passage that answers the exact question, from an identifiable source, saying something that other sources corroborate rather than contradict.
The tiebreakers, in order
1. Extractability
Can a span of text be lifted and still make sense? This dominates, and it is structural rather than qualitative.
- Wins: a heading that asks the question, with a direct answer in the first sentence or two beneath it.
- Loses: an answer assembled across several paragraphs, or one that depends on a preceding example to make sense.
- Loses badly: an answer that only exists in a table, an image, or an accordion that has to be opened.
A page can be the better piece of writing and lose here, which is uncomfortable but worth accepting.
2. Specificity
A page answering one question beats a page covering the topic that contains the answer somewhere. This is close to the opposite of what ranking has historically rewarded, and it is the adjustment most publishers find hardest.
The practical form: a comprehensive guide with clear question-shaped subheadings gets its individual sections cited. The same guide written as continuous prose does not.
3. Corroboration
Models weigh whether other sources agree. A claim appearing nowhere else is risky to repeat, however true — which is a real disadvantage for genuinely original work, and the reason to publish your evidence alongside your conclusion. A claim with visible reasoning behind it can be evaluated on its own terms.
4. Attribution
A named author, an identifiable organisation, a date. Not because these are ranking factors but because they let a system cite responsibly. Anonymous, undated content is harder to attribute and therefore gets attributed less.
What matters less than people expect
| Factor | Reality |
|---|---|
| Word count | Not a factor. Long pages get cited for their sections, not their length. |
| Keyword density | Not a factor, and has not been one anywhere for many years. |
| Domain age | Weaker here than in search — the page in front of the system carries more weight |
| Backlinks | Indirect at best. They influence what gets indexed rather than what gets extracted. |
| Publication recency | Matters for time-sensitive questions, irrelevant for stable ones |
| Page speed | Effectively irrelevant to citation. A crawler is patient. |
Making a page win the tiebreak
- Find the specific question your page answers, phrased as someone would actually ask it.
- Make that phrasing a heading.
- Answer it in one or two sentences immediately beneath, before any elaboration.
- Elaborate afterwards for the reader who wants more.
- Add something nobody else has — a measurement, a number, a first-hand detail.
- Name the author and the organisation, and date it.
Steps two and three are usually a restructuring of what you already wrote rather than new work, and they are where nearly all the gain is.
The competitive check
Rather than theorising, look. Ask several assistants the question your page answers and read the pages that get cited. The pattern is normally obvious within three or four examples, and it is more informative than any general advice — including this page.
What our audit reports about this
Every item below is measured directly, not inferred. Run it against your own site and the result names the exact rule or header responsible.
- Whether headings are question-shaped and whether an answer follows directly beneath each one.
- Whether content is in the delivered HTML, including anything inside accordions or tabs.
- Whether author, organisation and date are present and machine-readable.
- Whether structured data gives an engine an unambiguous account of what the page is and who published it.
For agents and scripts, the same measurement is at
/api/v1/ai?url=yoursite.com —
see the API documentation.
Related questions
Does length help or hurt?
Neither directly. Long pages get cited for individual sections, so the useful move is clear question-shaped subheadings with direct answers — which makes a long page into many extractable ones.
Do I need original research to be cited?
It helps considerably, and it does not have to be a study. A specific number, a first-hand observation, a worked example nobody else published — any of these makes you a primary source for something.
Does an FAQ section help?
Yes, when the questions are ones people genuinely ask and the answers are direct. It is close to the ideal extractable structure. It does not help when it is padding invented to have an FAQ.
Will restructuring hurt my search rankings?
Clear headings with direct answers followed by elaboration is good for both. The risk is stripping out depth to make a page more extractable — keep the substance, and lead with the answer.
Read next
How do I make a page quotable by AI?
The structural changes that make a passage extractable — most of them a rearrangement of what you have already written.
ReadWhat is the difference between crawling, indexing and citing?
Three separate stages, each with its own failure mode — and why fixing the wrong one wastes months.
ReadHow do I write content that AI assistants will quote?
Assistants extract passages, not pages. Everything follows from that one fact.
ReadDoes schema markup help with AI search?
What structured data actually does for answer engines — and why it is worth adding for a reason other than the one usually given.
Read