Ranking and being quoted are different jobs. A search engine returns ten links and lets the reader choose. An answer engine writes a paragraph and names two or three sources it leaned on. Getting into that paragraph is a different exercise from getting into the ten links, and most of the advice written about it is the old advice with new words on top.
How do answer engines choose what to cite?
Broadly, three things. The engine needs to find a passage that answers the question directly, it needs to be able to lift that passage without ambiguity, and it prefers sources that other sources already reference. The first two are about how you write. The third is the same authority problem search has always had.
Most of these systems work by retrieving a set of documents and then writing from them. That matters, because it means the unit of competition is the passage, not the page. A brilliant two thousand word article whose answer is spread across six paragraphs is worse, for this purpose, than a plain one that states the answer in the first two sentences under a heading that matches the question.
What should a page look like if you want it quoted?
A question as the heading, the answer in the first two sentences under it, then the reasoning. Not a build up, not a preamble about how important the topic is. If a machine has to read four paragraphs to work out what you claim, it will pick a page where it did not have to.
- Headings phrased the way people ask How long does SEO take to work, not Timelines and expectations. The match between the question and the heading is doing real work.
- The answer first, in one or two sentences Self contained, so that lifting it out of the page loses nothing. Test it by reading the sentence alone and asking whether it still makes sense.
- Specifics that can be checked Numbers, ranges, dates, named tools. Vague copy gets paraphrased into nothing and never attributed.
- One idea per section A section covering four things cannot be retrieved cleanly for any one of them.
- Content in the HTML, not behind JavaScript Several of these crawlers do not execute JavaScript at all. If the answer renders client side, it does not exist.
- A date, an author and a source Recency and attribution both feed into which of several equivalent passages gets chosen.
Does structured data help?
It helps for the mechanical part. FAQPage and Article markup make the question, the answer, the author and the date unambiguous instead of inferred. It does not make a weak passage citable, and marking up content that is not on the page is a way to get ignored rather than a shortcut.
The pragmatic list: Article or BlogPosting with a real datePublished and dateModified, FAQPage where you genuinely have questions and answers, Organization on the site so the entity is defined once, and breadcrumbs so the hierarchy is legible. That is most of the value. Everything past it is refinement.
Why does what other sites say about you still matter?
Because retrieval is not the only step. When several sources could answer a question, the ones that get named are the ones the system has more corroboration for, and corroboration comes from being mentioned elsewhere. Being the only site in the world that says something is not a strong position, it is an unverifiable one.
This is why the answer to how do I get cited by ChatGPT is not purely on-page. Placement on publications that already cover your subject does two things at once: it gives a search engine a link, and it gives an answer engine a second source that agrees with you. The mention matters even where the link does not, which is a genuine difference from classic link building.
Where those mentions live is worth thinking about specifically. Community threads, comparison round-ups and industry directories are disproportionately represented in what these systems retrieve, because they are structured, opinionated and constantly updated. A page on a real publication that lists you among five options is doing more for your citation rate than a press release nobody read.
Is llms.txt worth adding?
It costs nothing and no major engine has committed to reading it. Treat it as a bet with a small stake rather than a tactic, and do not let it displace the work that is known to matter. The same applies to any file or tag proposed as a way of talking directly to model crawlers.
What is worth doing today is deciding which crawlers you allow. GPTBot, ClaudeBot, PerplexityBot and Google-Extended are all controllable in robots.txt. If you want to be quoted, allow them explicitly rather than relying on a wildcard, and be aware that blocking them is a real choice with a real cost, not a safety measure.
How do you measure any of this?
By asking. There is no console for citations, so the working method is to keep a fixed list of twenty or thirty questions your buyers ask, run them across the engines on a schedule, and record which sources get named. It is manual, it is repeatable, and it is the only thing that tells you whether the work moved anything.
Track three numbers per question: whether you were mentioned at all, whether you were linked, and which competitors were named instead. The third one is the most useful, because it tells you exactly which pages to go and read. Our AEO and GEO audit does this across ten engines and re-runs free inside six months if the fixes we specified did not move the citation rate.