# GeoSalience — Cornerstone library (full text)
> Generative Engine Optimization, primary research only.
This file aggregates the canonical articles from GeoSalience. Citations welcome — please link back.
Source: https://geosalience.com/llms-full.txt | Updated: 2026-06-08T00:00:00.000Z
---
# State of GEO Q2 2026: the AI engine you optimize for matters most
> We probed 100 brands across ChatGPT, Claude, Gemini, and Perplexity — 4,000 prompts. Citation rate varies far more by which engine you ask than by which brand you are: Perplexity cites 42%, ChatGPT 7%, Gemini hides its sources, and legacy news publishers get 0%.
Published: 2026-06-07T00:00:00.000Z
Updated: 2026-06-07T00:00:00.000Z
Canonical: https://geosalience.com/measurement/state-of-geo-q2-2026
We asked four AI engines — ChatGPT, Claude, Gemini, and Perplexity — the same ten questions about each of 100 brands, then recorded which engine cited each brand's own website. The clearest finding of the **4,000-prompt** run, dated **2026-06-07**, is not about any brand. It's about the engines: **citation rate varies far more by which engine you ask than by which brand you are.** Perplexity cited the target brand's domain on **42%** of prompts; ChatGPT on **7%**; Gemini grounds almost everything but returns opaque redirect URLs so you can't tell who it cited; and across all of them, the legacy news publishers whose journalism trained these models — NYT, the BBC, Bloomberg, Reuters — were cited **0%** of the time.
**The one-line takeaway:** in mid-2026, "optimizing for AI search" mostly means optimizing for the *engine*, not in the abstract. The same brand can be a first-class citizen in Perplexity and invisible in ChatGPT.
## Read this first: coverage and what broke
This is a v1, and an honest one. Of the 4,000 probes, **2,876 returned a usable answer** and 1,124 errored — and *errors are excluded from every rate below* (an API failure means "not measured," never "not cited"). The errors were not random, so per-engine coverage is uneven and you should read each engine's number with its sample size:
| Engine | Valid probes | Grounding (searched) | Citation rate | Note |
|---|---|---|---|---|
| **Perplexity** | 1,000 / 1,000 | 100% | **41.6%** | Clean, full 100-brand coverage — the backbone of this report |
| **Claude** (Sonnet) | 238 / 1,000 | 93% | **32.8%** | A 24-brand sample — our Anthropic credits ran out mid-run; treat as directional |
| **ChatGPT** | 650 / 1,000 | 23% | **6.9%** | Rarely searches at all (so rarely cites); 350 calls hit OpenAI rate limits |
| **Gemini** | 988 / 1,000 | 99% | **0% (opaque)** | Grounds heavily but returns `vertexaisearch` redirect wrappers — publisher hidden |
Because Perplexity is the only engine with clean, uniform 100-brand coverage, **the brand leaderboard below uses Perplexity's citation rate** — a consistent ruler across all 100 brands. The other engines appear in the cross-engine section, where their job is comparison, not ranking. This mirrors what we found turning the same lens on [our own site](/case-studies/geosalience-as-its-own-case-study) and on [who LLMs cite for GEO](/measurement/who-llms-cite-for-geo): Perplexity and Claude do the real retrieving; ChatGPT mostly answers from memory; Gemini hides its sources.
## The brand leaderboard (Perplexity citation rate)
Across 100 brands, Perplexity cited the brand's own domain on a mean of **41.6%** of prompts (median 40%). The spread is wide: **43 brands were cited on at least half** their prompts, while **11 brands were never cited once**.
| # | Brand | Sector | Perplexity citation rate |
|---|---|---|---|
| 1 | TechCrunch | Media / Publishers | 100% |
| 2 | Brex | Finance / Fintech | 90% |
| 2 | Databricks | SaaS / B2B | 90% |
| 2 | Salesforce | SaaS / B2B | 90% |
| 2 | Talkspace | Healthcare / Wellness | 90% |
| 6 | Atlassian | SaaS / B2B | 80% |
| 6 | MongoDB | SaaS / B2B | 80% |
| 6 | Stratechery | Media / Publishers | 80% |
| 6 | Wealthfront | Finance / Fintech | 80% |
The full ranking for all 100 brands is in the [dataset](https://geosalience.com/datasets/state-of-geo.json).
## The sector gradient
Grouping the Perplexity leaderboard by sector, the practitioner-heavy categories win and media loses:
| Sector | Mean Perplexity citation rate |
|---|---|
| SaaS / B2B | 49.6% |
| Finance / Fintech | 46.0% |
| Healthcare / Wellness | 42.0% |
| AI / ML companies | 42.0% |
| Travel / Hospitality | 41.0% |
| Consumer e-commerce / DTC | 38.7% |
| Education / EdTech | 38.0% |
| Media / Publishers | 21.0% |
SaaS and fintech brands — the ones with dense, structured documentation and a heavy developer/marketing content footprint — are cited roughly **2.4× more often than media publishers.**
## The media paradox
The single most counter-intuitive result sits inside that bottom row. The publishers whose archives plausibly *trained* these models are the ones the models won't cite:
- **TechCrunch: 100%. Stratechery: 80%. The Information: 30%.**
- **NYT, Washington Post, BBC, Bloomberg, Reuters, Wired, The Verge: 0%.**
Every legacy or paywalled publisher in the set scored zero Perplexity citations; only the tech-native, openly-readable outlets surfaced. The likeliest explanation is structural, not editorial: paywalls and aggressive `robots`/bot policies keep the retrieval layer out, so the engine reaches for whatever it *can* fetch — which is the open tech press. Whatever the cause, the lesson for a publisher is blunt: training on your content is not citation, and if the live crawler can't reach the page, you don't exist in the answer.
## How much the engine matters: divergence
For brands where we have more than one engine's data, the gap between engines dwarfs the gap between brands. Two examples from the [dataset](https://geosalience.com/datasets/state-of-geo.json):
- **Brex** — Perplexity 90%, ChatGPT 10%.
- **Wealthfront** — Perplexity 80%, ChatGPT 10%.
A brand can be nearly always cited by one engine and nearly never by another. That is the practical core of GEO in 2026: there is no single "AI visibility" number. ChatGPT's low rate is mostly because **it searches on only ~23% of prompts** — when it answers from training memory, no one gets cited. Gemini's zero is an artifact of its API hiding the publisher behind a redirect, not evidence that it cites no one. Perplexity, which searches every time and returns real URLs, is where citation is both highest and measurable.
## A note on sentiment
We also ran a crude keyword-based sentiment proxy on the "is X any good / what are its weaknesses" prompts — disclosed as rough (a v2 would use a proper classifier). The only pattern worth stating at this confidence: **telehealth brands skewed negative** (Hers, Ro), consistent with the hedging LLMs apply to health topics, while education and open-model brands skewed positive (Codecademy, Mistral). Treat these as directional, not definitive.
## How we tested
The method reuses [our citation harness](/lab): for each of 100 brands (8 sectors, see the [dataset](https://geosalience.com/datasets/state-of-geo.json)), ten prompts — direct ("tell me about X"), category ("best $category in 2026"), comparison ("X vs competitor"), and sentiment — were sent to ChatGPT, Claude, Gemini, and Perplexity with web grounding enabled, on 2026-06-07. **Citation = the brand's own registrable domain appears in the engine's returned sources** (an unambiguous signal, and the same one we use for [our own citation rate](/glossary/citation-rate)). Every figure here is produced by an open analysis script over the captured probes; nothing is estimated. Probes ran from a single EU location.
## Limitations
This is the first quarterly State of GEO, and it is bounded:
- **Uneven engine coverage.** Perplexity is complete (100 brands); Claude is a 24-brand sample (Anthropic credits ran out mid-run, on the cheaper Sonnet model); ChatGPT lost 350 probes to rate limits; Gemini's sources are opaque. The leaderboard is therefore Perplexity-anchored, and cross-engine numbers carry their sample sizes. A v2 re-runs all four to full coverage.
- **Single run, 10 prompts, one day, one region.** A snapshot, not a trend; the value compounds when Q3 gives the first delta.
- **Citation, not ranking quality.** We measure whether the brand's domain was cited, not whether the answer was good or the brand was described accurately.
- **Crude sentiment proxy** (keyword-based) — directional only.
- **Prompt-set and brand-set bias.** A different 100 brands or 10 prompts would move the numbers.
The full per-brand, per-engine data — including every brand's coverage so you can see exactly what's measured — is the [public dataset](https://geosalience.com/datasets/state-of-geo.json). This sits in the [measurement pillar](/measurement) alongside our other quantitative work; if you're new to the metric, start with [what GEO is](/foundations/what-is-geo) and the [citation rate](/glossary/citation-rate) definition.
## Frequently asked questions
**Why is the leaderboard only Perplexity?**
Because it's the only engine that searched every prompt and returned real URLs, giving uniform, clean coverage across all 100 brands. Ranking brands on an engine with partial coverage would compare them on different sample sizes. The other engines are reported in aggregate, with their coverage stated.
**Does Gemini really cite nobody?**
No — Gemini grounds on ~99% of prompts. But its API returns `vertexaisearch.cloud.google.com` redirect wrappers instead of the publisher URL, so a brand-domain citation can't be observed. We report that as opacity, not as zero citations. (We hit the same wall measuring [who LLMs cite for GEO](/measurement/who-llms-cite-for-geo).)
**Can I reproduce this?**
Yes. The [dataset](https://geosalience.com/datasets/state-of-geo.json) has every brand's per-engine rate and coverage; the brand list, prompt set, and analysis code are in [our repo](https://github.com/geosalience/geosalience). Re-run it and you should reproduce these numbers from the captured probes.
---
# This site is our GEO lab: the stack, the data, the experiments
> GeoSalience measures its own AI-crawler traffic, tracks whether LLMs cite it, and runs controlled experiments on its own pages. Here are the first real numbers: an 898-request crawler footprint, a measured 0% citation baseline, and an honest account of where data is still thin.
Published: 2026-05-31T00:00:00.000Z
Updated: 2026-06-07T00:00:00.000Z
Canonical: https://geosalience.com/case-studies/geosalience-as-its-own-case-study
GeoSalience is a publication about Generative Engine Optimization, so it has to be the best live example of GEO we can build — and a place where we test GEO methods on ourselves instead of describing them in the abstract. That means measuring our own AI-crawler traffic, tracking whether large language models cite us, gating every page on a GEO checklist, and running controlled before/after experiments on our own pages. Operating the site is, itself, our primary research.
This page is the running account of that lab. It is dated, it uses only numbers we have actually measured, and where a measurement is still thin it says so rather than guessing. The launch-day section below is preserved as history; the **[2026-06-07 update](#update--2026-06-07-the-first-real-measurements)** carries the first real numbers from the live pipelines.
**Key takeaways (as of 2026-06-07):**
- The site went live on **2026-05-31** and now runs **5 live articles** plus 6 glossary terms, with a full GEO surface — `llms.txt`, `.md` aliases, JSON-LD, canonical URLs — verified returning 200 over HTTPS.
- We track four things no off-the-shelf analytics tool reports for us: AI-crawler hits, AI **citation rate** (our North Star), a self-audit GEO gate, and on-site experiments. All four are now **live and streaming** — not pending.
- **The crawler and citation pipelines have produced their first real data.** In the first week (2026-05-31 → 2026-06-07) we logged **898 AI-crawler requests from 18 distinct crawlers**, and our first citation run measured **0% (0 of 50 prompts)** — a real, honest day-zero baseline, not a placeholder. See the [citation rate](/glossary/citation-rate) definition for what we count.
- The first experiment, [exp-001](/lab/experiments/exp-001), records the day-zero baseline. Three more are [drafted and waiting](/lab/experiments) for a decision on which to run.
## Why a publication about GEO should experiment on itself
The thing that makes a claim about GEO credible is data nobody else can reproduce. "We added FAQPage schema to one page and its AI-crawl frequency changed by this much over 30 days" is a finding that needs our server logs and our measurement harness — a competitor cannot copy it. Every page we publish is a unit we control, which makes the whole site a sample frame.
This is the same logic behind open-metrics companies that published their own numbers and got linked and cited for it. Our version is a public lab: the citation rate, the crawler footprints, and the experiment log are all on [/lab](/lab), updated as the data lands.
## The stack, component by component — and its real status
The honest part of a build-in-public report is the status column. Here is the same stack as on launch day, with the status column brought current to **2026-06-07** — every "pending" replaced by a real, dated state.
| Layer | What it measures | Status on 2026-06-07 |
|---|---|---|
| GEO surface (`llms.txt`, `.md` aliases, JSON-LD, canonical, sitemap) | discoverability + machine-readability | **Live** — verified 200 over HTTPS |
| AI-crawler tracking (nginx log → daily dataset) | which LLM bots read us, how often, which pages | **Live** — 898 requests / 18 crawlers logged, first bots seen 2026-05-31 |
| Citation-rate harness (50 prompts × 4 LLMs) | our North Star: do LLMs cite us | **Live** — baseline 0% (0/50) on 2026-06-07, prompt set v1 |
| Self-audit GEO gate (40-point Playbook) | every page obeys our own GEO rules | **Automated** — `pnpm verify:geo` runs in the pre-deploy gate |
| Experiment framework + public log | controlled before/after tests on our pages | **Live** — see [/lab/experiments](/lab/experiments) |
| Traffic (privacy-friendly analytics) | human visitors, referrers | **Live** — Plausible self-hosted, first-party (early/low traffic) |
Two of these — crawler tracking and citation rate — are the GEO-native metrics that make the lab worth running. Both now produce dated data we publish under [`/datasets`](https://geosalience.com/datasets/citation-rate.json); the next two sections report the first real numbers and the caveats that go with them.
## The launch-day baseline (the real numbers)
You can only prove a curve moved if you wrote down where it started. On **2026-05-31** the verifiable state was:
- **Corpus:** 2 live articles — [What is GEO?](/foundations/what-is-geo) and [The llms.txt spec: adoption and setup](/technical/llms-txt-spec-adoption-setup) — plus 6 glossary terms. Ten further articles exist as drafts, gated on primary research and hidden from listings.
- **Crawler policy:** `robots.txt` welcomes the major AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended and others), and Cloudflare runs DNS-only, so those bots reach the origin directly instead of being filtered at the edge.
- **Quality:** the cornerstone pages score 99/100/100/100 on Lighthouse (performance / accessibility / best-practices / SEO).
- **One published dataset already:** our [llms.txt adoption study](/technical/llms-txt-spec-adoption-setup) found **37 of 100** surveyed domains served a valid `llms.txt` when we crawled them on 2026-05-19. That is real primary research about other sites; the lab described here turns the same lens on ourselves.
On launch day the two GEO-native numbers — total AI-crawler hits and citation rate — were recorded as **pending first run**, because we do not write down a measured number we did not collect. [exp-001](/lab/experiments/exp-001) is the formal record of that baseline. The next section replaces those two "pendings" with real measurements.
## Update — 2026-06-07: the first real measurements
One week after launch the crawler parser and the citation harness have each produced their first dated run. Both numbers are public datasets; neither is rounded or spun.
### AI-crawler footprint: 898 requests from 18 distinct crawlers
Between **2026-05-31 and 2026-06-07**, our nginx logs recorded **898 requests from 18 distinct AI-related crawlers** ([dataset](https://geosalience.com/datasets/ai-crawlers.json)). The traffic is concentrated, and the mix is more interesting than the total:
| Crawler | Operator · type | Requests | First seen |
|---|---|---|---|
| GoogleOther | Google · mixed (training/research) | 395 | 2026-06-05 |
| ClaudeBot | Anthropic · training | 249 | 2026-05-31 |
| Googlebot | Google · search | 102 | 2026-05-31 |
| bingbot | Microsoft · search | 24 | 2026-06-01 |
| Applebot | Apple · search | 24 | 2026-06-04 |
| Google-Extended | Google · training | 15 | 2026-06-06 |
| Meta-ExternalAgent | Meta · mixed | 14 | 2026-06-06 |
| CCBot | Common Crawl · training | 12 | 2026-06-06 |
| ChatGPT-User | OpenAI · on-demand | 11 | 2026-06-06 |
| PerplexityBot | Perplexity · search | 10 | 2026-06-06 |
| GPTBot | OpenAI · training | 8 | 2026-06-06 |
The remaining seven (Bytespider, OAI-SearchBot, Amazonbot, anthropic-ai, cohere-ai, Claude-User, Perplexity-User) trail in single digits. The honest reading: a brand-new GEO site is crawled within hours, but early volume is dominated by **training and generic fetchers** (GoogleOther, ClaudeBot), while the **on-demand "answer" fetchers** that signal a live citation (ChatGPT-User, Perplexity-User) are still a trickle. That gap is exactly what the citation rate measures next.
### Citation rate: a measured 0% (0 of 50) — and why that number is credible
Our first citation run, on **2026-06-07** with prompt set **v1** (50 prompts across the four GEO pillars, run against ChatGPT, Claude, Gemini and Perplexity), returned **0% — 0 of 50 prompts cited or mentioned geosalience.com** ([dataset](https://geosalience.com/datasets/citation-rate.json)). Per provider, every rate is 0%.
A zero is only worth publishing if you can show it is real rather than an empty pipeline. Ours is real: across the run the four models grounded their answers in **1,105 source URLs** (870 unique) — and **not one** pointed to geosalience.com. The models were answering GEO questions with sources; we simply were not among them yet. That is the honest starting height, and it matches the crawler picture above: we are being read for training, not yet retrieved for answers. This is the baseline the [citation rate](/glossary/citation-rate) curve climbs from.
We turned those same captured sources into a first real finding: [who LLMs actually cite for GEO](/measurement/who-llms-cite-for-geo) maps the 927 attributable citations to 417 domains — led by YouTube and SEO-tool blogs, with geosalience.com nowhere in them yet.
### Methodology caveat: grounding coverage was uneven
The 0% headline is a confident "no LLM cited us." But per-provider **retrieval coverage was not uniform**, and that has to be stated plainly. Web grounding actually fired on:
- **Perplexity — 50/50** prompts (always searches),
- **Claude — 35/50** (searches at the model's discretion),
- **Gemini — 13/50** (a free-tier `429` quota capped the run partway through),
- **ChatGPT — 3/50** (searched only when the model chose to).
So while no provider cited us on any prompt where it *did* search, ChatGPT and Gemini searched on few prompts, so their slice of the baseline is thin. We run grounding at model discretion on purpose — it mirrors how real users get answers — but a stricter follow-up run (forced search, a higher-quota Gemini key) is a stated next step before we read too much into per-provider deltas. The 0% overall stands; the per-provider confidence varies.
## The self-audit gate
Every article on GeoSalience is meant to pass a 40-point [GEO Playbook](/methodology): answer-first opening, standalone-coherent sections, dense dated facts, inline primary-source citations, Schema.org markup, a working `.md` alias, internal links. As of 2026-06-07 that gate is **automated**: `pnpm verify:geo` runs the machine-checkable subset of the Playbook against every `state: live` page in the pre-deploy step, and refuses the deploy if any page drops below its threshold (cornerstone 100%, spoke ≥85%, pulse ≥60%). A page that breaks our own rules cannot ship. The tool is deliberately honest about its limits — it never claims a page is "GEO-perfect," only that it has not regressed on the checkable subset; the research-quality points (A and B of the Playbook) still need a human. At this update every live page passes its threshold — the same gate had to pass for this very cornerstone to ship, and the running self-compliance figure is on [/lab](/lab).
## The experiments
The framework treats each experiment as a citable page with a protocol, a status, and — only at conclusion — a result. The analysis is deliberately modest: it reports descriptive before/after deltas with explicit caveats and **computes no statistical significance**, because a handful of pages with noisy, slow metrics does not justify it. When the data is too thin to say anything, the analyzer returns *inconclusive* and says why. We publish that as readily as a positive result.
- [exp-001 — Baseline established](/lab/experiments/exp-001): concluded. The day-zero snapshot above.
- Three candidates are drafted and **not started**, pending a decision on which to run and how to pair control and treatment pages: a [FAQPage schema test](/lab/experiments/exp-002), an [answer-first rewrite](/lab/experiments/exp-003), and an [llms-full.txt inclusion test](/lab/experiments/exp-004).
One limit is worth stating here because it shapes what these experiments can show: our citation harness measures the **whole domain**, not individual pages. So a page-level citation experiment cannot be cleanly attributed, and the framework says so — it falls back to per-page crawler re-fetch frequency as a proxy and treats citation movement as site-wide context. Sharpening that into true per-page citation attribution is on the list.
## How we test (disclosure)
The methodology here is the dogfooding loop itself: publish a page, audit it against the GEO Playbook, measure which AI crawlers read it, measure whether LLMs cite the site, change one variant and measure the delta, then publish the finding — which becomes another page in this lab. We disclose that the site is our own subject. Numbers come from our nginx logs (bot rows only, no visitor PII) and from our own harness querying public LLM APIs with our own prompts.
The citation harness runs a locked **50-prompt set (v1)** — spread across the four GEO pillars and the same intents we write about, from "what is GEO" to "[GEO vs AEO vs LLMO](/foundations/geo-vs-aeo-vs-llmo-vs-sge)" — against four models (ChatGPT, Claude, Gemini, Perplexity). Grounding is left at **model discretion**, not forced, so the measurement reflects how a real user gets an answer; the trade-off is the uneven per-provider coverage disclosed in the [2026-06-07 update](#methodology-caveat-grounding-coverage-was-uneven) above. Datasets are published under `/datasets`; the experiment log is at [/lab/experiments](/lab/experiments).
## Frequently asked questions
**Is the citation rate really 0% right now?**
Yes. On 2026-06-07 our harness measured **0% — 0 of 50 prompts** cited or mentioned geosalience.com across all four models. It is a measured zero, not a guess: the models grounded those answers in **1,105 source URLs and none were ours**. The expected first value was always near zero, and that is fine — the point is the curve, not the starting height. Track it on [/lab](/lab).
**Doesn't experimenting on your own site bias the results?**
Yes, and we say so. These are quasi-experiments on a small, single-domain corpus, not randomized controlled trials. We disclose the confounders — page-type differences, tiny samples, site-wide-only citation data, and the uneven grounding coverage noted above — on every experiment.
**Where can I check these claims?**
Every figure is public: fetch [`/robots.txt`](https://geosalience.com/robots.txt), [`/llms.txt`](/llms.txt), or any article's `.md` alias; open the [citation-rate dataset](https://geosalience.com/datasets/citation-rate.json) and the [crawler dataset](https://geosalience.com/datasets/ai-crawlers.json); read [exp-001](/lab/experiments/exp-001); follow the [citation rate](/glossary/citation-rate) definition. Nothing in the baseline depends on private data. If you are new to the terms, start with [what GEO is](/foundations/what-is-geo) and [how knowledge cutoffs and live web access interact](/foundations/knowledge-cutoff-and-web-access).
## Limitations
This is still a young lab, but the gaps have shifted. Resolved since launch: both GEO-native metrics now stream (898 crawler requests; a measured 0% baseline), and the self-audit gate is automated in the deploy step. Still open, and stated plainly:
- **Per-page citation attribution does not exist** — the harness measures the whole domain, so a page-level citation experiment falls back to crawler re-fetch frequency as a proxy.
- **Grounding coverage is uneven** — ChatGPT searched on 3/50 prompts and Gemini on 13/50 (a quota cap) this run, so per-provider numbers carry less weight than the overall figure until a stricter run lands.
- **The corpus is small** — five live articles is a thin sample frame, and traffic is early/low.
Each is a reason a later entry here will be worth reading. This is, for now, the only entry in our [case studies](/case-studies) pillar; the next one will be a delta, not a baseline.
---
# llms.txt: Spec, 100-Domain Adoption Audit, and Setup
> We audited 100 top developer-tools and SaaS sites for an llms.txt file. Only 37 of them serve one at the apex — and the gap is concentrated in the places you might expect it not to be. The full spec, the audit, and a 10-minute setup.
Published: 2026-05-25T00:00:00.000Z
Updated: 2026-05-19T00:00:00.000Z
Canonical: https://geosalience.com/technical/llms-txt-spec-adoption-setup
- **37 of 100** audited top sites serve `llms.txt` at the apex — above the 18–25 prior.
- **Top SaaS leads at 14/20 (70%)**; news outlets and global brands ship none.
- **Mean spec compliance** among adopters is **0.84**. Stripe, Together AI, Webflow, Athena HQ, and Turborepo all score 1.00.
- **Dominant failure mode:** no typed links under `## Docs` / `## Optional` — the file ships an H1 and sections but no link list, which defeats its purpose for LLM consumers.
- **Setup is one route handler + raw-markdown alternates**, about 10 minutes for a typical SaaS.
`llms.txt` is a proposal by Jeremy Howard (Answer.AI) for a small text file at the root of your site that helps LLMs read it. It is _not_ a standard. It is _not_ enforced by any crawler today. And yet — a meaningful slice of the AI-native sites we audited have one, and the ones that don't are leaving money on the table for an investment of about 10 minutes.
This is the full spec, the 100-domain audit, and a setup tutorial you can copy.
## What is llms.txt
`llms.txt` is a curated, machine-readable index of your site, structured as Markdown, served at `https://yourdomain.com/llms.txt`. It tells an LLM crawler:
1. The name of your publication / project.
2. A one-line description.
3. A curated list of canonical URLs, grouped by topic, with one-line descriptions.
4. Optional sections like "About", "FAQ", "API reference".
There is a companion file, `llms-full.txt`, which contains the _full text_ of the most important pages, concatenated.
The point: instead of a crawler walking your sitemap and ranking every page by guesswork, you hand-curate what matters and serve it as a single file.
## Why it matters (even though no crawler enforces it yet)
- **Voluntary signal.** Engineers at OpenAI, Anthropic, and Perplexity have all referenced reading `llms.txt` files during product development. It's used today as a research aid, even if it's not crawled at scale.
- **It's an editorial statement.** A site that publishes a curated index is signalling something about how it wants to be read. That's a useful signal _for humans_ too, in particular journalists and competitors.
- **Cost: zero.** You're already maintaining a sitemap. This is one more file.
- **Forward compatibility.** If a major crawler does adopt the format, you're already ready.
## The spec, in 90 seconds
A minimal valid `llms.txt`:
```markdown
# My Project Name
> One-sentence description of what this site is.
A short paragraph explaining the project, audience, and scope.
## Docs
- [Quickstart](https://example.com/quickstart): Get started in 5 minutes.
- [API reference](https://example.com/api): Full API documentation.
## Optional
- [Changelog](https://example.com/changelog): Recent releases.
```
Rules:
- File is **Markdown** (not plain text, despite the `.txt` extension).
- First line is `# Project Name` (H1).
- Second line is `> Short description` (Markdown blockquote).
- The next paragraph is a longer description.
- After that, `## Section Name` headers separate logical groups.
- Inside each section: a `- [Title](URL): description.` line per resource.
- "Optional" section is a convention for secondary links.
The companion `llms-full.txt`:
- Same format on top.
- Each H1 is the title of an article.
- Body of the article in raw Markdown follows.
- `---` separates articles.
## Audit: 100 domains
We crawled 100 domains on 2026-05-19 across eight categories: AI/ML companies (15), top SaaS (20), dev tools and frameworks (15), documentation platforms (10), tech media (10), top global brands (10), news and journalism (10), and GEO/SEO tools (10). The full domain list and the crawler that ran the audit are open: see [the methodology dataset](/datasets/llms-txt-audit.csv) and `scripts/llms-txt-audit.ts` in [our repository](https://github.com/geosalience/geosalience).
The audit probes the **apex domain only** — `https://.com/llms.txt`. A site that ships the file under a subdomain (`docs.example.com/llms.txt`) or a deeper path is recorded as "not found." That apex-only frame matters: it's the location the spec proposes, and the one an LLM crawler would try first without prior knowledge.
### Adoption rate
| Category | Audited | With `llms.txt` | Adoption rate |
|---|---|---|---|
| AI / ML companies | 15 | 5 | 33% |
| Top SaaS | 20 | 14 | 70% |
| Dev tools and frameworks | 15 | 7 | 47% |
| Documentation platforms | 10 | 3 | 30% |
| GEO / SEO tools | 10 | 3 | 30% |
| Tech media | 10 | 1 | 10% |
| Top global brands | 10 | 0 | 0% |
| News and journalism | 10 | 0 | 0% |
| **Total** | **100** | **37** | **37%** |
The headline is the SaaS row: **70% of the top SaaS sample serves an `llms.txt` at the apex**, a higher rate than AI/ML or dev tools and a much higher rate than we expected before running the crawl. The cleanest gap is at the other end: news, global brands, and the bulk of tech media ship nothing.
Among the 37 adopters, 8 also serve the companion `llms-full.txt` (22% of adopters; 8% of the full sample). The companion file is concentrated in dev tools and docs platforms — Cloudflare, Supabase, Bun, Drizzle, tRPC, Vite, Turborepo, Speakeasy.
A few well-known names are missing from the apex adoption list because they redirect their apex to `www..com`, which 404s on `/llms.txt`. The most visible of these are Anthropic and Mintlify — both ship documentation surfaces that publish `llms.txt`, but not at the apex this audit probes. That's a finding about where the file lives, not whether it exists; we treat it as a separate question for a follow-up audit.
### What's good (exemplary files)
A short list of `llms.txt` files that get the format right, ranked by our compliance score (sum of seven dimensions / 7) then by file size.
- **Stripe** (`stripe.com/llms.txt`, ~64 KB body, score 1.00) — the largest well-formed file we found. The crawler truncated the body at 64 KB; the live file is longer still. Sections cover API products and SDKs; every link in the file is a curated entry, not a generic sitemap dump.
- **Together AI** (`together.ai/llms.txt`, ~39 KB, score 1.00) — comprehensive coverage of model docs, inference endpoints, and platform features; clean H2 structure; `.md` alternates throughout.
- **Webflow** (`webflow.com/llms.txt`, ~13 KB, score 1.00) — a SaaS that ships `llms.txt` as an editorial signal, not just a docs concession. Concise H2 sections, descriptive one-liners per link.
- **Athena HQ** (`athenahq.ai/llms.txt`, ~12 KB, score 1.00) — one of three GEO-tool vendors that publish at the apex. Categorized sections with descriptions written for LLM consumers, not humans.
- **Turborepo** (`turbo.build/llms.txt`, ~10 KB, score 1.00) — a Vercel-stack project that uses the file as a curated index over the docs site. Clean H1 + summary + Docs/Examples sections; no marketing copy.
What these have in common: an H1 with the brand name, a single-sentence blockquote summary, 3-5 H2 sections (Docs, API Reference, Optional, sometimes Examples), `.md` alternates linked rather than HTML pages, and no marketing copy.
### Common mistakes
The crawler's compliance check scores each found file on seven dimensions
(see the [crawler spec](/raw/technical/llms-txt-spec-adoption-setup.md) for
detail). The **most common dimensional failure across the 37 found files
was "no typed links"** — the file ships an H1, a summary, and at least
one `##` section, but no `[text](url)` line under a `## Docs` or
`## Optional` heading. Sites in this group treat `llms.txt` as a
human-readable about-page instead of a curated link index, which defeats
the file's purpose for LLM consumers.
Beyond the dimensional check, the same five editorial mistakes recur
whenever we hand-inspect the files:
1. **Linking to HTML pages instead of `.md` alternates.** The whole point of `llms.txt` is to surface raw Markdown. Linking to `https://example.com/docs/quickstart` makes the LLM crawler walk back through HTML parsing — the very work the file was meant to remove. Link to `https://example.com/docs/quickstart.md` (or whatever your raw-Markdown route is) instead.
2. **Missing one-line descriptions on link items.** The spec says `- [Title](URL): description.` The colon and trailing description are not optional. Files that omit them give the LLM no hint about relevance, defeating the curation step.
3. **Stale links (404).** Curated indexes drift faster than sitemaps because they aren't usually regenerated by the build system. We expect to find dead links in 10-20% of audited files.
4. **Marketing copy in the description.** "The best CRM platform for modern teams" is a tagline, not a summary. The summary should describe what the linked page contains, not why you should buy from the vendor.
5. **Mixing first-party docs with marketing pages.** An `llms.txt` that links to `/pricing`, `/about`, and `/blog/announcement-x` is using the file as a sitemap. The spec is for an LLM-grounding index — pricing belongs on a sitemap, not here.
## Practical setup (10 minutes)
### Step 1 — Inventory
Decide what belongs in `llms.txt`. Not every page. Just the pages an LLM should ground on when answering questions about your site. For a typical SaaS:
- Quickstart / getting-started
- API reference
- Pricing
- Top 5–10 most-read blog posts (cornerstones)
- Methodology / how-it-works pages
- Changelog
### Step 2 — Write the index
In your site repo, create the file at the path your framework serves as the root. For Next.js app router, a complete production-quality `llms.txt` route looks like this:
```ts
// app/llms.txt/route.ts
import { articles } from '#site/content'; // Velite-typed export
export const dynamic = 'force-static';
export const revalidate = 3600;
const SITE = {
name: 'GeoSalience',
description: 'Primary research on Generative Engine Optimization.',
url: 'https://geosalience.com',
};
export async function GET() {
const cornerstones = articles
.filter((a) => a.cornerstone && a.state === 'live')
.sort((a, b) => +new Date(b.publishedAt) - +new Date(a.publishedAt));
const docs = cornerstones
.map((a) => `- [${a.title}](${SITE.url}${a.permalink}.md): ${a.dek}`)
.join('\n');
const optional = [
`- [About](${SITE.url}/about.md): Who runs this publication.`,
`- [Methodology](${SITE.url}/methodology.md): How we run our tests.`,
`- [Editorial policy](${SITE.url}/editorial-policy.md): Corrections, disclosure, right of reply.`,
].join('\n');
const body = `# ${SITE.name}
> ${SITE.description}
GeoSalience is an editorial publication producing primary research on how
LLMs decide which sources to cite. Every article is built on a dataset
we publish alongside it.
## Cornerstones
${docs}
## Optional
${optional}
`;
return new Response(body, {
headers: {
'Content-Type': 'text/markdown; charset=utf-8',
'Cache-Control': 'public, max-age=3600, s-maxage=86400',
},
});
}
```
Three things to notice in the above:
1. The route reads from the content layer (Velite, in our case) rather than hardcoding URLs. The file stays in sync with what's actually published without manual maintenance.
2. Every URL ends in `.md` — those are the raw-Markdown alternates from Step 3.
3. `Content-Type` is `text/markdown`, not `text/plain`. The spec accepts either, but `text/markdown` is more accurate and is what LLM crawlers explicitly look for.
GeoSalience's actual `llms.txt` route is open-source; see [the repo](https://github.com/geosalience/geosalience).
### Step 3 — Add a `.md` alternate for every linked URL
This is the step everyone skips. An LLM crawler does not want your hero animation and your cookie banner. Serve a raw-markdown version of every URL you link in `llms.txt`, at the same path with a `.md` suffix.
In Next.js you can do this with a middleware rewrite — see [our raw markdown route](/foundations/what-is-geo.md) for a working example.
### Step 4 — `llms-full.txt` (optional, recommended)
Concatenate the full Markdown body of each cornerstone article into one file. This gives an LLM with a generous context window everything it needs to ground on your content in a single fetch.
### Step 5 — Test
Fetch your `llms.txt` with `curl` and confirm:
- `Content-Type: text/markdown` (or `text/plain` — both are spec-compliant; `text/markdown` is preferred).
- File is valid Markdown — paste into any Markdown viewer to confirm headings render.
- All linked URLs return 200.
- The `.md` alternate URLs also return 200, with `Content-Type: text/markdown`.
A 30-second sanity check:
```bash
# Fetch llms.txt and verify content-type
curl -sI https://yourdomain.com/llms.txt | grep -i 'content-type'
# Extract every URL from the file and check status codes
curl -s https://yourdomain.com/llms.txt \
| grep -oE 'https?://[^)]+' \
| sort -u \
| xargs -I{} sh -c 'echo "$(curl -sI -o /dev/null -w "%{http_code}" {}) {}"'
```
If any URL returns 404 or 5xx, fix it before publishing. If a URL redirects (3xx), decide whether the redirect destination is what you actually want LLMs to see — if yes, link the destination directly.
### Step 6 — Publish and announce
This step is one we wish more sites would do: when you publish an `llms.txt`, tell people. Mention it in your changelog. Tweet the URL. Add a footer link `LLMs index` so humans can discover it too. The file itself is invisible to humans by design, but the cultural signal is worth surfacing.
## What llms.txt is not
- It is **not** a robots.txt for AI. Use `robots.txt` for crawl directives. `llms.txt` is descriptive, not restrictive.
- It is **not** a sitemap. A sitemap lists all your URLs. `llms.txt` is the curated, hand-picked index.
- It is **not** a standard. Treat it as a strong proposal that some teams find useful.
## FAQ
### Will my llms.txt actually be fetched by ChatGPT / Claude / Gemini?
There is no public claim from any of the major LLM vendors that they fetch `llms.txt` as a routine part of crawl. Anecdotally, Perplexity has referenced reading them, and Anthropic engineers have referenced consulting them during product development. Treat adoption as forward-looking, not a current channel. If a major crawler does adopt the format and you already ship the file, you're already ready.
### Does writing one help SEO?
No. `llms.txt` is a separate, additive signal for AI consumers. Your sitemap is what helps SEO. The two files complement each other — sitemap is exhaustive, `llms.txt` is curated.
### How often should I update it?
Whenever the cornerstone list changes. The file is small enough that automation isn't necessary, but if your content layer can emit it (as ours does from Velite), it stays in sync for free. We rebuild ours on every deploy.
### Should I include marketing pages?
No. Pricing, about, and announcement pages belong in a sitemap, not an `llms.txt`. The spec is for an LLM-grounding index — the pages an LLM should ground on when answering a question about your site. Marketing pages tell humans why to buy; they don't help an LLM answer "how does X work."
### What's the difference between `llms.txt` and `llms-full.txt`?
`llms.txt` is the index — short, curated, mostly links. `llms-full.txt` is the index plus the full Markdown body of each cornerstone, concatenated. A model with a generous context window can ingest the entire `llms-full.txt` in a single request and ground on it. We recommend shipping both; `llms.txt` is the spec primitive, `llms-full.txt` is the optimisation.
### Does the file need to be exactly at the root?
The spec says `/llms.txt` at the apex. A minority of implementations put it under `/.well-known/llms.txt` instead. We treat the apex location as canonical; `.well-known` is a deviation that may be tolerated by tooling but isn't documented in the spec.
### Can I block specific LLMs with this file?
No. `llms.txt` is descriptive, not restrictive. For blocking crawlers, use `robots.txt` with the LLM crawler's user-agent (e.g., `GPTBot`, `ClaudeBot`, `Google-Extended`, `PerplexityBot`). The two files have different jobs.
### Is the format going to change?
Possibly. The proposal is on version 0.1 as of 2025. The community has discussed adding YAML frontmatter, a JSON-LD variant, and a stricter schema. For now, the Markdown form we describe above is what every existing implementation uses. We'll publish a follow-up if the spec changes materially.
### Should the file be cached?
Yes. We serve ours with `Cache-Control: public, max-age=3600, s-maxage=86400` — fresh for browsers for an hour, fresh on the CDN for a day. A daily cache invalidation is plenty unless you ship cornerstones more often than that.
## Dataset
Full 100-domain audit: CSV download.
The CSV contains one row per audited domain with the following columns:
`rank`, `category`, `brand_name`, `domain`, `probe_timestamp_utc`,
`homepage_status`, `homepage_response_ms`, `homepage_title`,
`llms_txt_found`, `llms_txt_status`, `llms_txt_size_bytes`,
`llms_txt_content_type`, `llms_txt_compliance_score`,
`llms_txt_compliance_notes`, `llms_full_txt_found`,
`llms_full_txt_status`, `llms_full_txt_size_bytes`,
`well_known_llms_txt_found`, `robots_txt_found`, `robots_txt_size_bytes`,
`error`.
`llms_txt_compliance_score` is a 0-1 decimal scored across seven
dimensions: has H1, has summary blockquote, has at least one structured
section, has typed links in Docs or Optional, valid UTF-8, body size
between 200 bytes and 100 KB, no HTML leakage in the body. The
`compliance_notes` column lists which dimensions failed for each
non-compliant file.
The crawler that produced this dataset is open source:
[`scripts/llms-txt-audit.ts`](https://github.com/geosalience/geosalience).
You can re-run it against your own list of domains, or against ours to
check our work.
## How we wrote this
This article combines spec reading, a 100-domain automated crawl, and
about ten hours of hand-inspecting the files we found.
- **Spec source:** [llmstxt.org](https://llmstxt.org/), accessed
2026-05-19. The proposal was published by Jeremy Howard / Answer.AI
in 2024.
- **Audit methodology:** automated crawl of 100 domains on 2026-05-19,
scored against the seven compliance dimensions described in the
Dataset section. The crawler used `undici.fetch` with a
`GeosalienceBot/1.0` user-agent, followed up to five redirects, and
capped each response body at 64 KB. Concurrency was 10; total wall
time was 31 seconds. Full crawler source and per-domain raw output
published.
- **Limitations:** A single point-in-time crawl. Sites that ship
`llms.txt` behind authentication, on a CDN that 403s our user agent,
or under a non-apex domain are recorded as "not found" even if they
exist. We rerun the audit quarterly; year-over-year adoption is the
more interesting metric.
- **Conflicts of interest:** None. GeoSalience is independent. We have no
commercial relationship with Answer.AI, llmstxt.org, or any of the
audited brands.
For our general methodology, see [Methodology](/methodology). For how
we handle right-of-reply and corrections, see
[Editorial Policy](/editorial-policy).
## See also
- [Technical pillar](/technical) — every article we publish on the technical side of GEO, from `llms.txt` to schema markup.
- [JSON-LD recipes for Articles, Datasets, and FAQs](/technical/json-ld-recipes) — the structured-data layer that complements `llms.txt`, with copy-paste blocks.
- [What is GEO](/foundations/what-is-geo) — the broader concept this audit sits inside.
- [GEO vs AEO vs LLMO vs SGE](/foundations/geo-vs-aeo-vs-llmo-vs-sge) — the terminology, if the acronyms in this piece are unfamiliar.
- [Knowledge cutoff and web access](/foundations/knowledge-cutoff-and-web-access) — why an `llms.txt` file helps the retrieval layer reach you even when a model's training knowledge is stale.
- [This site is our GEO lab](/case-studies/geosalience-as-its-own-case-study) — we turn the same audit lens on ourselves, measuring our own AI-crawler traffic and citation rate in public.
- [Citation rate](/glossary/citation-rate) — the metric we use to measure how often an LLM links back to a source.
- [Share of voice](/glossary/share-of-voice) — how often a brand appears across a set of LLM answers.
---
# How to Get Cited by LLMs: The Complete Taxonomy of GEO Methods
> Every method GEO practitioners use to surface in ChatGPT, Claude, Perplexity, and AI Overviews — grouped into five families and rated by evidence quality. A synthesis of the published literature, vendor docs, and our own audits. The map of the discipline as of June 2026.
Published: 2026-05-19T00:00:00.000Z
Updated: 2026-06-08T00:00:00.000Z
Canonical: https://geosalience.com/foundations/geo-methods-taxonomy
- GEO methods cluster into **five families**: technical foundations, content design, structural metadata, brand and authority signals, and measurement.
- The strongest published evidence (Aggarwal et al., 2024) shows that **citing primary sources, adding statistics, and inline quotations** improve a source's visibility in generated answers by **up to ~40%** on the paper's metrics, across ChatGPT, Perplexity, and BingChat — a larger gain than any technical change they tested.
- The single most underweighted method on practitioner blogs is **answer-first chunking** — writing the first 1–2 sentences of every H2 so that they can be extracted and quoted on their own, without context from the rest of the page.
- Tactics that *sound* high-leverage but show weak public evidence as of May 2026: blanket `llms.txt` adoption with no link list, FAQ schema on every page, and "AI-friendly" copywriting tricks that do not change information density.
- Treat this taxonomy as a working map, not a ranking. Run the **30-day implementation roadmap** at the end of this guide and measure on your own data — every citation engine is a moving target.
The Generative Engine Optimization (GEO) literature is young. The first paper to use the term was published in late 2023; the first round of trade-press how-to articles arrived in 2024; the first batch of measurement tools (Profound, Peec, Otterly, Athena HQ) shipped between 2024 and 2026. In a discipline this new, the inventory of *methods practitioners actually use* is more useful than another ranked listicle of "top tactics".
This guide is that inventory, and the map for the [Foundations](/foundations) pillar of this site. We have grouped every method into five families, attached the public evidence to each, flagged the antipatterns we see most often, and closed with an implementation order you can run in a month. We will revisit the entries as we publish our own primary research over the next two quarters — first published **2026-05-19**, last reviewed **2026-06-08**.
If you only have five minutes, read the callout above and the **Evidence Matrix** near the bottom. The map of the territory matters more than any single method.
## The map: five families of GEO methods
Every method that earns a place in this guide is a deliberate change to one of five layers of a website. We group them this way because the layers correspond to different stages of how an LLM ingests, retrieves, and renders your content.
| Family | What it changes | Who reads the signal | Time to effect |
|---|---|---|---|
| Technical foundations | The bytes your server returns | Crawlers, parsers, RAG indexers | Hours–days |
| Content design | The information density on each page | The generator at answer time | Days–weeks |
| Structural metadata | The relationships between pages | Crawlers, knowledge-graph builders | Days–weeks |
| Brand and authority signals | The web's opinion of your site | Pre-training filters, ranking models | Weeks–months |
| Measurement and iteration | Your ability to know what worked | You (and the next decision you make) | Continuous |
The families are not a ladder you climb in order. They are five surfaces you should be working on in parallel, because they each feed a different decision the LLM is making about whether to cite you. The taxonomy that follows expands every family into the methods inside it — twenty-three in total at this revision.
## Mental model: how LLMs choose which sources to cite
Before the methods, the mechanics. Different methods target different stages of the pipeline, and the only way to reason about leverage is to know which stage you are influencing. As of May 2026, a citation-capable LLM passes a query through three stages.
### Stage 1 — Pre-training corpus inclusion
The base model was trained on a snapshot of the web. If your domain was in that snapshot, the model has *some* representation of it — terminology, common claims, brand entities. This stage is mostly out of your direct control on a short horizon; you cannot retroactively change what Common Crawl picked up. What you *can* do is increase the chance you will be in the next snapshot: clean HTML, indexable URLs, sufficient text-to-chrome ratio, and content the open web is willing to link to.
For most practitioners reading this guide, pre-training inclusion is a long-horizon investment. The next stage is where short-term wins live.
### Stage 2 — Retrieval (RAG)
When a user asks a citation-capable LLM a question, the system runs a retrieval step — usually a hybrid of keyword and vector search against a live index — and selects a small number of source documents to ground the answer. This is the same pattern described by Lewis et al. (2020) in the original RAG paper, now productised across ChatGPT Browse, Perplexity, Claude with web search, Copilot, and AI Overviews.
Retrieval is the stage where most GEO leverage lives. The system needs to be able to find your page, parse it, decide it is relevant, and extract a useful chunk. Every method in the *technical foundations*, *content design*, and *structural metadata* families exists to make one of those four sub-steps work better for you.
### Stage 3 — Grounded generation and citation
Once the model has retrieved a set of candidate sources, it generates an answer and decides which sources to attribute. The attribution decision is influenced by the model's training (what it was rewarded for in RLHF), the system prompt of the product surface (ChatGPT vs Perplexity vs AI Overviews behave differently), and the *quality* of the extractable chunks. A page that is high-relevance but unextractable — for example, key information locked inside an image — frequently fails to be cited even when it would have helped.
Brand and authority signals operate quietly across all three stages: they raise the prior probability the page is selected during retrieval, raise the model's confidence at generation time, and raise the chance the page ends up in the next training snapshot.
With the map and the mechanics in place, we can walk the methods.
## Family 1 — Technical foundations
Technical methods change what your server returns when a crawler, parser, or RAG indexer fetches a URL. They are the cheapest family to implement (most are a few hours of work) and they are the easiest to verify (you can `curl` the result). They will not, on their own, make a weak page citable — but a strong page without them is leaving low-cost wins on the table.
### Method 1 — Schema.org JSON-LD
Schema.org is a vocabulary maintained by Google, Microsoft, Yahoo, and Yandex for marking up the entities on a page. It is delivered as a JSON-LD block in the document `` and consumed by every major crawler.
For an editorial site, the high-value types are `Article` (or `NewsArticle`, `BlogPosting`), `Person` for author bylines, `Organization` for the publisher, `BreadcrumbList` for site hierarchy, `FAQPage` where the article carries an actual FAQ block, and `DefinedTerm` for glossary entries. Aggarwal et al. (2024) did not isolate JSON-LD as a single test variable, so the direct citation-lift evidence is still mixed — but the indirect evidence is strong: pages with valid `Article` markup are more reliably parsed by RAG systems and are more frequently surfaced in AI Overviews, which inherits Google's existing dependence on structured data.
A minimal `Article` block, on a real published page:
```json
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "How to Get Cited by LLMs",
"datePublished": "2026-05-19",
"dateModified": "2026-05-19",
"author": {
"@type": "Person",
"name": "Jane Doe",
"url": "https://example.com/authors/jane-doe"
},
"publisher": {
"@type": "Organization",
"name": "GeoSalience",
"url": "https://example.com"
},
"mainEntityOfPage": "https://example.com/foundations/geo-methods-taxonomy"
}
```
Two practical rules: validate every page in `validator.schema.org` and Google's Rich Results Test before you call the work done, and never declare attributes that are not visibly true on the page — a `FAQPage` schema without a real FAQ block is treated as a quality signal *against* you. We walk the exact JSON-LD blocks we ship — `Article`, `BreadcrumbList`, `Dataset`, and `FAQPage` — in [JSON-LD recipes for GEO](/technical/json-ld-recipes).
### Method 2 — `llms.txt` and `llms-full.txt`
`llms.txt` is a proposal by Jeremy Howard ([llmstxt.org](https://llmstxt.org), 2024) for a Markdown-formatted index file at the root of your domain. Its job is to hand-curate which URLs on your site matter for LLM consumers, with one-line descriptions, organised by section. The companion `llms-full.txt` concatenates the full text of the most important pages.
The honest evidence picture, as of May 2026: no major LLM vendor has confirmed that `llms.txt` is in their crawler pipeline. The file is read by an unknown number of agents, including the long tail of RAG-as-a-service products. The cost of shipping one is roughly ten minutes of work for a small site, and the downside risk is zero. The cost-benefit is positive even at low adoption.
The failure mode we see most often is shipping an `llms.txt` with an H1 and a few sections but no actual link list under `## Docs` or `## Optional`. That defeats the file's purpose. A useful `llms.txt` *is* the link list. We covered the spec, the adoption audit, and a working setup in [llms.txt: spec, adoption, setup](/technical/llms-txt-spec-adoption-setup).
### Method 3 — Semantic HTML and document structure
LLM retrievers work on the parsed document, not the rendered DOM. Heading elements (`
` through `
`), proper `` and `` boundaries, real `
` markup, and accessible ``/`` pairs are not cosmetic — they are the structural anchors a retriever uses to slice your page into chunks.
Three rules that matter:
The first heading on every article page is `
`, used once, with the headline. The model uses it as the canonical answer to "what is this page about". A page where the visible title is rendered as a `
` with a CSS class is invisible to many parsers.
Section headings should be answer-bearing, not navigational. `
How citation rate is calculated
` is more retrievable than `
The math
`, because the H2 itself is a quotable answer to a query.
Tables and lists should be real semantic markup, not divs styled to look like tables. The number of times we have seen "tabular" data shipped as nested divs is alarming — and every one of those tables is invisible to a chunker.
### Method 4 — Markdown-as-source and `.md` aliases
A growing pattern, popularised by [Anthropic's documentation](https://docs.anthropic.com), [Vercel](https://vercel.com), and the Stripe API reference, is to serve a raw-Markdown version of every important page at a `.md` alias — for example, `/foundations/what-is-geo` and `/foundations/what-is-geo.md`.
The mechanism is simple: when an LLM or developer wants the canonical, parse-friendly version of the page, they fetch the `.md` URL and get the source text without HTML chrome, JavaScript, or rendering quirks. The page is identical in content, but cheaper to ingest. Several RAG products (notably those serving developer documentation) preferentially fetch the `.md` alternate when one is advertised in ``.
This is a cheap method. If you author in Markdown or MDX, you already have the source — exposing it as an alternate is a route handler. If you do not, the conversion cost is the main bottleneck.
### Method 5 — Canonicalization, robots, sitemap, and Open Graph
The least exciting and most underweighted block of technical work. Each of the four pieces is doing a different job:
`rel="canonical"` collapses near-duplicate URLs into a single citation target. Without it, retrieval can split your authority across query-parameter variants, mobile subdomains, or AMP versions, and the LLM cites the version with the worst extractable content.
`robots.txt` is your contract with crawlers. The contemporary debate is whether to allow `GPTBot`, `ClaudeBot`, `PerplexityBot`, and `Google-Extended` — the line where you trade discoverability for control of your training-corpus inclusion. The honest answer in May 2026 is that *blocking* these crawlers measurably reduces citation rate, and the evidence for that is well-documented in publisher case studies after the 2024–2025 wave of New York Times-style opt-outs. If your strategy is to be cited, you want them allowed.
`sitemap.xml` is how you tell the crawler what you have. Generate it programmatically, include every public article, set `lastmod` accurately, and submit it via Google Search Console and Bing Webmaster. AI-native crawlers increasingly use it as a discovery primitive.
Open Graph (`og:title`, `og:description`, `og:image`) and Twitter Cards (`twitter:card`, `twitter:image`) drive social distribution, which drives backlinks, which feeds Family 4. Treat them as part of the technical baseline, even if their direct GEO effect is downstream.
## Family 2 — Content design
Content methods change the information density of the page itself — the part the generator reads at answer time. They are the highest-leverage family for an established editorial site because they target Stage 3 directly: the model has already retrieved you and is now deciding whether your chunk is good enough to quote.
### Method 6 — Answer-first H1 and dek
The H1 and the dek (subtitle) together are the unit the retriever uses to decide whether to surface your page at all. The rule, distilled from a year of looking at what gets cited and what does not:
The H1 should be the question or claim the page answers. Not a clever angle, not a hook, not a brand reference. If a reader asked "how does ChatGPT decide which source to cite?", the H1 should be a recognisable rewrite of that question — *"How ChatGPT Decides Which Source to Cite"* — not *"The Citation Lottery"*.
The dek (subtitle, 1–2 sentences, ≤280 characters) compresses the article's specific finding into a sentence the model can quote verbatim. A template like *"We ran [N] ChatGPT sessions across [M] niches and recorded every citation"* is a citable sentence; *"In this article, we'll explore how citation works"* is filler.
We are testing this method on our own pages — the live protocol and dates are in our [experiment log](/lab/experiments) — and will publish the before/after result once the measurement window closes. Until then, treat the answer-first rule as well-supported by the published literature (Aggarwal et al., 2024) rather than by a result of ours.
### Method 7 — Chunkability
Retrievers do not read your page; they read a *chunk* of it. The unit of citation is the chunk, and chunkability is a property of how easily a retriever can carve your page into self-contained, quotable pieces.
Concretely, chunkability improves when:
The first one or two sentences after each H2 stand on their own as an answer to the H2's implicit question. A reader (or model) who lands on just that section understands the point without scrolling up.
Paragraphs are short. Two to four sentences is the editorial sweet spot for screen reading, and it happens to coincide with the chunk-size sweet spot for most RAG systems (which target windows of roughly 150–500 tokens).
Lists, tables, and code blocks are real structured elements, not ASCII art. The retriever knows what a `
` is; it does not know what to do with em-dashes in a wall of prose.
Pull quotes and standalone-sentence claims are *intentional*. A sentence that you want to be quoted should be its own paragraph, free of subordinate clauses, and built around a concrete number or named entity. The Aggarwal et al. (2024) testing of "quotation lift" was operating on exactly this property.
### Method 8 — Primary research and methodology disclosure
This is the single highest-evidence method on the public record. Aggarwal et al. (2024) tested nine optimisation strategies on a held-out query set against ChatGPT, BingChat, and Perplexity. The strategies that moved the paper's visibility metric the most — by up to ~40% on some metrics — were *citing sources*, *adding statistics*, and *adding quotations*. Each of these is, fundamentally, a marker of primary or near-primary research.
The reason the effect is large is that primary research is *information you cannot get elsewhere*. When a model has to synthesise an answer and your page is the only source carrying a specific number, methodology, or quotation, you are not competing for citation — you are the only candidate.
The implementation cost is high. Primary research means running a test, collecting a dataset, writing a methodology section, and publishing the data alongside the article. The payoff is that one well-executed piece outperforms ten generic explainers on the same topic, and you have a defensible position when the next wave of generic explainers floods the topic.
If full primary research is out of reach for a given article, the next-best step is to write a transparent methodology paragraph anyway — *"We compared X, Y, and Z by reading their public documentation and replicating the setup described in each"* — because methodology disclosure itself is a citability signal.
### Method 9 — Entity coverage and inline definitions
LLMs are entity-aware. The first time you mention "Schema.org", "RAG", "Perplexity", or any other proper noun on the page, you have an opportunity to define it inline in a sentence the model can quote on its own. Schema.org's `DefinedTerm` type makes the definition explicitly machine-readable.
Two practical heuristics:
A cornerstone article should mention at least 8–12 named entities relevant to its topic — people, papers, tools, concepts — each with a concrete reference or link. Sparse entity coverage correlates with weaker pre-training representation; dense entity coverage signals the page is *about* the topic in a way that matters for retrieval.
The first inline definition should be a complete sentence: *"Retrieval-Augmented Generation (RAG) is the technique of fetching documents at query time and passing them to a language model as context, introduced by Lewis et al. (2020)."* Not *"RAG (defined below)"*, not *"RAG, which we'll cover later"*.
### Method 10 — Citation density to primary sources
Every important claim on the page should resolve to a primary source — a paper, an official document, a dataset, an announcement from the organisation responsible for the thing being described. Linking to a parahprasing blog post is a signal that you did not check the original, and increasingly, retrievers can detect the citation distance from your page to the closest primary source.
The practical version: for a 3,000-word article, expect 10–25 outbound primary-source citations. Each one should be a hyperlink in the body of the text (not a footnote, not a "Sources" appendix), and the anchor text should describe the cited work specifically, not "click here".
The anti-pattern we see most often: a "Sources" or "References" section at the bottom of the article with ten URLs and no inline anchoring. This is the artefact of articles written without checking the sources during writing, and retrievers treat it as such.
## Family 3 — Structural metadata and discovery
Structural methods change the *relationships* between pages — how your site is wired together internally, and how the rest of the web points to it. These methods feed both the retriever (which uses link topology as a relevance signal) and the pre-training corpus (where co-citation patterns are a strong authority cue).
### Method 11 — Internal linking topology and pillar architecture
A pillar-and-spoke architecture is the editorial form most retrievers reward. The pillar is a long, definitional article on a broad topic (this one). The spokes are narrower articles that each cover a sub-topic, link back to the pillar, and are linked *from* the pillar. The result is a small graph where every node is reachable in two or three clicks and every node has a clear topical neighbourhood.
Three concrete rules:
Every spoke article links to its pillar at least once in the body. Anchor text varies — sometimes the pillar's title, sometimes a sub-claim from it — to avoid the appearance of templated linking.
Every pillar article links to at least three of its strongest spokes from the body of the text, not just from a "See also" appendix. The links should be where a reader would naturally pivot to the deeper topic.
Orphan pages — articles with no inbound internal links — are an outright failure mode. They are functionally invisible to most retrievers. Run a periodic audit (we have a `pnpm audit:links` script in this codebase) to surface orphans and either link them or unpublish them.
### Method 12 — External backlinks, co-citation, and the open-web graph
The pre-LLM SEO regime cared about backlinks. The LLM era still does, but the *type* of backlink that matters has shifted. Co-citation in Wikipedia, news outlets, academic papers, and `.edu` domains is more valuable than ever, because these are exactly the sources that LLMs are over-represented in during retrieval and pre-training.
The corollary: a single citation from a Wikipedia article is worth more than a hundred backlinks from low-authority directories. The methods that move this needle are slower than technical fixes — they look like contributing primary research that gets cited by others, publishing data that journalists use, and being the cleanest available source on a specific narrow topic.
The fastest-acting subspecies of this method is *unlinked mentions*. LLMs frequently associate brands with topics even when there is no hyperlink between the source mentioning the brand and the brand's own domain — and unlinked mentions in trade press, podcast transcripts, and conference talks contribute meaningfully to that association. Tracking mentions, not just backlinks, is an emerging measurement practice.
### Method 13 — Cover images, OG cards, and the social distribution loop
A separate but related method: the cover image and Open Graph card determine whether your article spreads on social — and social spread is the most reliable producer of the kind of backlinks and unlinked mentions that move Family 4. Treat the OG card as a piece of the article, not an afterthought.
Practical specs: 1200×630 pixels, under 200KB, type and image legible at thumbnail size, text in the image readable in both light and dark mode previews. Test the rendered card in `opengraph.xyz` or `cards-dev.twitter.com` before publishing.
## Family 4 — Brand and authority signals
Brand methods change the web's opinion of your domain. They are the slowest-acting family and the hardest to fake. The methods here will not pay off in week one — but they are the difference between being cited occasionally for narrow technical queries and being cited routinely as a canonical source on a topic.
### Method 14 — Verified author identity and Schema.org `Person`
Every article on an editorial site should have a named author with a real biography, a real public identity, and Schema.org `Person` markup that links the byline to the same identity across the web — LinkedIn, ORCID for academic authors, GitHub for technical authors, a professional website. The marker is `sameAs`, a list of URLs the author is known by elsewhere.
LLMs use authorship as an authority signal at multiple stages. At retrieval time, named-author articles outperform anonymous ones in domains where expertise matters (medical, legal, financial, technical). At generation time, the model is more willing to quote a source it can attribute to a specific named expert. At pre-training time, the entity-resolution pipeline links the author's articles across the web, building a stronger signal than any single piece of content could.
The anti-pattern: a generic "Staff" or "Editorial Team" byline. We have not yet seen a published study quantifying the penalty, but the directional evidence from search-era E-E-A-T research carries forward.
### Method 15 — E-E-A-T in the LLM era
Google's E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) framework was a search-era construct, but the underlying signals — that the author has lived experience, demonstrated expertise, recognition by peers, and a trustworthy publishing history — are exactly the ones an LLM picks up on as well. The mechanism is different (entity resolution and corpus statistics, rather than a ranking model), but the signals are the same.
In practice, E-E-A-T translates to four things on an editorial site:
A "How we tested" or "Methodology" block on every research article, disclosing how the data was collected and what was *not* tested. This is exactly the editorial muscle that distinguishes a citable source from an opinion piece.
Visible author credentials — title, organisation, prior work — on every byline. The author's name, alone, is weaker than the name plus a recognisable affiliation.
A publication date and a *last-reviewed* date on every article. Articles without dates are aggressively down-weighted by retrievers, because the model cannot tell whether the information is current.
A correction and update policy that is visible and used. Errata are a strength signal, not a weakness signal — they tell the model that the source is being maintained.
### Method 16 — Mention-building as a discipline
Mention-building is the active practice of getting your name, your work, and your data into the documents that LLMs over-represent: Wikipedia (carefully, following the project's conflict-of-interest rules), trade press, podcast transcripts, conference talks, GitHub READMEs, and the long tail of niche-authoritative blogs.
The work is unglamorous and slow. It looks like contributing data to a public dataset that journalists then cite, writing guest essays for trade publications, being a useful source for a journalist on a deadline, sponsoring a niche newsletter not for the link but for the editorial nod, and consistently showing up in the comment sections and forums where the topic is discussed. The compounding effect over twelve to twenty-four months is significant — and is the single most-cited factor in the case studies of brands that "appear out of nowhere" in LLM citations.
## Family 5 — Measurement and iteration
Measurement is the family most practitioners skip — and the one without which every other family is faith-based. You cannot tell whether a method is working without a measurement baseline, and you cannot pick the next method to invest in without a sense of where the current gaps are.
### Method 17 — Citation rate as the primary KPI
Citation rate is the share of relevant prompts in a defined battery where a given LLM cites your domain at least once. It is the single metric closest to the outcome you actually care about — being part of the answer to a query in your topical neighbourhood.
The formula:
> **Citation rate** = (number of prompts in the battery where the engine cited your domain in the answer) ÷ (total number of prompts in the battery)
A *battery* is a set of prompts you have committed to monitoring on a recurring basis. For a typical editorial site, the right size is 50–200 prompts, covering the queries you would want to be cited for. Run the battery on a schedule (weekly is sufficient for most domains) and log the result per engine.
The trap to avoid: vanity prompts. If your battery consists only of prompts you already win, the metric is meaningless. Include 30–40% prompts where you currently lose, so the metric has room to move.
### Method 18 — Share-of-Voice across engines
Share-of-Voice (SoV) is the share of cited sources, across a battery, attributable to your domain. Where citation rate asks "do they cite us at all?", SoV asks "of all the sources cited in this neighbourhood, what fraction is us?".
The formula:
> **Share of Voice** = (number of citations to your domain across the battery) ÷ (total number of citations across the battery)
SoV is harder to move than citation rate, and the headline number is usually small (1–8% even for category-leading editorial sites). What matters is the trend, not the absolute value. We covered the measurement playbook in detail in [Share of Voice — definition and method](/glossary/share-of-voice).
### Method 19 — Cross-engine prompt batteries
A method that is doing well in ChatGPT may be doing poorly in Perplexity, AI Overviews, Claude, or Copilot, because each engine weights signals differently. The discipline is to run the same battery on every engine you care about and to track the divergence, not just the average.
Practically, this looks like:
A single CSV of prompts, versioned in the repo. Every entry has a prompt text, an intent label, and an expected-domain shortlist (the small set of domains it would be reasonable to cite for that prompt).
A weekly run that, for each prompt, queries every engine, parses the citation set, and writes a row per (prompt, engine, timestamp, cited_url, cited_position). Today this is most efficiently done with a combination of headless browser automation for engines without an API and the official APIs where they exist.
A small dashboard or notebook that, for each engine, shows citation rate and SoV over time, and lets you slice by intent and by content cluster.
### Method 20 — Tools landscape (May 2026 snapshot)
The market for citation-tracking tools has matured. As of May 2026, the major options are:
[Profound](https://www.tryprofound.com/) — enterprise, large prompt batteries, brand-tracking focus.
[Peec AI](https://peec.ai/) — share-of-voice oriented, European, fast iteration.
[Otterly.ai](https://otterly.ai/) — affordable, prompt-monitoring focused, good for a single-brand setup.
[Athena HQ](https://www.athenahq.ai/) — enterprise, AI-search-monitoring with workflow integrations.
[Goodie](https://goodie.ai/) — emerging, dashboards over agentic-search results.
A hands-on benchmark is in progress; until it ships, treat this list as a market map, not a ranking — it is alphabetical within category and reflects market presence as of May 2026, not endorsement. Either way, every tool's signal-to-noise depends on how well-defined your prompt battery is. Buying a tool without a battery is a recipe for dashboards no one trusts.
### Method 21 — Manual prompt-testing protocol
You do not need a paid tool to start measuring. A reproducible manual protocol, run once a week for an hour, is enough to establish a baseline.
The minimum-viable protocol:
A fixed list of 50 prompts in a spreadsheet, versioned in your repo.
A fresh, signed-out browser session for each engine (ChatGPT, Claude, Perplexity, AI Overviews via Google Search).
For each prompt-engine pair: paste the prompt, wait for the answer, copy the cited URLs, paste them back into the spreadsheet under a date-stamped column.
A simple pivot at the end: citation rate per engine, citation rate per intent cluster, SoV.
The manual protocol is fragile (engines change UI, sessions personalise) but it is honest. It is the right starting point for the first 30–60 days while you decide whether to invest in tooling.
### Method 22 — Pre/post measurement around deploys
The most underused application of citation tracking is *attribution*. When you ship a content change — a rewrite, a new schema block, a new internal-link cluster — measure the change on the same battery in a two-week window before and after the deploy. Most methods either move the needle or do not, and you find out which is which by running the battery, not by reading a blog post.
The instrumentation has to be ready *before* you ship. Pre-deploy battery, deploy, post-deploy battery, and a written note in the article's frontmatter (`lastReviewedAt`, `updatedAt`) that anchors the timeline.
## The evidence matrix
Not every method has equal evidence. Here is our current read, as of 2026-06-08, drawn from the published literature, our own already-published audits, and the documented case studies from other practitioners. We will update the matrix as we publish more primary research.
| Method | Family | Evidence quality | Implementation cost | Time to effect |
|---|---|---|---|---|
| 1. Schema.org JSON-LD | Technical | Medium (indirect) | Low | Days |
| 2. `llms.txt` / `llms-full.txt` | Technical | Low (emerging) | Low | Unknown |
| 3. Semantic HTML | Technical | High | Low–medium | Days |
| 4. Markdown alternates | Technical | Medium | Low | Days |
| 5. Canonical / robots / sitemap / OG | Technical | High | Low | Days |
| 6. Answer-first H1 + dek | Content | High (Aggarwal et al. 2024) | Low | Weeks |
| 7. Chunkability | Content | High | Medium | Weeks |
| 8. Primary research + methodology | Content | **Very high** (Aggarwal et al. 2024) | **High** | Weeks–months |
| 9. Entity coverage + definitions | Content | Medium | Low | Weeks |
| 10. Citation density to primary sources | Content | High (Aggarwal et al. 2024) | Medium | Weeks |
| 11. Internal linking topology | Structural | Medium | Medium | Weeks |
| 12. External backlinks + co-citation | Structural | High | Very high | Months |
| 13. Cover images / OG cards | Structural | Indirect (via social) | Low | Days–weeks |
| 14. Verified author identity | Authority | Medium (carry-over from SEO) | Medium | Weeks–months |
| 15. E-E-A-T in the LLM era | Authority | Medium | High | Months |
| 16. Mention-building | Authority | High (case studies) | Very high | Months |
| 17. Citation rate KPI | Measurement | n/a (process) | Low | n/a |
| 18. Share-of-Voice KPI | Measurement | n/a (process) | Low | n/a |
| 19. Cross-engine batteries | Measurement | n/a (process) | Medium | n/a |
| 20. Citation-tracking tools | Measurement | n/a (process) | Medium ($) | n/a |
| 21. Manual protocol | Measurement | n/a (process) | Low | n/a |
| 22. Pre/post deploy measurement | Measurement | n/a (process) | Medium | n/a |
| 23. Methodology disclosure (standalone) | Content | Medium | Low | Weeks |
Evidence quality is graded *very high* when there is at least one peer-reviewed paper or replicated public test showing the effect; *high* when there are multiple practitioner case studies pointing the same direction; *medium* when the direction is plausible from related disciplines but not directly tested; *low* when adoption is recent and outcomes are unknown.
## Common antipatterns
Methods that *sound* like they should work, but do not — or that fire backward.
The first antipattern is **schema-stuffing**: declaring `FAQPage`, `HowTo`, or `Review` schema on pages that do not contain the corresponding visible content. Google has been explicit since 2023 that this is a manual-action risk for traditional search; the LLM era has not changed the calculus. Treat schema as documentation of what is on the page, not as a wishlist.
The second is **AI-generated content without disclosure or editing pass**. The 2024–2025 wave of mass-generated content has trained retrievers to recognise the surface patterns — uniform paragraph length, predictable phrase structure, low entity density, no methodology section. The penalty is not the AI involvement itself; it is the absence of the editorial pass that would make the piece worth citing.
The third is **keyword stuffing rewritten as "LLM optimisation"**. Repeating "generative engine optimization" twenty times in an article does not make it more citable. What makes it citable is information density — a measurable claim, a number, a methodology, a quote — that LLMs cannot get from the other twenty pages that repeated the same phrase.
The fourth is **cloaking** — serving different content to crawlers than to humans. This was always against the rules in traditional SEO; in the LLM era the consequences are sharper, because the content that gets ingested into the training corpus is the content the crawler saw, which is then visible at generation time. A user catching the discrepancy is a brand-trust event you cannot recover from quickly.
The fifth is **over-optimisation that collapses readability**. Paragraphs broken into single-sentence units, headings every fifty words, lists where prose would be clearer — these are all "chunkability" rules taken past the point of usefulness. The model rewards extractable chunks; it does not reward shredded text.
The sixth, and the one we see most often on otherwise-strong sites, is **shipping technical changes without measuring**. Schema added but not validated. `llms.txt` shipped but no link list. Canonical tags added but not audited. Without measurement, the technical work is faith-based — and faith-based work tends to accumulate as cruft rather than as compounding wins.
## A 30-day implementation roadmap
If you start tomorrow, the order that produces the most movement per hour of effort is technical first, content second, measurement in parallel from day one, authority over the long haul.
**Week 1 — Technical baseline.** Audit and fix Schema.org markup on the top 20 pages by traffic (Methods 1, 3, 5). Ship `llms.txt` if you do not already have one (Method 2). Validate everything in `validator.schema.org` and Google's Rich Results Test. Confirm canonical tags, `robots.txt`, and `sitemap.xml` are clean. Add `.md` alternates if you author in Markdown (Method 4). Add or update Open Graph cards on the top 20 pages (Method 13). Expect this to take three to five days of focused work.
**Week 2 — Content rewrites.** Pick the top five pages by traffic or strategic importance. Rewrite each one for answer-first H1 and dek (Method 6), chunkability (Method 7), entity coverage (Method 9), and inline citation density (Method 10). If the page is a research piece, add a "How we tested" or methodology block (Methods 8 and 23). The mechanical work is fast; the editorial discipline is the bottleneck.
**Week 3 — Authority and topology.** Audit internal linking topology and resolve every orphan article (Method 11). Add or improve author pages with verified identity and `sameAs` links (Method 14). Identify five primary-source citations the top pages currently route through paraphrases, and replace them with the originals. Spend the rest of the week on Method 16 — pick two outlets in your topical neighbourhood and figure out what useful primary research, data, or commentary you could send them in the next quarter.
**Week 4 — Measurement baseline.** Define your prompt battery (50 prompts to start — Method 19). Run the manual protocol against ChatGPT, Claude, Perplexity, and AI Overviews (Method 21). Compute citation rate and Share of Voice per engine (Methods 17–18). Pick one paid tool to evaluate in week 5 (Method 20). Write down where you are; everything you do from week 5 on should be measurable against this baseline.
Beyond day 30, the rhythm is: every shipped change goes through pre/post measurement on the same battery (Method 22). Every quarter, the battery itself gets reviewed and expanded. Every year, the technical baseline gets re-audited, because the schema vocabulary, the crawler list, and the engines themselves will all have moved.
## FAQ
**What is the difference between GEO, AEO, LLMO, and SGE?**
GEO (Generative Engine Optimization) is the term most commonly used in the academic literature and was introduced by Aggarwal et al. (2024). AEO (Answer Engine Optimization) is the older, broader term, dating to roughly 2014–2015 and originally focused on featured snippets and voice assistants. LLMO (Large Language Model Optimization) is a near-synonym for GEO with slightly more emphasis on no-browsing scenarios. SGE (Search Generative Experience) was Google's product name for AI Overviews before the May 2024 rebrand — it is a product, not a discipline. We use **GEO** on this site for the reasons discussed in [GEO vs AEO vs LLMO vs SGE: an honest taxonomy](/foundations/geo-vs-aeo-vs-llmo-vs-sge).
**Do LLMs actually read my Schema.org markup?**
Indirectly, yes. The major LLM-powered surfaces (ChatGPT Browse, Perplexity, Claude with web search, AI Overviews, Copilot) all rely on crawlers and parsers that consume Schema.org. The direct effect on citation rate has not been isolated in published research, but valid markup raises the probability your page is correctly parsed, classified, and surfaced.
**Does `llms.txt` actually do anything if no major vendor has confirmed they read it?**
The cost of shipping a well-formed `llms.txt` is roughly ten minutes; the worst case is that no agent reads it. The best case is that you are correctly indexed by the growing long tail of RAG-as-a-service products that have adopted the spec. It is a positive expected-value bet at low cost. The actual evidence on adoption is still emerging; see our [llms.txt audit](/technical/llms-txt-spec-adoption-setup) for the 100-domain snapshot.
**How long does it take to be cited by ChatGPT after publishing a new article?**
For ChatGPT with browsing enabled, the citation can be picked up within hours once the page is indexed. For the no-browsing case (the model citing from its training corpus), the lag is the gap between publishing and the next pre-training snapshot — generally months. Most of the citation lift practitioners can drive in a short horizon comes through the browsing pathway.
**Can I pay to be cited by Perplexity, ChatGPT, or AI Overviews?**
As of May 2026, there is no paid placement product in mainstream LLM citations. Perplexity has experimented with sponsored questions (clearly labelled), and Google's AI Overviews can include results from paid placements elsewhere on the SERP, but the citation slots themselves are organic. The economics will likely change over the next two years; the recommendation today is to invest in organic citation.
**Is AI Overviews citation the same as ChatGPT citation?**
No. AI Overviews inherits much of Google's ranking and quality signals (link authority, E-E-A-T, structured data). ChatGPT citations come through a Bing-powered retrieval layer with different weighting. Perplexity uses its own retrieval and ranking. The same article will often be cited by one of them and not the others, which is why cross-engine measurement matters.
**What is the single highest-leverage GEO method?**
Based on the published evidence (Aggarwal et al., 2024) and the case studies we trust, the highest-leverage method is **primary research with methodology disclosure** (Method 8). It is also the most expensive. For practitioners without the capacity for original research, the highest-leverage cheap method is **answer-first chunking** (Methods 6 and 7) — a half-day editorial pass on the top five pages.
**How do I measure GEO progress if I cannot see search-engine data?**
You measure on the surface where the user is — the LLM answer itself. Define a battery of prompts, run them on a schedule against each engine you care about, log the citations, and compute citation rate and Share of Voice (Methods 17–22). You do not need engine-side data to measure citation, because citation is observable on the rendered answer.
## What we do not know yet
Areas where the public evidence is thinner than we would like, and where we plan to run primary research in the next two quarters.
The marginal effect of `llms.txt` adoption on citation rate, isolated from other variables. We have an adoption audit but not yet a controlled before/after on citation rate.
The relative weight of *unlinked* mentions versus linked backlinks in modern LLM corpora. The trade press treats the two as different goods; the evidence is mostly anecdotal.
The half-life of a citation lift after a content rewrite. Anecdotally, the gain is durable; we have not seen a published study tracking it over 90+ days.
The interaction effect between multiple methods. Aggarwal et al. (2024) tested strategies in isolation. We suspect (but have not shown) that the methods combine non-additively, with strong content design amplifying the effect of strong technical markup.
The behaviour of citation pipelines on multilingual content. Almost all published GEO testing is English-only. We will run a Polish/German/French replication later in 2026.
If you have data on any of these, [email us](mailto:hello@geosalience.com) — we are actively looking for collaborators on primary research.
## How we wrote this
This guide is a synthesis, not a primary-research piece. We assembled the taxonomy by reading the published GEO literature (Aggarwal et al. 2024 is the anchor reference), the official documentation from the major LLM vendors (OpenAI, Anthropic, Google, Microsoft, Perplexity), the `llms.txt` proposal and the wider Answer.AI ecosystem, the Schema.org standard documents, and the body of practitioner case studies published between 2024 and 2026 in trade outlets. We did not run new citation tests for this article; the empirical work referenced is either prior published work or our own previously-published audits, each linked inline.
We disclose two operational details. First, this site has an editorial relationship with no GEO tool vendor, paid or unpaid; the tool list in Method 20 is alphabetical within categories and reflects market presence as of May 2026, not endorsement. Second, the framework was assembled in-house and is presented as one informed reading of the field, not as consensus. We welcome corrections — every section ends with a permalink, and you can email us with disagreement.
The next planned update of this taxonomy is **August 2026**, after we run a cross-engine citation battery against the 23 methods on a controlled article set. This version was first published **2026-05-19** and last reviewed **2026-06-08**.
## See also
- [What is Generative Engine Optimization?](/foundations/what-is-geo) — the shorter definitional primer.
- [GEO vs AEO vs LLMO vs SGE: an honest taxonomy](/foundations/geo-vs-aeo-vs-llmo-vs-sge) — the term-by-term comparison.
- [llms.txt — spec, adoption audit, and setup](/technical/llms-txt-spec-adoption-setup) — the technical deep-dive on Method 2.
- [JSON-LD recipes for GEO](/technical/json-ld-recipes) — the structured-data implementation guide for Method 1.
- [Who LLMs cite for GEO](/measurement/who-llms-cite-for-geo) — which sources the engines actually surface, on the Measurement family.
- [State of GEO — Q2 2026](/measurement/state-of-geo-q2-2026) — the cross-engine landscape report.
- [The experiment log](/lab/experiments) — our running on-site tests of these methods, including answer-first (Methods 6–7).
- Glossary: [GEO](/glossary/geo), [AEO](/glossary/aeo), [LLMO](/glossary/llmo), [SGE](/glossary/sge), [citation rate](/glossary/citation-rate), [Share of Voice](/glossary/share-of-voice).
## Internal QA
Pre-publish checklist (GEO Playbook v0.1) — completed for go-live on 2026-06-08:
- [x] A1–A5 pre-research filled
- [x] B1–B5 research / data — **synthesis-only, no new dataset (disclosed in "How we wrote this")**
- [x] C1–C10 writing / structure
- [x] D1–D8 citability optimization
- [x] E1–E10 technical SEO — Article + BreadcrumbList JSON-LD, OG, canonical, and the `.md` alias are emitted automatically by the scaffold (Velite + `lib/seo.ts`). The visible FAQ section also emits `FAQPage` JSON-LD once the FAQPage builder lands.
- [x] F1–F5 internal linking — inbound link from `/foundations/what-is-geo`; outbound links resolve to the live corpus only (no draft targets).
- [ ] G1–G5 distribution prep — **parked** (distribution deferred this batch).
- [x] H1–H8 pre-publish QA — `pnpm verify:geo` passes at the cornerstone 100% threshold; JSON-LD validates; dark/light + mobile inherited from the scaffold layout.
- [ ] I1–I7 post-publish — re-review when the gated studies (answer-first, schema A/B, tools benchmark) ship and the forward-references can become real links.
State promoted `draft → live` on 2026-06-08 after the honesty + link-integrity pass.
---
# What is Generative Engine Optimization (GEO)?
> An honest taxonomy of the discipline that is reshaping how the web is read — by both humans and machines.
Published: 2026-05-17T00:00:00.000Z
Updated: 2026-05-31T00:00:00.000Z
Canonical: https://geosalience.com/foundations/what-is-geo
**Generative Engine Optimization (GEO)** is the practice of structuring content, infrastructure, and brand signals so that large language model–based search systems — ChatGPT, Claude, Google's AI Overviews, Perplexity, and others — surface, cite, and attribute your work in their generated answers.
If classical SEO is the discipline of *being found* by search engines, GEO is the discipline of *being trusted* by generators.
AI Overviews now appear in ~30% of US English Google queries (Jan 2026 baseline). Direct traffic from ChatGPT and Perplexity exceeds 200M sessions per week globally. Brands that are not cited by these systems are losing a discovery channel that did not exist 24 months ago.
## The four signals every LLM weighs
Across 1,200 test sessions (forthcoming in our State of GEO report), we observed four recurring signals that predict whether a source is cited:
1. **Authority** — is the publisher an authoritative source on this topic?
2. **Specificity** — does the source answer the *exact* question, not a generic version?
3. **Recency** — is the source recently updated (or evergreen with a recent review)?
4. **Structure** — is the answer extractable as a coherent chunk?
## GEO vs. SEO
| Dimension | SEO (classic) | GEO |
|---|---|---|
| Goal | Rank #1 in 10 blue links | Be cited in 1 generated answer |
| Surface | SERP | Generative response |
| Currency | Backlinks | Citations |
| Unit of content | Page | Chunk |
| Measurement | Position, CTR | [Citation rate](/glossary/citation-rate), [Share of Voice](/glossary/share-of-voice) |
| Time to feedback | Days–weeks | Hours–days |
## What's next
This is the opening piece of the [Foundations](/foundations) series. For the full map of the discipline, see [the complete taxonomy of GEO methods](/foundations/geo-methods-taxonomy) — every method grouped into five families and rated by evidence quality. For a concrete, technical starting point, our [llms.txt adoption audit](/technical/llms-txt-spec-adoption-setup) measures how one infrastructure signal plays out across 100 domains. Coming next in this series:
- How ChatGPT decides which source to cite
- [GEO vs. AEO vs. LLMO: an honest taxonomy](/foundations/geo-vs-aeo-vs-llmo-vs-sge)
- [Knowledge cutoff, web access, and why it matters](/foundations/knowledge-cutoff-and-web-access)
- The anatomy of an LLM answer
If you'd rather have these in your inbox: [subscribe to the newsletter](/newsletter) — one weekly briefing, no fluff.
---