Original research · 7 September 2026

The 2026 AEO benchmark.

We pointed our scanner at 75 well-known websites and recorded which signals answer engines rely on were actually present. Here is what came back, with the corpus and the method published so you can check it.

How AEO-ready are large websites in 2026?

We ran an answer engine optimization scan across 75 well-known websites in September 2026 and got results from 55. The median score was 55 out of 100. Not one site — zero of 55 — had an explicit AI crawler policy in robots.txt. Only 4% carried FAQ structured data and 4% described what they sell in Service or Product schema. These are large organisations with in-house SEO teams; the long tail is worse.

55

median score /100

13–91

full range

55

sites scanned of 75

0

with an AI crawler policy

What stood out

Four findings.

0of 55

have an explicit AI crawler policy

Not one site in the corpus names GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot or Google-Extended in robots.txt. Every one of them is letting AI crawlers in — or out — by accident rather than decision. This was the most one-sided result in the study and the cheapest thing on this list to fix.

20of 75

blocked our scanner entirely

Roughly a quarter returned 403 or refused the connection to a politely identified, single-page request. Bot protection does not usually distinguish between our scanner and an AI search crawler, so there is a real chance these sites are invisible to answer engines for the same reason they were invisible to us — and nobody inside those companies would know.

4%

say what they sell in structured data

Service or Product schema is how a machine learns what a business actually offers. Almost nobody publishes it, which means engines are inferring the category from prose — and inference is where competitors get substituted for you.

24%

have an H1 that names their category

Three quarters lead with a slogan. A headline that describes a feeling gives a model nothing to attach the brand to when somebody asks for that category by name — the single most common failure we see, and the one most likely to be defended as branding.

Every check

Pass rate by signal, worst first.

Warn means partially present — a robots.txt that exists but names no AI crawler, an H1 that reads as a slogan. Fail means absent.

Explicit AI crawler policy

0% pass

0 pass · 41 warn · 14 fail

robots.txt names GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot or Google-Extended — allowing or blocking them deliberately rather than by omission.

FAQPage structured data

4% pass

2 pass · 53 warn · 0 fail

Question-and-answer markup, the format engines lift most readily.

Service or Product schema

4% pass

2 pass · 0 warn · 53 fail

Structured data saying what the business actually sells.

Category-bearing H1

24% pass

13 pass · 32 warn · 10 fail

A headline that names what the business is, not only what it promises.

sameAs corroboration

38% pass

21 pass · 4 warn · 30 fail

Organization markup linking to independent profiles that confirm the entity exists.

llms.txt

38% pass

21 pass · 34 warn · 0 fail

A plain-text file describing the site for AI systems.

Organization schema

44% pass

24 pass · 0 warn · 31 fail

Structured data identifying the business as an entity.

Meta description

78% pass

43 pass · 5 warn · 7 fail

Often the sentence an engine quotes back when describing a business.

Content without JavaScript

85% pass

47 pass · 2 warn · 6 fail

Body text present in the served HTML, which is what retrieval reads.

Page title

85% pass

47 pass · 6 warn · 2 fail

A descriptive, non-placeholder title element.

Score distribution

31 of 55 scored under 60.

0–39
11
40–59
20
60–79
19
80–100
5

Median by sector

Directional only — four to eight sites each. A hint about where to look, not a league table.

SaaS n=8
79
B2B services n=4
72
Financial services n=4
67
Education n=4
67
Real estate n=4
62
Accounting n=5
60
Law firms n=4
55
Recruitment n=4
55
Agencies n=5
47
Ecommerce n=5
45

Excluded for too few results: Healthcare (n=2), Manufacturing (n=3), Home services (n=2), Hospitality (n=1).

Method

How this was run, and what's wrong with it.

Published in full, including the limitations. A statistic nobody can reproduce is an assertion.

What exactly was measured?
Ten checks on each site's homepage, plus its robots.txt and llms.txt: Organization schema, Service or Product schema, FAQPage schema, sameAs links, a category-bearing H1, a meta description, a page title, server-rendered body text, an explicit AI crawler policy, and the presence of llms.txt. Each is weighted and combined into a score out of 100. The same scan runs free at rankvyze.com for any URL.
How was the corpus chosen?
75 well-known organisations across fourteen sectors, chosen to mirror the industries we publish guides for. The full list is public in the repository, so anyone can re-run the study and check these numbers.
What are the limitations?
Three worth stating. The corpus skews to large, well-resourced companies, so it flatters the wider web rather than representing it. 20 sites could not be scanned and are excluded from every figure, which may bias results if blocked sites differ systematically. And only the homepage was scanned — a site can have excellent structured data on inner pages and score poorly here.
Why are some sectors missing from the table?
Sectors with fewer than 4 successfully scanned sites are excluded, because a median over one or two sites is not a median. Even the sectors shown are directional rather than statistically robust — treat them as a hint about where to look, not a league table.
Is this a one-off?
It is a snapshot, dated. The plan is to re-run it against the same corpus periodically so the interesting number becomes the change rather than the level. If you want to be told when it updates, the contact page reaches a person.

The corpus

All 75 domains, in the order they were scanned. Per-site scores are deliberately not published — the point is the aggregate, and nobody here agreed to be an example.

stripe.com · notion.so · figma.com · linear.app · vercel.com · intercom.com · asana.com · airtable.com · ogilvy.com · dentsu.com · publicisgroupe.com · wpp.com · rga.com · linklaters.com · freshfields.com · dlapiper.com · cliffordchance.com · hoganlovells.com · allbirds.com · glossier.com · warbyparker.com · gymshark.com · everlane.com · casper.com · mayoclinic.org · clevelandclinic.org · hopkinsmedicine.org · teladoc.com · onemedical.com · pwc.com · deloitte.com · kpmg.com · ey.com · bdo.com · grantthornton.com · zillow.com · redfin.com · compass.com · savills.com · knightfrank.com · angi.com · thumbtack.com · houzz.com · checkatrade.com · indeed.com · roberthalf.com · hays.com · michaelpage.com · kornferry.com · revolut.com · monzo.com · wise.com · schwab.com · fidelity.com · coursera.org · udemy.com · khanacademy.org · edx.org · duolingo.com · masterclass.com · salesforce.com · zendesk.com · docusign.com · workday.com · servicenow.com · marriott.com · hilton.com · hyatt.com · ihg.com · airbnb.com · siemens.com · caterpillar.com · honeywell.com · 3m.com · bosch.com

Citing this

Free to quote with attribution (CC BY 4.0). If you use a figure, please link the page so readers can see the method it came from.

RankVyze, “The 2026 AEO Benchmark”, 7 September 2026. https://rankvyze.com/research/aeo-benchmark

Curious where your own site sits against that median?

Ready to become the answer?

Free scan in ten seconds. $99 to fix it — refunded in full if we don't.