Original research · 7 September 2026
The 2026 AEO benchmark.
We pointed our scanner at 75 well-known websites and recorded which signals answer engines rely on were actually present. Here is what came back, with the corpus and the method published so you can check it.
How AEO-ready are large websites in 2026?
We ran an answer engine optimization scan across 75 well-known websites in September 2026 and got results from 55. The median score was 55 out of 100. Not one site — zero of 55 — had an explicit AI crawler policy in robots.txt. Only 4% carried FAQ structured data and 4% described what they sell in Service or Product schema. These are large organisations with in-house SEO teams; the long tail is worse.
55
median score /100
13–91
full range
55
sites scanned of 75
0
with an AI crawler policy
What stood out
Four findings.
0of 55
have an explicit AI crawler policy
Not one site in the corpus names GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot or Google-Extended in robots.txt. Every one of them is letting AI crawlers in — or out — by accident rather than decision. This was the most one-sided result in the study and the cheapest thing on this list to fix.
20of 75
blocked our scanner entirely
Roughly a quarter returned 403 or refused the connection to a politely identified, single-page request. Bot protection does not usually distinguish between our scanner and an AI search crawler, so there is a real chance these sites are invisible to answer engines for the same reason they were invisible to us — and nobody inside those companies would know.
4%
say what they sell in structured data
Service or Product schema is how a machine learns what a business actually offers. Almost nobody publishes it, which means engines are inferring the category from prose — and inference is where competitors get substituted for you.
24%
have an H1 that names their category
Three quarters lead with a slogan. A headline that describes a feeling gives a model nothing to attach the brand to when somebody asks for that category by name — the single most common failure we see, and the one most likely to be defended as branding.
Every check
Pass rate by signal, worst first.
Warn means partially present — a robots.txt that exists but names no AI crawler, an H1 that reads as a slogan. Fail means absent.
Explicit AI crawler policy
0% pass0 pass · 41 warn · 14 fail
robots.txt names GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot or Google-Extended — allowing or blocking them deliberately rather than by omission.
FAQPage structured data
4% pass2 pass · 53 warn · 0 fail
Question-and-answer markup, the format engines lift most readily.
Service or Product schema
4% pass2 pass · 0 warn · 53 fail
Structured data saying what the business actually sells.
Category-bearing H1
24% pass13 pass · 32 warn · 10 fail
A headline that names what the business is, not only what it promises.
sameAs corroboration
38% pass21 pass · 4 warn · 30 fail
Organization markup linking to independent profiles that confirm the entity exists.
llms.txt
38% pass21 pass · 34 warn · 0 fail
A plain-text file describing the site for AI systems.
Organization schema
44% pass24 pass · 0 warn · 31 fail
Structured data identifying the business as an entity.
Meta description
78% pass43 pass · 5 warn · 7 fail
Often the sentence an engine quotes back when describing a business.
Content without JavaScript
85% pass47 pass · 2 warn · 6 fail
Body text present in the served HTML, which is what retrieval reads.
Page title
85% pass47 pass · 6 warn · 2 fail
A descriptive, non-placeholder title element.
Score distribution
31 of 55 scored under 60.
Median by sector
Directional only — four to eight sites each. A hint about where to look, not a league table.
- SaaS n=8
- 79
- B2B services n=4
- 72
- Financial services n=4
- 67
- Education n=4
- 67
- Real estate n=4
- 62
- Accounting n=5
- 60
- Law firms n=4
- 55
- Recruitment n=4
- 55
- Agencies n=5
- 47
- Ecommerce n=5
- 45
Excluded for too few results: Healthcare (n=2), Manufacturing (n=3), Home services (n=2), Hospitality (n=1).
Method
How this was run, and what's wrong with it.
Published in full, including the limitations. A statistic nobody can reproduce is an assertion.
- What exactly was measured?
- Ten checks on each site's homepage, plus its robots.txt and llms.txt: Organization schema, Service or Product schema, FAQPage schema, sameAs links, a category-bearing H1, a meta description, a page title, server-rendered body text, an explicit AI crawler policy, and the presence of llms.txt. Each is weighted and combined into a score out of 100. The same scan runs free at rankvyze.com for any URL.
- How was the corpus chosen?
- 75 well-known organisations across fourteen sectors, chosen to mirror the industries we publish guides for. The full list is public in the repository, so anyone can re-run the study and check these numbers.
- What are the limitations?
- Three worth stating. The corpus skews to large, well-resourced companies, so it flatters the wider web rather than representing it. 20 sites could not be scanned and are excluded from every figure, which may bias results if blocked sites differ systematically. And only the homepage was scanned — a site can have excellent structured data on inner pages and score poorly here.
- Why are some sectors missing from the table?
- Sectors with fewer than 4 successfully scanned sites are excluded, because a median over one or two sites is not a median. Even the sectors shown are directional rather than statistically robust — treat them as a hint about where to look, not a league table.
- Is this a one-off?
- It is a snapshot, dated. The plan is to re-run it against the same corpus periodically so the interesting number becomes the change rather than the level. If you want to be told when it updates, the contact page reaches a person.
The corpus
All 75 domains, in the order they were scanned. Per-site scores are deliberately not published — the point is the aggregate, and nobody here agreed to be an example.
stripe.com · notion.so · figma.com · linear.app · vercel.com · intercom.com · asana.com · airtable.com · ogilvy.com · dentsu.com · publicisgroupe.com · wpp.com · rga.com · linklaters.com · freshfields.com · dlapiper.com · cliffordchance.com · hoganlovells.com · allbirds.com · glossier.com · warbyparker.com · gymshark.com · everlane.com · casper.com · mayoclinic.org · clevelandclinic.org · hopkinsmedicine.org · teladoc.com · onemedical.com · pwc.com · deloitte.com · kpmg.com · ey.com · bdo.com · grantthornton.com · zillow.com · redfin.com · compass.com · savills.com · knightfrank.com · angi.com · thumbtack.com · houzz.com · checkatrade.com · indeed.com · roberthalf.com · hays.com · michaelpage.com · kornferry.com · revolut.com · monzo.com · wise.com · schwab.com · fidelity.com · coursera.org · udemy.com · khanacademy.org · edx.org · duolingo.com · masterclass.com · salesforce.com · zendesk.com · docusign.com · workday.com · servicenow.com · marriott.com · hilton.com · hyatt.com · ihg.com · airbnb.com · siemens.com · caterpillar.com · honeywell.com · 3m.com · bosch.com
Citing this
Free to quote with attribution (CC BY 4.0). If you use a figure, please link the page so readers can see the method it came from.
RankVyze, “The 2026 AEO Benchmark”, 7 September 2026. https://rankvyze.com/research/aeo-benchmark
Curious where your own site sits against that median?