Methodology · Version 1.0
AG AgentReady Research Methodology
How website readiness is scored, how the research corpus is assembled from anonymised scan results, and exactly how each published figure is calculated. Figures produced with this methodology appear on the State of AI Customer Readiness benchmark.
What the scanner measures
A scan reads up to five publicly accessible pages of the submitted website: the homepage plus, where they can be discovered from it, pages such as services, contact, about and pricing. Only published, publicly reachable content is analysed — nothing behind a login is read.
Deterministic checks defined in code extract observable signals from that content — headings, page titles and descriptions, action wording, contact routes, forms, review and credential references, semantic landmarks, JSON-LD structured data, agent-discovery signals and basic technical hygiene. Each check returns a 0–1 result, which may include partial credit.
Every score is computed in code from those checks. A language model is used only to rewrite the plain-language summary and recommendations; it never influences a score, a severity or a check result.
The four scores
- Revenue Readiness (45% of Overall) — whether a visitor can understand what is sold, where it is sold, and act on it: proposition and service clarity, service area, pricing signals, primary and above-the-fold calls to action, contact routes, quote, booking and lead-capture paths.
- Agent Readiness (30% of Overall) — whether a machine can parse and act on the site: semantic HTML landmarks, heading hierarchy, machine-readable contact details, labelled and parseable forms, unambiguous actions, JSON-LD structured data including LocalBusiness and Service types, agent-tool/WebMCP discovery signals, and technical hygiene.
- Trust & Clarity (25% of Overall) — whether the business looks credible and legible: testimonials, reviews, credentials, evidence of past work, complete contact details, business identity, policy pages, navigation clarity and content depth.
- Overall Website Readiness — the weighted combination of the three pillars above (45% revenue, 30% agent, 25% trust), rounded to a whole number.
Each pillar score is the weighted average of its checks, expressed 0–100. Check weights reflect how much a signal typically affects a real buying or agent interaction, and they are fixed in code so two scans of the same content produce the same score.
Score bands
- 90–100 Excellent
- 70–89 Good
- 50–69 Fair
- 0–49 Needs Improvement
Bands are presentation only. They group scores for reporting and never change how a score is calculated.
Corpus eligibility
A scan enters the research corpus only when its status is completed and all four scores — Overall, Revenue, Agent and Trust & Clarity — are present. Queued, in-progress and failed scans are excluded, as are completed scans missing any score.
One website, one row
Websites are frequently rescanned. To stop repeat scanning from skewing an average, the corpus keeps only the most recent completed scan per website, ranked by completion time. Every aggregate on the research pages therefore has a denominator of distinct websites, not scans.
Domain normalisation
Before deduplication, each domain is lowercased and a leading “www.” is removed, so example.com, WWW.Example.com and www.example.com are treated as one website. Other subdomains are treated as distinct websites, because they usually publish different content.
Exclusions
An internal exclusion list holds domains that must never appear in research aggregates — our own sites, preview environments and obvious test targets. Excluded domains are removed after normalisation and before any figure is calculated. Adding a domain to that list is an internal operation and does not affect the owner's report.
Means and medians
For each score we publish both the arithmetic mean and the median across the corpus, rounded to one decimal place. Both are shown because a small corpus with a few very low or very high sites moves the mean much more than the median; where the two diverge, the distribution is skewed.
Finding-category prevalence
Every finding belongs to one category: revenue, conversion, trust, clarity, structured data, agent readiness, WebMCP or technical. A website counts towards a category when it has at least one finding in that category that is not marked positive — that is, an unresolved issue rather than a passed check. Prevalence is that count of websites divided by the corpus size. A website with five unresolved findings in a category still counts once.
Stable check failure rates
Each deterministic check carries a stable, namespaced identifier in code — for example conversion.primary_cta, structured_data.local_business or agent_readiness.semantic_structure. Only findings carrying such an identifier, and flagged as benchmark-eligible, are aggregated per check.
For a given check, the denominator is the number of corpus websites where that check was actually observed, and the numerator is the number where it did not pass. We publish the denominator alongside every rate, and we do not rank checks observed on very few websites.
Free-form finding titles are never aggregated as checks. Only code-defined identities are counted, so a wording change to a finding title cannot silently create or split a statistic.
Check versioning
Each stored finding records the check version that produced it, currently version 1.0. If a check's logic changes materially, the version is incremented rather than the identifier being reused, so historical results stay interpretable. Findings recorded before stable identifiers existed carry no identifier and are excluded from per-check rates while still counting towards score aggregates.
Privacy and anonymisation
Research pages publish aggregates only. Domains, website addresses, report links, email addresses, IP addresses, finding evidence, enquiry submissions and payment records are never published, and are never queried by the browser: aggregation runs server-side and returns only computed figures.
We are not claiming the underlying records are anonymous. Operational data — the scanned address, the report and its findings — remains stored internally so your report link keeps working, as described in our privacy policy. What is anonymised is the research output: nothing published identifies an individual scanned website.
Sample bias
This is a self-selected corpus of websites whose owners or users chose to scan them. It should be read as a directional benchmark, not a census of the entire web.
Websites reach the corpus because somebody chose to scan them, often because they suspected a problem. That plausibly biases the corpus towards weaker sites, and towards the sectors and regions we reach. Nothing corrects for that bias, so figures should be read as directional.
Changelog
- 1.0 — Initial published methodology. Three weighted pillars (45/30/25), four score bands, latest-completed-scan-per-domain corpus, domain normalisation with www stripping, internal exclusion list, mean and median reporting, category prevalence counted per website, and stable check failure rates with published denominators.
Run a scan against this methodology
Every free scan uses exactly the checks described above and returns your Overall Website Readiness score with the three pillar scores.
Run Free Website Scan