Skip to content

Grading

The scale

Every score falls between 1.0 and 5.0, where 5.0 is best.

Score Meaning
5.0 Excellent, follows every recommendation
4.0 Good, some minor improvements possible
3.0 Acceptable, works, but should get better
2.0 Needs improvement, clear shortcomings
1.0 Critical, serious problems

Three levels

  1. Individual test: each test sets its own score by its own rules, described on that test's page.
  2. Category score: tests measuring the same thing are weighed together. The engine keeps sub-scores for integrity and security, performance, accessibility and standards compliance, and a test only contributes to the areas it actually measures.
  3. Overall score: the average of the individual tests.

A test that couldn't run isn't counted. It doesn't lower the score. It doesn't show either.

Why does the score vary between runs?

Small variations are normal. Common causes:

  • The website itself: dynamic content, campaign banners, A/B tests and personalisation serve different pages at different times
  • The network: latency and server load vary over the day
  • The test environment: tests run from different servers
  • External services: CDNs, ad networks and APIs respond at different speeds

The performance test therefore runs the page several times and uses the median. That removes the worst outliers, but not all variation.

Look at the trend

A single measurement says less than the movement over time. Going from 3.2 to 3.4 between two runs is probably noise; sitting steadily around 3.2 for six months and suddenly landing on 2.1 is not.

Which part of the website is tested?

In most cases the start page. The exceptions are:

  • The 404 test, which requests an address that doesn't exist
  • Standard files, which fetches robots.txt, sitemap and security.txt
  • The accessibility statement, which searches a number of pages to find the statement

For websites with an intermediate page, such as a language chooser, the real start page is tested.

See also