New to technical SEO on bigger sites? This tutorial walks you through when to choose Site Audit Pro over the standard audit, how to run your first crawl, and how to read the results without getting lost.
📊 Run Site Audit Pro Full Guide →Use the standard Site Audit for monthly health checks on sites up to 50 pages. Choose Pro when your site is larger, when you're preparing for a migration, or when you suspect duplicate content across many pages.
Enter your homepage URL and start the crawl. A 50-page audit finishes in about 5–8 minutes; a 200-page site takes 20–30. The crawler runs pages in parallel while respecting your server.
Pro scores LCP, CLS and INP for each page. Sort failing pages worst-first and start with the ones that get the most traffic — fixing a slow high-traffic page beats fixing ten pages nobody visits.
Rewrite any duplicate title tags so each describes its page uniquely. Consolidate or canonicalise duplicate content. Fix any structured-data errors so those pages become eligible for the rich results that still exist — bearing in mind that FAQ rich results were restricted to government and health sites in 2023 and HowTo rich results were removed entirely, so not every error is worth a sprint. And check the thing no validator checks: that the values in the markup match the values on the page.
After making changes, run Pro again and check the failing counts have dropped. Running it quarterly (with the standard audit monthly in between) keeps a large site healthy over time.
A domain-level score is a mean, and a mean is where the interesting cases hide. A site averaging 78 can contain a checkout page at 30 and forty blog posts at 95 — and the average will reassure you while the one page that takes money is the broken one.
Per-page results invert that in exactly the right way. Instead of a number to feel vaguely good or bad about, you get a list — and a list can be sorted by something the crawler cannot know but you can: which of these pages matters to the business.
That sort is the whole discipline. Severity ranks pages by how badly they fail. Commercial value ranks them by what the failure costs. Fixing the worst page on a site is rarely the right first move; fixing the worst page that people buy from always is, and no tool can make that call because no tool knows what you sell.
Step 4 asks you to rewrite duplicate titles, and it is right — but almost everybody does it for the wrong reason, which is fear.
There is no duplicate content penalty. There is no duplicate title penalty. Google consolidates near-identical signals; it does not punish the site that has them. Nothing is being deducted from you.
What duplication costs is precision. Six pages announcing the same title are six pages telling Google they cover the same ground, which forces a guess about which to show for a query — and it will sometimes guess wrong, or show none of them confidently. Each duplicate is also a wasted opportunity: a title is the strongest short claim a page makes about itself, and spending it on a template string says nothing.
Worth knowing while you rewrite them: Google frequently replaces titles in the results when it judges the supplied one a poor fit for the query, drawing instead on your H1 or other on-page text. So the title you write is a strong suggestion rather than a guarantee, and the way to make it stick is to make it accurately describe the page. Which was the only good reason to write one anyway.
Step 4 also asks you to fix structured-data errors so pages become eligible for rich results. Two refinements make that advice much more useful.
Not every rich result still exists. Product and Recipe enhancements do. FAQ rich results were restricted by Google to recognised government and health sites in 2023, and HowTo rich results were removed entirely. So a schema error on an FAQ block is worth correcting for accuracy and is not worth a sprint — there is no enhancement waiting on the other side of the fix.
No validator checks whether the markup is true. A Product block asserting £29.99 while the page displays £34.99 parses cleanly, has every required property, and passes every automated test in existence. It is also a false statement about your own product in machine-readable form — grounds for item disapproval in Google's shopping surfaces, and a broken promise to the shopper who clicked expecting the price they were shown.
So the check that earns its keep is a comparison rather than a validation: price against the rendered price, availability against the state of the Add to Cart button, headline against the H1. Missing markup costs you a feature. Contradictory markup costs you the listing.
Per-page Core Web Vitals are the most valuable thing Pro gives you, and there is one limitation to hold in mind when acting on them.
A crawler produces lab data: a simulated device on a simulated network, measured once. Google's ranking systems use field data — what real Chrome users, on real connections, actually experienced, aggregated over a rolling window. These diverge routinely in both directions. A page can post an excellent lab score while your customers on mobile networks suffer, and a mediocre lab score can sit on top of perfectly acceptable field data.
Neither is wrong; they are different measurements. Use the lab results for what they are good at — telling you which pages to investigate and which specific thing on them is slow — then check field data before concluding the problem is real, and remember that field data lags a fix by weeks because of that rolling window.
And keep the ranking claim honest while you do it. Core Web Vitals are a genuine ranking consideration and a small one, acting as a tie-breaker between pages of comparable relevance. Nobody outranks a better answer by being faster than it. The stronger argument for the work is that people abandon slow pages and lose their place when the layout shifts — which you can measure in your own analytics, without anybody's theory of the algorithm.
The most common outcome of a diligent technical programme is a site that scores well and ranks no better. That is not a failure of the crawl. It is what happens when the crawl has already found everything it can find.
A page can be fast, indexable, canonical, well-marked-up and correctly linked — and still be thinner than the page above it, vaguer where that page is specific, and silent where that page admits a limitation. None of that is machine-detectable, and none of it is fixed by another pass of the crawler.
So treat a clean audit as the floor rather than the achievement. Fix the defects, because they are cheap and their cost is real. Then close the tool, open the three pages that outrank you for the query you care about, and read them properly. That is where the answer is, and it is the one part of this process nobody can automate for you.
Because a domain-level score is an average, and averages hide the pages that are costing you money. A site scoring 78 overall can contain a checkout page at 30 and forty blog posts at 95, and the average tells you to relax while your most valuable page is the broken one. Per-page results give you a list, and a list can be sorted by something no crawler knows: which of these URLs matters to the business.
Only if they matter. Severity ranks pages by how badly they fail; it says nothing about what the failure costs. A blog post from 2021 scoring 38 is a low-priority page with a low score; a pricing page scoring 68 is a high-priority page with a middling one, and it deserves attention first. That judgement is the one input you have to supply, and supplying it reorders almost every audit.
No. There is no duplicate content penalty and no duplicate title penalty — Google consolidates near-identical signals rather than punishing the site that has them. What duplication actually costs is precision: several pages announcing the same title are several pages claiming the same ground, which forces Google to guess which to show, and it will sometimes guess wrong. Fix them because each page deserves an accurate description, not because a sanction is hanging over you.
Only where a rich result still exists. Product, Recipe and several others do. FAQ rich results were restricted to recognised government and health sites in 2023, and HowTo rich results were removed entirely — so schema errors on those types are worth correcting for accuracy, and are not worth a sprint. And valid is not the same as true: markup asserting a price the page does not show passes every validator and is grounds for item disapproval in Google's shopping surfaces.
Not necessarily. A crawler produces lab data — a simulated device on a simulated network, measured once. Google's ranking systems use field data: what real Chrome users on real connections actually experienced, aggregated over a rolling window. The two diverge routinely, so a page can post a strong lab score while your mobile customers suffer. Use the per-page lab results to decide which pages to investigate, and field data to decide whether the problem is real.
Quarterly is enough for most sites, with a lighter check monthly. The exception is any period of change — a migration, a replatform, a redesign — when you want a crawl before and after, because those are the moments when canonicals, hreflang, robots directives and structured data quietly disappear. Those four are what migrations lose most often, precisely because they are invisible in a browser and nobody clicks them during testing.
Whether your pages are any good. A page can be fast, indexable, canonical, well-marked-up and correctly linked, and still be thinner than the page above it, vaguer where that one is specific, and silent where that one admits a limitation. None of that is machine-detectable. This is the uncomfortable finding at the end of most technical programmes: the scores are green, and the reason a competitor outranks you is that somebody there wrote a better answer.
Pay as you go, no subscription. See exactly which pages need work.
Run Site Audit Pro →aiwebpageseo.com is a data-driven SEO and AEO (Answer Engine Optimisation) platform providing a free suite of technical website tools. Rather than relying on AI-theorised assumptions, the platform analyses live URL performance, delivering objective diagnostics, page speed metrics, CLS debugging, and site crawl data alongside actionable technical tutorials.