Site Crawler Example: Multi-Page SEO Crawl with Scores
This example shows the aiwebpageseo Site Crawler for a 20-page crawl of a marketing site. The average SEO score is 71/100. Three pages score below 50 and need urgent attention. Eight pages score above 90. The report shows every page with its score, error count and warning count.
| Page | Score | Errors | Warnings |
|---|---|---|---|
| /blog/old-post-no-meta | 34 | 5 | 8 |
| /landing/campaign-2024 | 41 | 4 | 6 |
| /about/team | 47 | 3 | 5 |
| URL | Score | Errors | Warnings |
|---|---|---|---|
| / | 94 | 0 | 1 |
| /features | 88 | 0 | 2 |
| /pricing | 91 | 0 | 1 |
| /blog/seo-guide-2026 | 78 | 1 | 3 |
| /blog/technical-seo | 72 | 2 | 4 |
| + 15 more pages... |
Twenty pages, average score 71. That single figure is compatible with two completely different sites: twenty uniformly mediocre pages, or seventeen strong pages and three that are badly broken. The work implied by each is entirely different, and the mean cannot distinguish them.
What a crawl is actually for is the list. A list of URLs with specific findings against each is a thing you can sort — and the sort that matters is not the one the tool applies. Tools sort by severity, because severity is the only thing they can see. You should sort by what the page is worth, because that is the thing that determines whether fixing it pays.
A blog post from three years ago scoring 38 is a low-value page with a low score. A pricing page scoring 68 is a high-value page with a middling one. Every automated report will put the first at the top. Almost every business should start with the second.
A crawler reports everything it can measure, which means a long list in which a handful of items genuinely matter and the rest are decoration. The test for each: does this cost something real?
Fix these — they have a cost and no upside
Pages returning errors. Pages accidentally noindexed or blocked. Canonicals pointing somewhere wrong. Internal links that lead nowhere. Structured data that contradicts the visible page. Each of these is something broken, and each has a concrete consequence for a user or a crawler.
Fix where it improves the page for a human
Duplicate or missing titles and descriptions. There is no penalty for duplicates — Google consolidates rather than punishing — but a title is the strongest short statement a page makes about itself, and six pages sharing one is six pages declining to say what makes them different. Worth writing properly; not worth panicking about.
Do not chase these
Word-count thresholds, keyword density, heading-count rules, meta keyword tags, and any check that scores a page against an arbitrary numeric ideal. Keyword density in particular is not a ranking factor and is a ratio that rises when you delete content — an optimisation target that rewards writing less is not a target.
The most common outcome of a diligent technical programme is a site that scores well and ranks no better. That is not a failure of the crawl. It is what happens when the crawl has already found everything it can find.
A page can be fast, indexable, canonical, well-marked-up, correctly linked and comprehensively described — and still be thinner than the page above it, vaguer where that page is specific, and silent where that page admits a limitation. None of that is machine-detectable, and none of it is fixed by another pass of the crawler.
So treat a clean crawl as the beginning rather than the achievement. Fix the defects, because they are cheap and their cost is real. Then close the tool, open the three pages that outrank you for the query you care about, and read them. That is where the answer is, and it is the only part of this process that nobody can automate for you.
Three items appear in nearly every crawl report and all three are routinely misread. Getting them right saves more time than any other change to how you use a tool like this.
Duplicate titles and descriptions. There is no duplicate content penalty and no duplicate title penalty. Google consolidates near-identical signals rather than punishing the site that has them. What duplication actually costs is precision: several pages announcing the same title are several pages telling Google they cover the same ground, which forces a guess about which to show — and it will sometimes guess wrong. Worth fixing so each page says what it is. Not worth fearing.
Missing canonical tags. A self-referential canonical is cheap insurance against parameter and variant URLs being treated as separate pages. But a canonical is a hint, not a directive: Google can and does select a different canonical when the signals disagree with your instruction. If two URLs genuinely need to be one page, make them one page; a canonical is not a substitute for architecture.
Thin pages. There is no word-count threshold, no minimum, and no penalty in the technical sense — nothing is applied to your site and nothing appears in Search Console. What happens is simpler: a short page competing against thorough answers is a worse answer, and it ranks accordingly. A 300-word page that answers a narrow question completely is not thin. A 1,500-word page that says nothing is.
Fix what is broken, on the pages that earn money
Errors, accidental noindex directives, canonicals pointing somewhere wrong, internal links to nothing, and any structured data that disagrees with the visible page. Start on the commercially important URLs, not on the ones with the lowest scores.
Give each page a title and description that describe it
Not to satisfy a checker, and not because duplicates are punished. Because a title is the strongest short claim a page makes, and it should be true and specific. Note that Google frequently rewrites titles in results when it judges yours a poor fit for the query, so the way to make yours stick is to make it accurate.
Link the pages that nothing links to
A sitemap says a URL exists; an internal link says it matters and what it is about. If a page is worth having, something on the site should point at it. If nothing does, that is a decision worth making explicitly rather than by neglect.
Then stop, and go and read your competitors
Once the defects are gone, the remaining difference between you and the pages above you is almost always the writing. No further crawl will surface it, and no score will tell you it is there.
A crawler follows links, which means the crawl itself is a map of what your site says about its own priorities. Two findings from that map are worth more attention than they usually get.
An orphan page is one that nothing on the site links to. It may sit in the sitemap and be perfectly good, but a sitemap only tells Google that a URL exists. An internal link does three things a sitemap cannot: it passes ranking signal from a page that already has some, it describes the destination through anchor text and surrounding context, and it places the page in a hierarchy. Orphans get discovery and none of the rest, which is why they consistently underperform even when the content is strong.
Click depth is the same story more gently. No threshold has ever been published — anybody quoting one is inventing it — but depth is a fair reading of what your architecture believes. Two clicks from the homepage says a page matters; six clicks, reachable only by paginating through a listing, says it does not, and search engines generally take the site at its word.
Both fixes are architectural rather than technical: a hub page, a category listing, a related-content module rendered in the served HTML rather than after a user interaction. A carousel that only populates client-side creates no crawlable link at all, which is how pages become orphans on sites that believe they have linked them.
How many pages can the Site Crawler audit?
Free plan crawls up to 5 pages. Silver plan crawls up to 50 pages. Gold plan crawls up to 500 pages. Diamond plan crawls up to 5,000 pages. For enterprise sites with tens of thousands of pages, contact aiwebpageseo for a custom crawl solution.
How does the crawler discover pages?
The crawler starts at the URL you provide, follows all internal links it finds, then follows links on those pages, recursively discovering the site. It respects robots.txt directives and does not crawl pages blocked by Disallow rules. It also checks your XML sitemap for additional URLs.
What is the site health score?
The site health score is the average of all individual page SEO scores across the crawled pages. A score of 80+ indicates a technically healthy site. Below 60 indicates significant issues affecting multiple pages. The crawler highlights the lowest-scoring pages first so you can prioritise fixes where they have the most impact.
What does an average score of 71 across 20 pages actually mean?
Very little on its own, because an average conceals exactly the pages you need to find. Twenty pages averaging 71 could be twenty mediocre pages, or seventeen strong ones and three that are badly broken — and those two situations call for entirely different work. The useful output of a crawl is never the mean; it is the list. Sort the list by how much each page matters commercially, not by how badly it scores, because a nine-point gain on a page nobody visits is worth less than a two-point gain on the page where people buy.
Are the three low-scoring pages necessarily the ones to fix first?
Only if they matter. Severity ranks pages by how badly they fail; it says nothing about what the failure costs. A blog post from 2021 scoring 38 is a low-priority page with a low score. A pricing page scoring 68 is a high-priority page with a middling one, and it deserves attention first. No crawler can make this judgement, because no crawler knows which of your URLs generate revenue. That is the one input you have to supply, and supplying it changes the order of nearly every audit.
Does fixing every crawler warning improve rankings?
No, and chasing a perfect score is one of the most reliable ways to waste a quarter. A crawl finds defects — broken links, missing canonicals, duplicated titles, invalid markup — and fixing defects is worth doing because defects have costs and no upside. But once a site is technically clean, further polish has sharply diminishing returns: there is no ranking bonus for a score of 100, and the pages above you in the results are almost never above you because their meta descriptions were better. The remaining headroom is in the content, and a crawler cannot see it.
What can a crawl not tell me?
Whether your pages are any good. It can confirm that a page is fast, indexable, well-marked-up and correctly linked, and every one of those can be true of a page that is thin, vague, and second-best to the one that outranks it. This is the uncomfortable finding at the end of most audits: the technical work is finished, the scores are green, and the reason a competitor ranks above you is that somebody there wrote a better answer. Twenty minutes reading the pages that beat you will tell you more than another crawl ever will.
Related Demo Reports
Run Site-Wide Crawler on Your Own Site
Get your real audit report with specific issues, fixes and actionable improvements.