Site crawl
Discover and analyse a whole site, and find the problems that repeat.
A crawl starts at one URL, follows the internal links, and analyses what it finds. It is how you audit a site rather than a page, and how you find the pages nobody remembered were there.
Starting a crawl
Give it a starting URL, usually the home page, and a page limit. It follows internal links only: external links are checked for whether they resolve but never crawled, so a crawl cannot wander off your site.
| Plan | Pages per crawl | Crawls per month |
|---|---|---|
| Free, Basic | Not available | |
| Pro | 100 | 10 |
| Agency | 500 | 30 |
Why the counts do not match
A crawl in progress often reads something like "157 of 90 pages", which looks wrong and is not.
The first number is URLs discovered, the second is your page limit. A crawler finds links faster than it can fetch pages, so the queue runs ahead of the work. It stops fetching at your limit; the extra discovered URLs are listed as found but not analysed, which is itself useful, because it tells you how much bigger the site is than the slice you looked at.
Inside one crawl
Opening a crawl from the list is where the work is. The list says a crawl happened; this says what it found.
Four figures sit across the top, and three of them are the ones to read first.
| Tile | Means |
|---|---|
| Average quick score | The mean across every page that could be scored. Useful as a baseline to watch over time rather than as a number to chase. |
| Broken | Pages that answered 4xx or 5xx. These are the first thing to fix: a broken page wastes every link pointing at it and every crawl that reaches it. |
| Need work | Pages scoring under 50, or carrying something critical. A short list on a healthy site, and the one worth working down. |
| Orphans | Pages in your sitemap that nothing on the site links to. Search engines treat a page nothing links to as a page nothing vouches for. |
Site issues
Below the tiles is a set of checks that only exist across pages. A single report cannot tell you that nothing links to a page, or that two pages claim the same title, because both of those are facts about the site rather than about the page.
| Looks for | Why it needs a crawl |
|---|---|
| Pages nothing links to | Being in the sitemap is a claim; being linked is a signal. Only a crawl knows which pages have neither. |
| Duplicate titles and descriptions | Two pages with the same title compete with each other. One report sees one title and has nothing to compare it against. |
| Broken internal links | The page that is broken cannot tell you what links to it. The crawl can, because it arrived from there. |
| Pages buried deep | Depth is measured from the starting URL, so it only exists once you have followed the links. |
Quick scores
Each crawled page carries a quick score on the same 0 to 100 scale as a full report. It is an estimate from the single fetch the crawler made, which is what makes crawling hundreds of pages affordable, and it is marked as an estimate wherever it appears.
A quick score is not a full report and does not replace one. It ranks pages against each other so you know where to look; pressing Details runs the full analysis on that page, with every check and the evidence behind it.
Reading the results
| Column | What it shows |
|---|---|
| URL | The page, linking to its full report. |
| Status | The HTTP status it returned. Non-200 pages are listed rather than hidden. |
| Score | Its overall score, for pages that were analysed. |
| Findings | Critical and recommended counts. |
| Depth | How many links from the starting URL. Deep pages are crawled rarely by search engines too. |
What a crawl is for
A single report tells you about a page. A crawl tells you about a site, and the useful findings are the ones that repeat.
| Pattern | Means |
|---|---|
| The same finding on nearly every page | A template or site-wide configuration problem. One fix, whole site. These are the best value findings a crawl produces. |
| Pages at depth 4 and beyond | Content search engines will crawl rarely. If it matters, it needs a shorter path to it. |
| Discovered pages you did not know about | Old campaign pages, staging pages that got linked, parameter variants of the same content. Each one is either worth fixing or worth removing. |
| A cluster of 404s | Internal links pointing at pages that have gone. Cheap to fix and directly wasteful of crawl budget. |
How we crawl
We obey robots.txt | Pages you have disallowed are not fetched. If a crawl finds far less than you expected, read your robots.txt first. |
| We rate limit ourselves | Requests are paced so a crawl does not behave like an attack on your own server. A large crawl takes minutes rather than seconds, deliberately. |
| We identify ourselves | The crawler sends its own user agent, so you can see it in your logs and allow or block it as you choose. |
Re-crawling
Crawls are kept, so you can re-crawl and compare against the last one: pages that appeared, pages that went, and pages whose score moved. Each crawl counts against your monthly crawl allowance, and the pages it analyses count against your report allowance.