Sign inCreate a free account
Theme

Site crawl

Discover and analyse a whole site, and find the problems that repeat.

A crawl starts at one URL, follows the internal links, and analyses what it finds. It is how you audit a site rather than a page, and how you find the pages nobody remembered were there.

A finished crawl: every page found, with its score and finding counts.

Starting a crawl

Give it a starting URL, usually the home page, and a page limit. It follows internal links only: external links are checked for whether they resolve but never crawled, so a crawl cannot wander off your site.

PlanPages per crawlCrawls per month
Free, BasicNot available
Pro10010
Agency50030

Why the counts do not match

A crawl in progress often reads something like "157 of 90 pages", which looks wrong and is not.

The first number is URLs discovered, the second is your page limit. A crawler finds links faster than it can fetch pages, so the queue runs ahead of the work. It stops fetching at your limit; the extra discovered URLs are listed as found but not analysed, which is itself useful, because it tells you how much bigger the site is than the slice you looked at.

Inside one crawl

Opening a crawl from the list is where the work is. The list says a crawl happened; this says what it found.

One crawl opened: the summary tiles, the site-wide checks, and every page found.

Four figures sit across the top, and three of them are the ones to read first.

TileMeans
Average quick scoreThe mean across every page that could be scored. Useful as a baseline to watch over time rather than as a number to chase.
BrokenPages that answered 4xx or 5xx. These are the first thing to fix: a broken page wastes every link pointing at it and every crawl that reaches it.
Need workPages scoring under 50, or carrying something critical. A short list on a healthy site, and the one worth working down.
OrphansPages in your sitemap that nothing on the site links to. Search engines treat a page nothing links to as a page nothing vouches for.

Site issues

Below the tiles is a set of checks that only exist across pages. A single report cannot tell you that nothing links to a page, or that two pages claim the same title, because both of those are facts about the site rather than about the page.

Looks forWhy it needs a crawl
Pages nothing links toBeing in the sitemap is a claim; being linked is a signal. Only a crawl knows which pages have neither.
Duplicate titles and descriptionsTwo pages with the same title compete with each other. One report sees one title and has nothing to compare it against.
Broken internal linksThe page that is broken cannot tell you what links to it. The crawl can, because it arrived from there.
Pages buried deepDepth is measured from the starting URL, so it only exists once you have followed the links.

Quick scores

Each crawled page carries a quick score on the same 0 to 100 scale as a full report. It is an estimate from the single fetch the crawler made, which is what makes crawling hundreds of pages affordable, and it is marked as an estimate wherever it appears.

A quick score is not a full report and does not replace one. It ranks pages against each other so you know where to look; pressing Details runs the full analysis on that page, with every check and the evidence behind it.

Reading the results

ColumnWhat it shows
URLThe page, linking to its full report.
StatusThe HTTP status it returned. Non-200 pages are listed rather than hidden.
ScoreIts overall score, for pages that were analysed.
FindingsCritical and recommended counts.
DepthHow many links from the starting URL. Deep pages are crawled rarely by search engines too.

What a crawl is for

A single report tells you about a page. A crawl tells you about a site, and the useful findings are the ones that repeat.

PatternMeans
The same finding on nearly every pageA template or site-wide configuration problem. One fix, whole site. These are the best value findings a crawl produces.
Pages at depth 4 and beyondContent search engines will crawl rarely. If it matters, it needs a shorter path to it.
Discovered pages you did not know aboutOld campaign pages, staging pages that got linked, parameter variants of the same content. Each one is either worth fixing or worth removing.
A cluster of 404sInternal links pointing at pages that have gone. Cheap to fix and directly wasteful of crawl budget.

How we crawl

We obey robots.txtPages you have disallowed are not fetched. If a crawl finds far less than you expected, read your robots.txt first.
We rate limit ourselvesRequests are paced so a crawl does not behave like an attack on your own server. A large crawl takes minutes rather than seconds, deliberately.
We identify ourselvesThe crawler sends its own user agent, so you can see it in your logs and allow or block it as you choose.

Re-crawling

Crawls are kept, so you can re-crawl and compare against the last one: pages that appeared, pages that went, and pages whose score moved. Each crawl counts against your monthly crawl allowance, and the pages it analyses count against your report allowance.

Type to search. Try a check name from your report, an endpoint, or what you are trying to do.

↑↓ to move ↵ to open Esc to close