Indexability and crawling
73 checks: crawl directives, canonical, redirects, security and site files.
The first question about any page is whether a search engine can reach it, read it, and is allowed to keep it. Nothing else on the report matters if the answer is no, which is why this category counts for more than any other.
73 checks, in four groups.
Crawl directives and canonical
What you have told search engines they may do with this page, and whether those instructions agree with each other.
| Check | What it looks for |
|---|---|
robots.indexable | Whether the page is indexable at all, from the meta robots tag, the X-Robots-Tag header and robots.txt together. A noindex on a page you want ranked is the single most expensive finding we report. |
canonical.missing | No canonical tag. Search engines will pick one for you, and they do not always pick the one you wanted. |
canonical.not_self_referential | The canonical points somewhere other than this page. Deliberate on a duplicate, a serious error anywhere else, because it tells search engines to rank the other page instead of this one. |
canonical.cross_domain | The canonical points at another domain entirely. |
canonical.protocol_mismatch | The canonical uses a different scheme from the page serving it, usually http on an https page. |
lang.missing, lang.invalid | The lang attribute on <html>, and whether it is a real language code. |
hreflang.* | For multilingual sites: invalid codes, duplicate codes, and the missing self-reference that quietly breaks a whole hreflang cluster. |
charset.missing, charset.not_utf8 | A declared character set, and whether it is UTF-8. |
viewport.missing, viewport.not_responsive | The viewport meta tag. Without it a page is treated as not mobile ready, and indexing is mobile first. |
viewport.zoom_disabled | A viewport that blocks pinch zoom. An accessibility failure, and one browsers increasingly ignore anyway. |
favicon.missing | A favicon, which appears beside your result in some search surfaces. |
Canonical and robots findings are worth reading together. A page that is noindexed and also canonicalised elsewhere is sending two different instructions, and search engines resolve that conflict in ways that are hard to predict. Pick one.
Transport and redirects
| Check | What it looks for |
|---|---|
http.status | The status code the page finally returned. |
https.not_served | The page is served over plain http. |
redirect.chain_length | How many hops before the page arrived. Every hop costs crawl budget and a little of the signal being passed along. |
redirect.loop | A redirect that never resolves. The page is unreachable. |
redirect.protocol_downgrade | A redirect that steps from https back to http, even briefly. |
mixed_content.active | Scripts or stylesheets loaded over http on an https page. Browsers block these, so part of the page is simply not running. |
mixed_content.passive | Images or media over http on an https page. Not blocked, but it removes the padlock. |
Security headers and certificate
These are not ranking factors in the direct sense, but a browser warning between a visitor and your page costs you more than any tag on it. HTTPS itself is a confirmed signal; the rest is hygiene a crawler and a visitor both notice.
| Check | What it looks for |
|---|---|
security.certificate_expired | An expired certificate. Every visitor gets a full page browser warning. |
security.certificate_expiring | A certificate close enough to expiry to act on now. |
security.hsts_missing, security.hsts_short | The Strict-Transport-Security header, and whether its max-age is long enough to be worth having. |
security.csp_missing, security.csp_unsafe | A content security policy, and whether it permits unsafe-inline or unsafe-eval, which removes most of the point of having one. |
security.clickjacking_risk | Nothing stopping the page being framed by somebody else. |
security.no_sniff_missing | No X-Content-Type-Options: nosniff. |
security.referrer_policy_missing | No referrer policy, so full URLs leak to other sites. |
security.version_disclosure | Server or framework version numbers advertised in headers. |
Site level files
These belong to the site rather than the page, and are reported on every page of it because their effect lands on every page.
| Check | What it looks for |
|---|---|
robots_txt.missing | No robots.txt at the site root. |
robots_txt.blocks_everything | A robots.txt disallowing everything. Usually a staging file that reached production, and it removes the whole site from search. |
sitemap.missing | No XML sitemap found. |
sitemap.invalid_xml | A sitemap that does not parse, so nothing in it is read. |
sitemap.not_in_robots | A sitemap that exists but is not referenced from robots.txt, which is how crawlers find it without being told. |
llms_txt.missing | No llms.txt. An emerging convention for telling AI crawlers what your site is and which pages matter. A notice, not a problem. |
URL shape
| Check | What it looks for |
|---|---|
url.too_long | Over 100 characters. Truncated in results and awkward to share. |
url.too_deep | More than five path segments from the root. Depth correlates with how rarely a page is crawled. |
url.underscores | Underscores as word separators. Hyphens are read as word breaks and underscores are not. |
url.uppercase | Uppercase letters, which invite duplicate URLs on case sensitive servers. |
url.encoded_characters | Percent encoding in the path, usually spaces or non-ASCII characters. |
url.many_parameters | Several query parameters, each a way for the same content to have several addresses. |
URL findings are notices for a reason. Changing a live URL means redirects, lost history and updated links, and that is rarely worth it for an underscore. Use these when choosing new URLs, not as a reason to rewrite old ones.