Sign inCreate a free account
Theme

Indexability and crawling

73 checks: crawl directives, canonical, redirects, security and site files.

The first question about any page is whether a search engine can reach it, read it, and is allowed to keep it. Nothing else on the report matters if the answer is no, which is why this category counts for more than any other.

73 checks, in four groups.

Crawl directives and canonical

What you have told search engines they may do with this page, and whether those instructions agree with each other.

CheckWhat it looks for
robots.indexableWhether the page is indexable at all, from the meta robots tag, the X-Robots-Tag header and robots.txt together. A noindex on a page you want ranked is the single most expensive finding we report.
canonical.missingNo canonical tag. Search engines will pick one for you, and they do not always pick the one you wanted.
canonical.not_self_referentialThe canonical points somewhere other than this page. Deliberate on a duplicate, a serious error anywhere else, because it tells search engines to rank the other page instead of this one.
canonical.cross_domainThe canonical points at another domain entirely.
canonical.protocol_mismatchThe canonical uses a different scheme from the page serving it, usually http on an https page.
lang.missing, lang.invalidThe lang attribute on <html>, and whether it is a real language code.
hreflang.*For multilingual sites: invalid codes, duplicate codes, and the missing self-reference that quietly breaks a whole hreflang cluster.
charset.missing, charset.not_utf8A declared character set, and whether it is UTF-8.
viewport.missing, viewport.not_responsiveThe viewport meta tag. Without it a page is treated as not mobile ready, and indexing is mobile first.
viewport.zoom_disabledA viewport that blocks pinch zoom. An accessibility failure, and one browsers increasingly ignore anyway.
favicon.missingA favicon, which appears beside your result in some search surfaces.

Canonical and robots findings are worth reading together. A page that is noindexed and also canonicalised elsewhere is sending two different instructions, and search engines resolve that conflict in ways that are hard to predict. Pick one.

Transport and redirects

CheckWhat it looks for
http.statusThe status code the page finally returned.
https.not_servedThe page is served over plain http.
redirect.chain_lengthHow many hops before the page arrived. Every hop costs crawl budget and a little of the signal being passed along.
redirect.loopA redirect that never resolves. The page is unreachable.
redirect.protocol_downgradeA redirect that steps from https back to http, even briefly.
mixed_content.activeScripts or stylesheets loaded over http on an https page. Browsers block these, so part of the page is simply not running.
mixed_content.passiveImages or media over http on an https page. Not blocked, but it removes the padlock.

Security headers and certificate

These are not ranking factors in the direct sense, but a browser warning between a visitor and your page costs you more than any tag on it. HTTPS itself is a confirmed signal; the rest is hygiene a crawler and a visitor both notice.

CheckWhat it looks for
security.certificate_expiredAn expired certificate. Every visitor gets a full page browser warning.
security.certificate_expiringA certificate close enough to expiry to act on now.
security.hsts_missing, security.hsts_shortThe Strict-Transport-Security header, and whether its max-age is long enough to be worth having.
security.csp_missing, security.csp_unsafeA content security policy, and whether it permits unsafe-inline or unsafe-eval, which removes most of the point of having one.
security.clickjacking_riskNothing stopping the page being framed by somebody else.
security.no_sniff_missingNo X-Content-Type-Options: nosniff.
security.referrer_policy_missingNo referrer policy, so full URLs leak to other sites.
security.version_disclosureServer or framework version numbers advertised in headers.

Site level files

These belong to the site rather than the page, and are reported on every page of it because their effect lands on every page.

CheckWhat it looks for
robots_txt.missingNo robots.txt at the site root.
robots_txt.blocks_everythingA robots.txt disallowing everything. Usually a staging file that reached production, and it removes the whole site from search.
sitemap.missingNo XML sitemap found.
sitemap.invalid_xmlA sitemap that does not parse, so nothing in it is read.
sitemap.not_in_robotsA sitemap that exists but is not referenced from robots.txt, which is how crawlers find it without being told.
llms_txt.missingNo llms.txt. An emerging convention for telling AI crawlers what your site is and which pages matter. A notice, not a problem.

URL shape

CheckWhat it looks for
url.too_longOver 100 characters. Truncated in results and awkward to share.
url.too_deepMore than five path segments from the root. Depth correlates with how rarely a page is crawled.
url.underscoresUnderscores as word separators. Hyphens are read as word breaks and underscores are not.
url.uppercaseUppercase letters, which invite duplicate URLs on case sensitive servers.
url.encoded_charactersPercent encoding in the path, usually spaces or non-ASCII characters.
url.many_parametersSeveral query parameters, each a way for the same content to have several addresses.

URL findings are notices for a reason. Changing a live URL means redirects, lost history and updated links, and that is rarely worth it for an underscore. Use these when choosing new URLs, not as a reason to rewrite old ones.

Type to search. Try a check name from your report, an endpoint, or what you are trying to do.

↑↓ to move ↵ to open Esc to close