Content and structure
63 checks: metadata, headings, keywords, answer readiness and AI crawler access.
Once a page can be indexed, the question is whether it says something worth ranking, in a shape a machine can follow. 63 checks, covering the tags that become your search result, the structure of the page itself, and whether AI crawlers can read any of it.
Title, description and social tags
These become your search result. Everything else on the page decides whether you rank; these decide whether anyone clicks.
| Check | What it looks for |
|---|---|
title.missing | No title, or an empty one. Critical: the title is the strongest on-page signal there is. |
title.length | A title long enough to say something and short enough to survive. |
title.pixel_width | The rendered width of the title in the font search results use. This is what decides truncation, and it is why a character count is not enough. |
description.missing | No meta description. Search engines will write one from the page, and it is usually worse than one you wrote. |
description.length | A description that fits the space it is given. |
opengraph.incomplete | Missing Open Graph tags, which decide how the page looks when it is shared into social apps and chat. |
twitter.card_missing, twitter.card_type | The Twitter card tags and whether the declared type matches what the page actually provides. |
Titles are measured in pixels, not characters. "Illinois Wine Institute" and "llllllll llll llllllllll" are the same number of characters and nowhere near the same width, so a character limit will tell you a title is safe when it is about to be cut in half.
Headings and depth
| Check | What it looks for |
|---|---|
heading.h1_missing | No H1. The page has no stated subject. |
heading.h1_multiple | More than one H1. Allowed in HTML5, still ambiguous about what the page is for. |
heading.level_skipped | An H2 followed by an H4. The outline has a hole in it, which matters for screen readers as much as for crawlers. |
heading.empty | A heading element with no text, often left by a template or an editor. |
content.thin | Not enough substantive text to rank for anything. Judged against what the page appears to be, not a fixed word count. |
content.low_text_ratio | Far more markup than text. Often a sign the real content arrives later via JavaScript. |
content.depth | Whether the page covers its subject or only mentions it. |
Keywords and topic
Not keyword density, which stopped being a useful measure a long time ago. These checks look at whether the page is recognisably about something.
| Check | What it looks for |
|---|---|
keywords.cloud | The terms the page actually emphasises, extracted from its text and headings. |
keywords.placement | Whether those terms appear where they count most: title, H1, opening paragraph. |
keywords.stuffing | A term repeated past the point of reading naturally. Actively penalised, and obvious to a reader long before it is obvious to an algorithm. |
keywords.natural | Term distribution consistent with writing rather than with optimising. |
Answer readiness
Search results increasingly extract an answer rather than sending a click. These checks look at whether your page is in a shape that can be quoted, and whether the quote would be attributed to you.
| Check | What it looks for |
|---|---|
readiness.scannability | Whether the page is broken into parts a machine can lift a passage from: headings, short paragraphs, lists. |
readiness.sparse_headings | Long stretches of text with no heading to name them. |
aeo.answer_paragraphs | Passages that directly answer a question near the heading that asks it. This is the shape a featured snippet is taken from. |
readiness.no_date, readiness.stale | A visible published or updated date, and whether it is recent enough to be trusted for a subject that changes. |
readiness.no_attribution | No author or organisation credited. Attribution is part of how an answer engine decides whether to cite you. |
readiness.no_answer_schema | No FAQ or QAPage structured data on a page shaped like questions and answers. |
Entities
| Check | What it looks for |
|---|---|
entity.subject_named | Whether the page names its subject explicitly enough to be linked to a known thing rather than guessed at. |
entity.no_publisher | No identified publisher. Search engines and answer engines both use publisher identity as a trust signal. |
AI crawler access
Large language models fetch pages both to train on and, increasingly, to
answer a question live with a citation. Your robots.txt decides
which of them may. We report what you have allowed rather than telling you
what to allow, because that is a business decision and not a technical
one.
| Crawler | Belongs to | What it does |
|---|---|---|
GPTBot | OpenAI | Collects pages for training. |
ChatGPT-User | OpenAI | Fetches a page live when somebody asks about it. |
OAI-SearchBot | OpenAI | Builds the index behind ChatGPT search. |
ClaudeBot, Claude-Web | Anthropic | Collects pages, and fetches live for citations. |
Google-Extended | Controls Gemini and AI Overviews use, separately from Google Search. | |
PerplexityBot | Perplexity | Indexes for answers that cite sources. |
Applebot-Extended | Apple | Controls Apple Intelligence use, separately from Siri and Spotlight. |
CCBot | Common Crawl | Builds the open dataset many models train from. |
Bytespider | ByteDance | Collects pages for training. |
ai.live_fetch_blocked is the one worth a second look. Blocking training crawlers is a reasonable position. Blocking the live fetchers as well means that when somebody asks an assistant about your product, it cannot read your page to answer, and it answers from whatever else it can reach. Those are different decisions and they are often made by accident in one line.
Client-side rendering
We read what the server sent. If the content only appears after JavaScript runs, these checks tell you, because a crawler that does not execute your JavaScript sees the same empty page we did.
| Check | What it looks for |
|---|---|
hydration.content_client_only | The server sent an effectively empty page and the content is assembled in the browser. |
hydration.content_mostly_client | Some content server-rendered, most of it not. |
hydration.title_client_only | The title is set by JavaScript. Crawlers that do not run it see whatever the template shipped, which is often the framework default. |
hydration.schema_client_only | Structured data injected after load, so it may never be read. |
hydration.server_rendered | The page arrived complete. This is the result you want. |