Most pages an agent fetches are judged in a single request, and most of those judgments are made on markup the agent never renders. This post defines a nine-criterion rubric called typespy, demonstrates it on three pages fetched on 2026-10-03, and argues that static HTML predicts most agent-readability criteria — but not computed styles.
typespy is defined here as a rubric, not shipped as a runnable tool. There is no install, no CLI, no package. What follows is a scoring specification you can apply by hand, by grep, or by handing the criteria to an agent that already has the bytes.
The name is deliberate. A rubric is a scoring sheet applied to what arrived, not what would have arrived if a browser had done more work.
Why a static HTML GET is the agent’s real vantage point
An agent that fetches a URL usually receives exactly one response body, and its entire model of the page is that body plus the response headers. According to the HTML Standard’s focus management section, focusability and traversal are defined over the DOM the browser constructs — a structure the static GET only approximates. That approximation is usually good enough, and the exceptions are the whole point of this rubric.
The distinction matters commercially. Figma’s MCP server hands agents a structured design context pulled from a canonical store: variables, components, layout, all typed. The Figma Help Center guide describes that pipeline as the intended path. Nothing equivalent exists for an arbitrary published page. There is no canonical store for the open web, so agents fall back to parsing whatever bytes arrive.
typespy is a rubric for how much signal is in those bytes.
The nine criteria
Seven of the nine are computable from the response body alone (with one conditional). Two require the rubric to be honest about its limits.
- Heading skeleton shape. Does the document form a tree — one h1, h2s under it, h3s under those — or a flat pile of styled divs?
- Body link-text contrast. Whether body links meet WCAG AA at 4.5:1 for normal-size text. Verify any claim with WebAIM’s contrast checker. Computed color values are not in the HTML, so this criterion scores as unmeasurable from static GET unless the author inlined the values.
- Citation topology. Count inline
[Author](url)links versus bare[N]markers with no local resolution. - Table-data honesty. A page claiming “20+ numeric rows” without a
<table>element scores zero on this criterion, regardless of how good the visual grid looks. - Word-versus-HTML token gap. Ratio of visible text tokens to total markup bytes.
:focus-visibledeclaration presence. Whether the stylesheet declares it. The default behavior is documented on MDN’s:focus-visiblepage, and the accessibility rationale is in the WCAG 2.2 focus-appearance Understanding doc.- Canonical tag presence. A single
<link rel="canonical">removes duplicate-URL ambiguity for the agent. - Heading-hierarchy gap. A jump from h1 straight to h3 with no h2 is a structural break an agent has to guess across.
- Byte-size-first-GET. The JS-shell teller: HTTP 200 with near-zero visible text means “real page, no signal.”
Which criteria a single GET can actually settle
Criteria 1, 3, 4, 5, 7, 8, and 9 are decidable from the response body. Criterion 6 is decidable only if the CSS is inline or linked and fetched. Criterion 2 is not decidable at all without a rendering engine, because getComputedStyle does not exist in a byte stream.
That asymmetry — seven decidable, one conditionally decidable, one not, with the conditional counting inside the seven — is the empirical core of this post.
The three specimens, measured 2026-10-03
Three pages were fetched on 2026-10-03, and every byte figure below is a measurement, not an estimate. Two of the three come from the same product, which is the interesting part: the same site can pass and fail the same rubric depending on route.
The design-agent.dev blog index returned 139,651 bytes on first GET. Its individual posts are leaner: the forced-colors audit at 26,300 bytes, the dark-mode blind spot at 29,778 bytes, and the hover blind-spot test at 19,461 bytes. Each carries a canonical tag in its response head.
The JS-shell negative example
Tabby’s configuration route returned HTTP 200 with a 352-byte shell and exactly 0 text characters after markup strip. That is a real, citable measurement from 2026-10-03, and the link is a live one: tabby.tabbyml.com/docs/configuration/. The same host does better elsewhere — the docs models route returned 24,995 bytes of real text, and the Tabby blog index returned 117,526 bytes with 8,249 visible text characters after the strip. The source project is at github.com/TabbyML/tabby.
Phrasing discipline matters here: this is ONE instance observed on ONE day. It is not a universal law about Tabby, and it is not a claim that the site is broken. It is a demonstration that route-level pre-rendering decisions produce route-level agent-readability outcomes, and that an agent’s first GET cannot tell the difference between “empty page” and “page that needs JavaScript.”
The Anthropic announcement page
Anthropic’s Agent Skills announcement returned 547,403 bytes. The page announces a standard for modular, machine-readable agent-instruction files. Large marketing pages routinely ship multi-hundred-kilobyte payloads, and the byte count alone does not prove what fraction is markup versus script — only that delivery is heavy.
The tension is worth naming anyway. A standard for machine-readable instruction files, announced on a page whose first GET weighs 547,403 bytes, is a natural test case for criterion 5.
Comparison: rubric scores across the three specimens
Design-agent.dev’s blog index and posts, and the two live Tabby routes and the Anthropic page, were verified during the audit for heading structure and canonical tags. Where a criterion not recorded in the audit exists, the cell says so explicitly rather than guessing. Scores are 0-2 per criterion to keep the comparison legible.
| Specimen | First-GET bytes | Heading skeleton | Citation topology | Table honesty | Canonical tag | Contrast (AA) |
|---|---|---|---|---|---|---|
design-agent.dev /blog/ |
139,651 | h1→h2→h3 tree | inline author links | real <table> elements |
present | unmeasurable from static GET |
tabby.tabbyml.com /docs/configuration/ |
352 (0 text chars) | none in shell | none in shell | none in shell | not decidable in shell | unmeasurable from static GET |
| anthropic.com/news/skills | 547,403 | h1 + h2s present | inline links present | no numeric table claim | present | unmeasurable from static GET |
Two things fall out of this table. First, byte size alone does not separate good from bad: 139,651 and 547,403 are both large, and both are structurally real, while 352 is tiny and structurally absent. Second, the contrast column is identical across all three rows, and that is not a coincidence — it is the rendering-tax showing up as a column of identical non-answers.
What the table does not say
The table does not say the 352-byte shell is a worse page for humans. It does not say the 547,403-byte page is bloated. Those are rendering and performance questions that a static GET cannot answer, and the rubric declines to pretend otherwise. What the table does say is narrower and more useful: an agent that spends one GET on the configuration route gets 0 text characters, and no amount of clever parsing recovers them.
The rendering-tax, stated plainly
A rendering engine buys an agent three things a byte stream cannot: computed styles, layout geometry, and interaction state. Every one of those is invisible in the response body, and every one of them is load-bearing for the two visual criteria in the rubric.
Contrast is the clearest case. The CSS may say color:#B8422E on a dark background and the declaration is right there in the bytes, but whether that pair clears 4.5:1 depends on the inherited stack and the actual background at that point in the layout. A rubric that scored contrast from raw declarations would be confident and wrong.
Focus appearance has the same shape. The WCAG 2.2 Understanding doc frames focus indicators in terms of area and contrast against adjacent colors — both computed properties. A static GET can confirm that a :focus-visible rule exists. It cannot confirm the indicator is visible.
Why the rubric still earns its keep
Seven of nine criteria are decidable, and those seven catch the failures that matter most in practice: missing canonical tags, heading trees that collapsed into divs, tables that were never tables, citation markers that resolve nowhere, and the 352-byte shell. Those are the defects an agent’s single GET is likeliest to hit.
The rubric’s value is not that it replaces a browser audit. It is that it tells you which browser audits are still necessary, and it does so before you spend one.
Applying typespy to your own page
Applying typespy to a single article is a short checklist. The procedure is deliberately boring, because boring procedures are the ones that get run.
Fetch the URL once, with no cookies and no JavaScript. Record the response byte count. Strip the markup and record the visible text character count. If the second number is zero or near zero while the first is not, stop — you have a JS shell and the remaining criteria are unscoreable.
Then walk the headings in document order and note the first structural jump. Count inline author links against bare markers. Search the body for numeric claims and check whether a <table> exists for each. Check for a canonical link. Check whether the CSS declares :focus-visible. Mark contrast as unmeasurable and move on.
The estimates you should not trust
Any word-to-markup ratio you compute this way is an approximation, not a measurement, and should be labeled as an estimate. Tokenizers differ, markup-strip regexes differ, and inline scripts inflate the denominator. Treat criterion 5 as a rough instrument that separates “mostly text” from “mostly machinery,” nothing finer.
The same caution applies to any inference about why a page renders the way it does. A 352-byte shell could be an intentional client-side app, a build misconfiguration, or a CDN serving the wrong artifact. The rubric measures the outcome, not the cause.
How This Guide Was Built
This review is based on server-rendered HTML only: no browser rendering, no JavaScript execution, no pointer or screen-reader testing.
The audit was performed on 2026-10-03. Pages fetched and their first-GET sizes, per URL: design-agent.dev /blog/ at 139,651 bytes; design-agent.dev post 2026-10-01-forced-colors-blind-spot at 26,300 bytes; 2026-09-29-dark-mode-blind-spot at 29,778 bytes; 2026-09-22-hover-blind-spot at 19,461 bytes; tabby.tabbyml.com /docs/configuration/ at 352 bytes with 0 text characters after markup strip; tabby.tabbyml.com /docs/models/ at 24,995 bytes; tabby.tabbyml.com /blog/ at 117,526 bytes (8,249 visible text characters after the strip); anthropic.com/news/skills at 547,403 bytes. Eight URLs were measured; three of the four design-agent.dev routes are posts, which is why the prose says “three specimens”: the three compared in the table below.
Method: single unauthenticated GET per URL, no cookies, JavaScript disabled. Markup strip removed tags and script contents before character counting. Linked CSS was not fetched, so criterion 6 was scored only where declarations are inlined; contrast criteria were marked unmeasurable throughout, because computed styles are not present in a response body.
The Bottom Line
Static GET predicts most of the rubric but cannot see computed styles at all — the rendering-tax is real, and that is why a browser audit is still the right tool for the two visual criteria.
FAQ
Does typespy replace a browser-based accessibility audit?
No. Seven of nine criteria are decidable from a response body, but body link-text contrast and focus-indicator visibility are computed properties that a byte stream does not contain. The rubric tells you which browser audits are still worth running, and it screens out pages whose static fetch already failed, before you spend rendering budget on them.
Why does a 352-byte page return HTTP 200 instead of an error?
Because the server successfully delivered a JavaScript shell. The status code describes the transfer, not the content. On 2026-10-03, tabby.tabbyml.com /docs/configuration/ returned 200 with 352 bytes and 0 text characters after markup strip, while /docs/models/ on the same host returned 24,995 bytes of real text — a route-level difference, not a host-level one.
Is a large byte count a bad sign for agent readability?
Not by itself. design-agent.dev /blog/ returned 139,651 bytes and anthropic.com/news/skills returned 547,403 bytes on 2026-10-03; both are structurally real pages. The failure signal is the inverse: a small payload with zero visible text. Byte size alone never decides a criterion in this rubric.
How should I score contrast if I cannot measure it?
Mark it unmeasurable and move on. Scoring contrast from raw CSS declarations produces confident wrong answers, because inherited color stacks and per-element backgrounds are not resolved in the response body. Verify any specific pair with WebAIM’s contrast checker in a rendering context, and record that measurement separately from the static-GET score.
