Hero image for Accessibility Compliance for AI Design Agents: EAA, EN 301 549, WCAG 2.2 AA, and What Your Tools Can't See

Accessibility Compliance for AI Design Agents: EAA, EN 301 549, WCAG 2.2 AA, and What Your Tools Can't See

Introduction

An AI design agent cannot improve what it cannot measure. Ask a model to “design an accessible dashboard” and you get a plausible-looking interface; ask it to satisfy a specific, checkable rule and you get something a pipeline can verify. The gap between those two requests is the subject of this post. The European Accessibility Act (EAA) and the harmonised standard EN 301 549 give the web a legal floor, and WCAG 2.2 AA gives that floor a technical vocabulary. But only part of that vocabulary is computable. If we want agents to design better, we need to be precise about which criteria an agent can enforce upstream, which it can only probe, and which it must hand to a human. The rest of this piece stays inside that frame: the legal floor, the computable subset, what tooling actually sees, and why an agent should behave like a constrained design system rather than an overlay.

The EAA is Directive (EU) 2019/882 of 17 April 2019, published in the Official Journal (EUR-Lex). Its application date is reported as 28 June 2025 by compliance-vendor guides such as AllAccessible and accessiBe, both vendor-authored. We did not check that date against the Official Journal text, so treat it as reported rather than confirmed here.

The harmonised standard referenced for the web is EN 301 549. The version cited here is V3.2.1 (2021-03) (ETSI EN 301 549). Its Annex E.5 notes that Directive 2016/2102 covers public-sector bodies as a minimum, and that clauses 9, 10 and 11 address websites, documents and software. Separately, W3C states that the 2026 version of EN 301 549 uses WCAG 2.2 (W3C WCAG overview). So the practical target converges on WCAG 2.2 AA, even though the standard’s own revision cycle lags the W3C Recommendation.

WCAG 2.2 was published on 5 October 2023, with an update published on 12 December 2024 (W3C WCAG 2.2; W3C WCAG overview). Per W3C, it adds nine success criteria over 2.1 and leaves existing criteria unchanged, except 4.1.1 Parsing, which is made obsolete. W3C also notes that WCAG 2.2 is ISO/IEC 40500:2025, based on the October 2023 text, with the December 2024 text expected as ISO/IEC 40500:2026.

What WCAG 2.2 AA asks of a design agent: the computable subset

“Accessible” is not one property. It is a list of success criteria, each with a testable statement. The six below are the subset this post works with. We did not re-verify each number and level line by line for this post, so check the W3C WCAG 2.2 text before quoting one.

  • 1.4.3 Contrast (Minimum), AA — text contrast against its background.
  • 1.4.11 Non-text Contrast, AA — contrast for user-interface components and graphical objects.
  • 2.4.7 Focus Visible, AA — a visible keyboard focus indicator.
  • 2.4.11 Focus Not Obscured (Minimum), AA — new in 2.2; the focused component is not entirely hidden by author-created content.
  • 2.5.8 Target Size (Minimum), AA — new in 2.2; pointer targets meet a minimum size, with stated exceptions.
  • 4.1.2 Name, Role, Value, A — user-interface components expose name, role and value to assistive technology.

For an agent, the useful cut is: which of these can be computed from tokens and component contracts before anything ships? Contrast ratios are arithmetic over color pairs. Target size is geometry. Non-text contrast is arithmetic over component boundaries. Focus visibility and “not obscured” are partly computable: you can check whether a focus ring exists and whether a sticky layer overlaps the focused element’s box, but judging whether a focus indicator is visible in context, or whether content obscures it, depends on runtime state. Name, role and value is the least computable: it is about what assistive technology announces, which is a semantic contract, not a pixel measurement.

The automation spectrum: what tools can and cannot see

Tooling is usually described in three tiers. Deque, for example, distinguishes automated rules-engine checks (axe-core, open source since 2015), guided semi-automated tests, and manual expert review with assistive technology (Deque axe). The table maps those tiers onto our six criteria. The mapping is analytical, not a measurement.

Automated Semi-automated Manual
What it evaluates Computable properties of rendered output: contrast ratios, geometry, presence of a focus indicator, target sizes. Properties needing a human or agent prompt to resolve: whether the focus indicator is visible in context, whether an accessible name is meaningful. Behavior with assistive technology and real content: whether name, role and value are announced and operable.
Example SCs from our set 1.4.3, 1.4.11, 2.5.8. 2.4.7, 2.4.11, 4.1.2 (name meaning). 4.1.2 end-to-end; 2.4.7 in dynamic layouts.
What it misses Context, semantics, meaningful names, dynamic obscuring by sticky headers or cookie banners, and content that changes after render. Screen-reader announcements, keyboard traps in real flows, and edge cases across assistive-technology and browser combinations. Scale, speed and regression coverage; it cannot run on every commit.
Role for an AI design agent Cheap upstream gate: lint tokens and components before render; fail fast in CI. Ask targeted questions, collect the missing signal, and attach evidence to the design ticket. Route to a human with a precise, reproducible brief; never substitute for expert review.

Two concrete references anchor the ends of this spectrum. Pa11y is an open-source CLI for automated page checks, per its own site (Pa11y). Stark, a design-tool plugin for contrast and focus checks, is mentioned here without any coverage claim; we did not evaluate its coverage.

What the coverage numbers actually say (and what they don’t)

Three numbers circulate in accessibility tooling conversations, and they are not interchangeable.

First, secondary roundups attribute to a Deque study, reportedly more than 2,000 audits across 13,000 pages, a figure of roughly 57% automated issue coverage by volume (Alphonso Labs; a11yflow). We did not fetch the primary study, so treat 57% as reported, not verified.

Second, a fixture-based benchmark reports that tools detected between 40% and 63% of 30 seeded defects, with 5 of 30 (17%) found by none of the four tools tested (accessibility.build benchmark). This is a single fixture page, not a web-wide sample; it tells you how tools behave against deliberately planted problems, not what share of real-page issues they find.

Third, Deque’s “up to 80%” phrasing appears in its own materials (Deque axe). That is vendor marketing: a claim, not a measurement. It should not be quoted as a coverage statistic.

Read together, the honest takeaway is narrow: automated checks catch a meaningful but bounded share of issues, a residual share is not machine-detectable at all, and the marketing ceiling is not evidence. For an agent, a green lint result is a necessary condition, never a sufficient one.

Why overlays fail — and why an agent shouldn’t become one

The architectural difference matters more than any feature list. An accessibility overlay patches the rendered page after the fact: it injects scripts, adjusts attributes, and attempts to remediate a site the user already loaded. An AI design agent operates upstream: it constrains tokens, component APIs and layout rules before the interface exists. The first is post-render patching; the second is constraint enforcement.

The enforcement record is thin, and the sources matter. A competing audit vendor’s write-up reports an FTC action against an accessibility overlay vendor over unsubstantiated “instant compliance” claims, reportedly with a USD 1M order (ADA Fail). We did not check that amount or the order against ftc.gov, and we are not naming the vendor. Separately, a paid article sponsored by accessiBe, an overlay vendor, reports that two French disability groups sent formal notices to four French grocery retailers after the EAA took effect (Europe Says). That article is vendor-sponsored, so we cite it only for the fact of the notices. Neither item tests overlays. They show enforcement is starting. They do not show that overlays fail in court.

The argument here rests on architecture, not on that enforcement record. An agent that mimics overlay logic would inherit the same failure mode: a post-hoc patch that cannot see the design intent that produced the defect. An agent that enforces constraints changes what gets generated in the first place.

Implications for design agents and design systems

If the computable subset is the agent’s territory, the design system is where that territory gets defined. Contrast becomes a token-level assertion: every foreground/background pair in the palette carries a computed ratio, and 1.4.3 and 1.4.11 are checked at the token, not the screenshot. Focus becomes a contract: a single focus-ring token, a defined offset, and a rule that author-created layers (sticky headers, toasts, consent banners) must not fully cover the focused element, which is the practical form of 2.4.11. Target size becomes a minimum in the component API, which is how 2.5.8 gets enforced before a designer drags a 24-pixel icon button onto a canvas.

Name, role and value is the boundary. An agent can require that every interactive component declare an accessible name and an explicit role, and can flag a button whose label is an unlabelled icon. It cannot confirm that a screen reader announces that name usefully. So the agent’s job is to make the semantic contract explicit and then stop, surfacing uncertainty rather than manufacturing confidence. In practice, design tickets carry computable checks as pass/fail, semi-computable checks as prompts with evidence, and manual checks as a named human step with a reproducible brief.

FAQ

Does meeting WCAG 2.2 AA make us EAA compliant? Not by itself. The EAA is the legal instrument (EUR-Lex), EN 301 549 is the harmonised standard (ETSI), and W3C states the 2026 version of EN 301 549 uses WCAG 2.2 (W3C). WCAG is the technical core, not the whole obligation. The EAA application date of 28 June 2025 is reported by vendor compliance guides (AllAccessible) and was not verified here.

Can an AI design agent guarantee accessibility? No. Automated rules catch a subset. Secondary roundups attribute roughly 57% coverage by volume to a Deque study we did not fetch (Alphonso Labs; a11yflow), and a single-fixture benchmark found 40%–63% detection, with 17% found by none of the tools (accessibility.build). A guarantee would require covering the undetectable residual, which tooling cannot do.

Is an accessibility overlay a shortcut to compliance? Treat that claim skeptically. A competing vendor’s report describes an FTC action against an overlay vendor over “instant compliance” claims, reportedly involving USD 1M, which we did not check against ftc.gov (ADA Fail). Post-render patching also cannot fix defects that originate in design decisions, which is why an agent should enforce constraints upstream instead.

Which WCAG 2.2 AA criteria are computable? From our set, 1.4.3, 1.4.11 and 2.5.8 are largely computable from tokens and geometry. 2.4.11 is partly computable: overlap geometry can be checked, but judging what counts as obscured depends on runtime state. 2.4.7 and 4.1.2 need context or assistive-technology confirmation. Check each number and level against the W3C WCAG 2.2 text before quoting it.

What actually changed in WCAG 2.2? It adds nine success criteria over 2.1 and makes 4.1.1 Parsing obsolete, with existing criteria otherwise unchanged (W3C). Among our six, 2.4.11 and 2.5.8 are new in 2.2.

Do we still need manual testing if we automate? Yes. Manual review with assistive technology is the only tier that confirms behavior such as name, role and value being announced and operable, and it is the tier automated and semi-automated checks miss by design (Deque). Automation reduces volume; it does not replace judgment.

The Bottom Line

An AI design agent gets better at accessibility when “accessible” becomes a set of computable constraints. The legal floor is set by the EAA and EN 301 549, the technical vocabulary is WCAG 2.2 AA, and the agent’s realistic scope is the subset it can compute: contrast, non-text contrast, target size, focus presence, overlap geometry, and explicit semantic contracts. Everything else is semi-automated probing or human review. Coverage numbers, whether reported, fixture-based, or vendor marketing, all point the same way: a passing automated check is necessary, not sufficient. Overlays patch what already shipped; agents can shape what gets made.

How This Guide Was Built

This guide relies on official documentation: the EAA text on EUR-Lex, the harmonised standard from ETSI, and WCAG 2.2 from W3C. It also uses secondary reports for dates, coverage figures, benchmark results and enforcement items, including vendor-authored and vendor-sponsored sources, which are labelled where they appear. The authors did not run any tool hands-on and ran no audits or benchmarks. Every quantitative claim is attributed to its source and labelled as verified, secondary, fixture-based, or vendor marketing.