Est.

Browser Automation for Accessibility Auditing

Automated tools catch only 20-40% of accessibility issues, requiring manual review for the rest.

Correspondent · · 10 min read · Updated
Cover illustration for “Browser Automation for Accessibility Auditing”
Browser Automations · August 17, 2026 · 10 min read · 2,322 words

Nearly every homepage on the internet fails at least one accessibility check. Testing across the top one million domains found detectable WCAG failures on 94.8% of them, averaging 51 errors per page. That's not a rounding error or an edge case: it's the baseline condition of the web as most people build it.

  • The scale problem is now also a legal one. More than 8,000 ADA lawsuits landed in 2025, and over 5,000 of those named websites or mobile apps specifically, a jump of 37% from the year before. Settlements average between $50,000 and $125,000, which makes accessibility debt one of the few line items in a codebase that can turn into a wire transfer.
  • The suits are not evenly spread. New York accounted for 637 in the first half of 2025 (31.6%), Florida followed with 487 (24.2%), California brought 380 (18.9%), and Illinois went from 28 cases to 237, a 745% jump that suggests plaintiffs' firms found a new venue worth working. New York remained the dominant venue for filings throughout 2025.
  • Filing a lawsuit used to require a lawyer who specialized in this niche. Self-represented filings rose 40% in 2025, tracking almost exactly with mainstream ChatGPT adoption. Drafting a complaint got easier, and the barrier to suing over a missing alt tag dropped with it.
  • Regulation is catching up on a separate track. The European Accessibility Act became enforceable on June 28, 2025, and it applies to any digital product sold to EU consumers, regardless of where the company is headquartered. A French court ordered Carrefour to make its website and app fully accessible within six months, with daily fines kicking in after the deadline, in what stands as the first major EAA enforcement ruling. Meanwhile, the DOJ's Interim Final Rule gave state and local governments an extra year on the Title II deadline, but the underlying ADA obligation, and the exposure to lawsuits, didn't move. A market roughly 1.3 billion people wide, about 1 in 6 people worldwide living with a significant disability, drives this: businesses that ignore accessibility are shutting out that population, a fact borne out by the adoption figures below. In the US, 18.7% of the population reports some form of disability, and 75% of those adults go online daily. An inaccessible checkout flow is a closed door to a customer base that's already online and ready to buy, on top of being a legal liability. It's a closed door to a customer base that's already online and ready to buy.

What browser automation catches and misses in a WCAG audit

Automated tools catch between 20% and 40% of accessibility issues, according to Crystal Preston-Watson speaking on Test Guild Automation Episode 515, with the manual share falling in the 60% to 80% range. That's the number to hold onto before buying any tool or writing any CI pipeline. Only about 30% of WCAG success criteria can even be evaluated by software in the first place, since the rest hinge on judgment calls about intent and meaning that no DOM parser can make.

A scanner can tell you an image has an alt attribute. It cannot tell you the attribute says "image1234.jpg" instead of describing what's actually in the picture. It can tell you headings exist in the markup. It cannot tell you whether jumping from an <h2> to an <h4> confuses a screen reader user trying to skim the page structure. The gap between "technically present" and "actually useful" is exactly where automation runs out of road.

  • Reliable catches: missing or empty alt text, unlabeled form fields, wrong ARIA roles and attributes, color contrast ratios that fall below WCAG thresholds, invalid semantic HTML, keyboard traps and focus failures that only appear once a user actually interacts with the page, and screen reader announcements tied to dynamic content updates.
  • What slips through: whether alt text is meaningful rather than merely present, whether heading order makes logical sense, whether an error message is clear enough for someone with a cognitive disability to act on, and whether a real screen reader user, on VoiceOver, NVDA, JAWS, or TalkBack, can complete an actual task from start to finish. The interaction layer catches keyboard traps and multi-step task failures that static scans miss entirely, because it tests real screen reader behavior on VoiceOver, NVDA, JAWS, or TalkBack as a task unfolds. A one-off page load might miss a keyboard trap that only appears after a modal opens. Tools that run inside a live browser session catch that; tools that just fetch and parse HTML do not.

The point in the development cycle where automated checks deliver the most value

Every defect gets cheaper to fix the closer it's caught to the change that introduced it. That's not specific to accessibility, it's true of software defects generally, but it has a particular bite here because accessibility regressions tend to hide in plain sight until a user with a disability actually hits them.

Unit tests give the best return on the effort. Checking component behavior, focus management APIs, and ARIA state at this layer is fast, deterministic, and catches problems before they ever reach a rendered page. Integration tests are where axe-core paired with a browser automation framework does its best work, since document-level rules and color contrast checks need an actual rendered DOM to evaluate against. End-to-end tests should carry a lighter load: reserve them for multi-page workflows and form submissions, because if the E2E suite already flakes on ordinary functional assertions, bolting accessibility checks onto it just adds noise nobody trusts.

  • CI/CD as a gate, not a suggestion. Running accessibility checks on every pull request catches problems exactly when they're cheapest to fix, and enforcing WCAG 2.1 AA as a gate condition means a build simply doesn't merge with a known violation in it.
  • The engineering lift is small. Adding axe-core into an existing Playwright suite takes under 20 lines of code, and the checks run fast enough in parallel CI pipelines that they don't meaningfully slow a build down.
  • Design systems compound the benefit. A button component gets reused across dozens of features, so a single regression caught in one automated check protects every place that component appears, instead of quietly propagating until someone notices in production. Rushed pre-launch passes leave no time to fix interpretation differences between browsers and assistive technology, so catching them early is better than letting the gap surface then.

The Playwright and axe-core stack: how the dominant open-source combination works

Playwright inspects the DOM as it actually renders in a real browser, capturing dynamic content and states triggered by user interaction. Dynamic content and states triggered by user interaction are testable in a way they simply aren't with older, request-based scanners. axe-core, built by Deque Systems and currently at version 4.11.1, takes that rendered DOM and checks it against a rule set covering WCAG 2.0 through 2.2, returning a structured list of violations with severity levels and direct references to the specific success criteria involved. Between tens of millions of weekly npm downloads, axe-core functions as something close to the industry's default accessibility testing engine.

The @axe-core/playwright package is the lowest-friction way to wire the two together: it inspects the rendered DOM, applies the WCAG 2.2 rule set, and surfaces violations tied to actionable style selectors a developer can go fix immediately. Configured well, with interaction-triggered checks included rather than just static page loads, the combination can catch up to 50% of WCAG issues automatically, compared to the 20-to-40% range that applies to automated testing generally. That's still short of full coverage, but it shrinks the feedback loop from "found in a manual audit three sprints later" to "flagged in the pull request that broke it."

  • Setup is genuinely light. Adding the check to an existing Playwright suite takes under 20 lines of code.
  • Microsoft released @playwright/mcp in early 2025. It's a Model Context Protocol server that gives AI agents browser control through structured accessibility snapshots, opening a path toward agent-driven testing that can navigate complex, multi-step interaction flows on its own.
  • Scope has to be chosen deliberately. Page-level checks running on every route answer a different question than component-level checks running inside unit tests, and conflating the two produces either redundant noise or blind spots.
  • Severity triage keeps the gate useful rather than annoying. axe-core returns violations at multiple severity levels, and most teams gate merges on critical and serious findings while logging moderate and minor issues to a backlog instead of blocking a release over them.
  • Browser version and operating system belong in the test configuration itself, since accessibility behavior is not identical across environments and a passing test on one combination doesn't guarantee the same result on another.

When to move beyond axe-core: enterprise and specialized tooling

The open-source stack has edges. Auditing a domain with a few dozen routes is one problem; auditing a domain with a few thousand pages is a different one, and writing individual Playwright scripts for each route stops being practical well before you get there. Crawl-based scanning fits that scale better. Login-walled content, the kind sitting behind an authenticated session, needs tooling built to handle that specifically. And regulated sectors, higher education, government, healthcare, tend to need an audit trail and expert review sitting alongside the raw scan output, in addition to a list of violations in a JSON file.

  • Siteimprove scans large web estates and ties accessibility findings into broader content quality signals, useful for organizations already tracking content health alongside compliance.
  • BrowserStack's Website Scanner runs scheduled scans across an entire domain, including pages that sit behind a login.
  • axe Monitor, from Deque, runs crawl-based monitoring paired with expert services for organizations that need audit trails alongside scan output.
  • Level Access pairs ongoing monitoring with expert review, serving roughly the same regulated-sector need.

Choosing among them comes down to a handful of weighted criteria that BrowserStack itself uses in its own scoring: detection coverage (missing labels, contrast failures, ARIA errors, keyboard traps), CI/CD integration and scalability, cross-browser and device support, and reporting quality alongside pricing and community support. A major standards body maintains a filterable list of evaluation tools organized by purpose, standard, scope, type, and platform, which is a more neutral starting point than any single vendor's comparison page.

Building the manual layer that automation cannot replace

The 60% to 80% of issues automation misses is the entire category of judgment calls, assistive technology behavior, and real-world usability that no rule engine can evaluate. It's the entire category of judgment calls, assistive technology behavior, and real-world usability that no rule engine can evaluate, because none of it reduces to a pass-or-fail check against the DOM.

Screen reader testing has to happen with the actual tools disabled users rely on: JAWS, NVDA, VoiceOver, and TalkBack are the tools disabled users rely on, and thorough testing requires exercising each of them. Switch device navigation and voice control flows fall into the same category. Emerging programmatic approaches are starting to close part of this gap, but expert-operated assistive technology sessions remain the standard that audits get measured against.

  • Alt text either communicates what's in an image or it doesn't, and only a human reviewing it in context can tell the difference.
  • Heading hierarchy either tells a logical story about the page's structure or it doesn't; software can confirm the tags exist, not that the order makes sense.
  • Error messages either give a user with a cognitive disability enough information to fix the problem, or they leave them stuck.
  • A modal, a multi-step form, a checkout flow, any of these can pass every automated check individually and still be impossible for a disabled user to complete end to end.

Paying disabled testers to run through real usability sessions is the only way to confirm that a stack of automated and manual checks actually adds up to something usable. Expert consensus across the accessibility field points the same direction: this work has to live inside design decisions, code review, and product requirements from the start, not get bolted on at the QA stage where fixing a bad pattern means redoing it everywhere it was already copied.

What an audit report must document to hold up legally and operationally

An audit report does two jobs at once. It's a technical diagnostic showing exactly where a site or app falls short of WCAG, and it's a paper trail that stands in for due diligence if a lawsuit asks why an accessibility problem existed. Skimp on either function and the report fails at the moment it's needed most.

Different readers pull different things from the same document. Legal teams read it for compliance risk. Developers need specific selectors and failure references they can act on directly. Executives need enough of a summary to justify resourcing a fix. Accessibility specialists need a baseline they can compare against the next audit to track whether things are actually improving.

State the overall WCAG 2.2 conformance level, whether A, AA, or AAA, in an executive summary that doesn't bury the headline finding.

  • Name the tools used, and be exact about it: axe DevTools 4.7, WAVE 3.2, or whatever manual methods were layered on top, rather than a vague reference to "automated testing."
  • Record the browser versions and operating systems the testing ran on, since results vary across environments and an unlabeled report can't be reproduced or defended later.
  • Date the testing precisely. Sites and apps change constantly, and a report without a testing date loses its evidential value fast once a plaintiff's attorney asks when the findings were actually current.
  • Assign a severity level to every violation and tie it back to the specific WCAG success criterion it fails, so nobody downstream has to guess how urgent a given fix is. Close with a prioritized remediation roadmap that turns each gap into a specific, measurable task, since a list of problems without a path to fixing them is a diagnostic.

Sources

  1. accessibility.huit.harvard.edu
  2. browserstack.com
  3. qamadness.com

More in Browser Automations