A label is not a verdict

“AI label” has become an umbrella term for several very different assurances. That ambiguity is not harmless. It allows a product to borrow credibility from one kind of assurance while implying another. A provenance marker can imply product safety. A governance certificate can imply effectiveness. A declaration of WCAG 2.1 conformance can imply that a particular disabled person can complete a particular task. None of those implications follows automatically1.

The Inclusive AI Label should begin from this distinction. Its purpose should not be to create a new logo in a crowded trust market. Its purpose should be to make a limited, auditable claim about product access: for a stated population, a stated use case, and a stated version of a system, does the available evidence show that users can complete meaningful tasks without being excluded or forced into an inferior route?2

That question is harder than “was AI involved?” and harder than “does the organization have an ethics policy?” It cannot be answered by a self-attestation or a generic benchmark. It requires evidence from the conditions in which people will actually encounter the system: the language they use, the assistive technology they rely on, the stakes of a failure, the time available to recover, and the alternative available when the system does not work.1

The label should be conservative by design. It should not promise universal accessibility, legal compliance, or absence of discrimination. It should not represent an assessment of one product feature as a certification of an entire company. It should describe what was tested, what was not tested, and what users can expect. The strength of the label would come from restraint: the Institute makes only claims that its protocol can support.2

Three categories that should not be conflated

The first category is content provenance. C2PA Content Credentials are cryptographically bound records of a digital asset’s provenance: assertions can describe origin, modifications, tools used, and AI involvement1. Their value is real. They can make an asset’s declared history tamper-evident and help a recipient examine the basis for trust. But C2PA itself is explicit that a valid credential does not establish whether the content is true, accurate, or factual. Nor does it establish that the creation tool was usable by a Deaf person, a screen-reader user, or a person with a speech disability. Provenance is evidence about an asset’s history; accessibility is evidence about a person’s ability to use a product.1

The second category is organizational governance. A management system, policy, or risk framework can help an organization decide who owns a risk, document decisions, and establish internal controls. NIST describes the AI Risk Management Framework as voluntary, rights-preserving, non-sector-specific, and use-case agnostic2. Those are sensible features for a cross-sector governance framework. They also explain why such a framework cannot, by itself, establish that a caption stream remains usable in a noisy meeting, that an avatar produces intelligible signing, or that a hiring assessment measures a job skill rather than an applicant’s disability.2

The third category is product-access evidence. This is the category the Inclusive AI Label should occupy. It asks whether a concrete product version enables a defined user group to perform defined tasks, and whether the evidence was collected in a way that merits public confidence. A product-access assessment can draw on provenance and governance, but it cannot be replaced by either.1

This taxonomy should shape the Institute’s public language. It should say: a C2PA credential may establish tamper-evident provenance; a governance framework may establish that an organization has a risk-management process; an Inclusive AI Label, if earned, would establish only the scope stated in its evidence record. The label would complement these tools rather than compete with them.2

Captions are essential—but they do not settle access

Captions are not a trivial accommodation. WCAG 2.1’s prerecorded-media criterion requires captions for synchronized audio content3. Its live-media criterion requires captions for live synchronized audio content4. W3C explains that captions convey dialogue, speaker identification, and relevant non-speech sound—not merely a transcript. A product that produces no captions where captions are necessary has an obvious access failure.3

But captions are not the whole access question. W3C’s sign-language criterion explains why: written text is often a second language for sign-language users, while sign-language interpretation can carry intonation, emotion, and other information in ways captions do not5. The criterion therefore calls for sign-language interpretation for prerecorded synchronized media at Level AAA. This does not make captions optional. It means that caption presence cannot, by itself, demonstrate equivalent access for every user, language, or task.5

The distinction matters especially for AI systems because their access claims often depend on real-time interaction. A meeting agent might record, transcribe, summarize, assign action items, and answer follow-up questions. A live captioning service might display speech rapidly enough that users have no practical opportunity to correct its errors. A voice interface might require a user to speak in a way the model recognizes. A tool can conform to some web-content requirements and still fail at the interaction that determines whether the person can participate.6

The relevant research points away from a single metric. Kafle and Huenerfauth studied caption usability with 30 Deaf or hard-of-hearing participants and found that a captioning-focused measure aligned with participant judgments better than conventional word error rate6. Their result is not a universal pass/fail threshold. It is a methodological warning: a system can optimize a technical error statistic without optimizing what users need to understand or do.6

Kuhn and colleagues tested 11 common ASR services with higher-education lecture recordings7. They found wide variation across vendors and individual audio samples, and lower quality for streaming systems used in live events. Those findings should not be read as a claim that all automatic captions are unusable. They show that vendor-level claims and offline benchmarks do not establish performance in every environment. The Institute should therefore require product-specific evidence: real-time systems should be assessed in real-time conditions, with the audio, latency, turn-taking, correction mechanisms, and consequences that users actually face.7

A separate study of DHH users’ views of ASR captions in one-to-one meetings illustrates another problem with relying only on a model score8. Participants were concerned about accuracy; some also worried that displaying word-level confidence information would distract from communication. An access evaluation should therefore examine not just whether the system detects uncertainty, but whether its interface helps the person recover from an error without shifting the burden onto them.8

The practical conclusion is modest but important. Captions should be treated as a baseline feature where relevant. Automatic captions should be assessed as a user-facing system, not as a transcript generator. And a label should never turn a general captioning claim into a broad assertion of language access without evidence from the users and tasks covered by that assertion.3

Sign-language AI deserves a higher evidentiary bar

Sign-language recognition, generation, and translation are often described as if they were an ordinary extension of speech technology. They are not. An interdisciplinary review by Bragg and colleagues emphasizes that successful systems require expertise across computer vision, graphics, natural-language processing, human–computer interaction, linguistics, and Deaf culture9. The authors describe a field fragmented across separate portions of the processing pipeline.

That has direct implications for certification. A fluent-looking signing avatar is not evidence of adequate translation. A model that recognizes isolated signs is not necessarily usable in a real conversation. A demonstration on a curated video set does not show what happens with regional variation, non-manual signals, signing space, mixed signing and speech, camera placement, poor lighting, or a user who needs to correct the system quickly.9

Nor should the Institute assume that technical performance exhausts the relevant question. Tran, Ladner, and Bragg report that Deaf community perspectives and requirements for automatic sign-language translation have been poorly understood10. Their survey addresses desired scenarios, performance expectations, interface preferences, and perceived benefits and harms. The label should make those questions part of its protocol. It should ask users not only “was the translation accurate?” but “was this a setting in which you would choose to use it?”, “what harm follows from a plausible error?”, and “what alternative is available when the model fails?”10

A credible initial scope would be narrower than “sign-language AI is accessible.” It might read: For a specified language, task, device configuration, and user group, the product met the Institute’s published usability and error-recovery criteria in the evaluated scenario. The statement would remain meaningful because it is bounded. It could be revised as the protocol matures and evidence expands.10

Accessibility is not a compliance checkbox

Accessibility standards and civil-rights law matter. They establish obligations and offer useful technical baselines. They should not be diminished to make room for a voluntary label.11

For covered state and local government web content and mobile applications, DOJ identifies WCAG 2.1 Level AA as the technical standard under its Title II rule11. DOJ’s current guidance states that compliance dates are 26 April 2027 for entities with populations of 50,000 or more and 26 April 2028 for smaller entities and special districts. The same guidance explains that effective-communication and equal-participation duties remain relevant even when a piece of content is excepted from the technical standard or a person cannot use technically conforming content.11

These facts support a constructive position. WCAG 2.1 conformance should be a required baseline where applicable. It should not be marketed as proof that an AI product works for every disabled user in every context. The gap is visible even in the employment setting.

DOJ guidance says that employers choosing a hiring technology must ensure its use does not cause unlawful disability discrimination, including when the technology comes from another company12. It further advises employers to evaluate whether tools screen out qualified individuals with disabilities and to examine them before and regularly during use.12

The key legal lesson for the label is not that every accessibility failure has the same legal consequence. It is that an AI tool cannot be evaluated in the abstract. A video interview system, for example, can burden an applicant with a speech impairment, autism, or a vision impairment in different ways; DOJ explicitly notes that tools must be assessed for their impact on different disabilities. A broad “bias tested” claim is therefore weaker than a documented evaluation of the particular job-related task, disability-related barriers, accommodation pathway, and error consequences.12

The Inclusive AI Label should never claim to confer legal compliance or a safe harbor. It should instead present product-access evidence that buyers, users, procurement teams, and regulators can examine alongside their own obligations.11

What the Institute should test

A credible label begins with a published protocol. The protocol must be versioned, publicly readable, and specific enough that an outside party can understand how a decision was reached. It should define the product boundary, the supported functions, the users for whom a claim is made, and the evidence needed for a pass, a conditional pass, or no label.

The following requirements would make the program evidence-led rather than performative.

1. Defined intended use.

Each assessment should state the product version, interfaces, languages, deployment setting, and user tasks. “Accessible meeting assistant” is too vague. “English-language live meeting captions, speaker attribution, and transcript correction in a browser, evaluated with DHH participants using specified assistive technologies” is testable.

2. Cohort-relevant evaluation.

Testing should include people who share the access needs named in the claim. Recruitment, eligibility, compensation, and participant support should be disclosed. Disability inclusion cannot be delegated to a generic usability panel. Interviews with AI practitioners have found gaps in data about disabled users and friction between responsible-AI and accessibility practices13. A label should correct for that gap with decision-relevant evaluation, not a one-time consultation.

3. User-facing outcomes.

The protocol should measure task completion, time, comprehension, error recovery, workload, and participant-reported usability where they are relevant. Technical metrics can be reported, but they should not substitute for outcomes that matter to users. For high-stakes functions, the evidence packet should describe the consequence of a wrong output, not simply its frequency.

4. Failure analysis.

A company should report material failure modes and the conditions in which they arise: background noise, accent or dialect variation, visual occlusion, latency, screen-reader incompatibility, keyboard traps, model refusal, unclear correction flows, or a system’s inability to handle a language it claims to support. Results should be broken out by relevant cohorts when sample sizes permit, while avoiding false precision from tiny subgroups.

5. Fallback and accountability.

A user must have a meaningful route through the product when the AI feature fails. Depending on the setting, this may mean a human interpreter, a human review channel, an accessible non-AI pathway, a correction mechanism, or a way to obtain help without losing the opportunity at issue. In hiring, DOJ specifically identifies notice about the technology, information sufficient to seek an accommodation, and a clear accommodation procedure as relevant practices.12

6. Privacy and data practices.

An access feature may handle unusually sensitive data: voice characteristics, video of signing, disability-related information, or records of accommodation requests. The evidence record should state what data are collected, retained, shared, and used to train or improve the system. It should distinguish what is technically necessary from what is commercially useful.

7. Disabled-user governance.

Disabled people should have paid, decision-relevant roles in writing criteria, reviewing test plans, interpreting findings, and updating standards. Participatory design is not a decorative advisory board. The Institute should publish how it recruits reviewers, compensates them, handles disagreement, and prevents a small group from being treated as representative of every user.

8. Public documentation and recertification.

Every issued label should link to a plain-language evidence summary: criteria version, product version, scope, methods, known limitations, decision date, and re-evaluation date. Major product changes should trigger reassessment. A claim that cannot survive version change is not a durable label claim.

A decision rule people can inspect

The label needs a decision rule, not only a list of virtues. The protocol should separate eligibility, evidence sufficiency, and certification outcome. Eligibility asks whether the product and access claim are specific enough to assess. Evidence sufficiency asks whether the evidence packet contains the required participants, tasks, outcomes, and failure analysis. Certification asks whether the results meet the published threshold for the stated claim. Keeping these decisions separate prevents a polished submission from being mistaken for a successful product.

The public record should make the claim legible. It should identify the scope of the certification, not merely the product name: for example, “live English-language captions for planned meetings of up to a stated duration, evaluated with a stated participant group and device configuration.” It should also list exclusions. A product that has not been evaluated with a sign-language user should not carry language suggesting that it has. A product that works only with a particular browser or assistive-technology configuration should say so. Boundaries are not embarrassing disclosures; they are what make an assurance interpretable.

The evidence packet should distinguish between an absence of observed failure and evidence of acceptable performance. A small pilot that happens not to reveal a serious problem cannot establish that the problem is absent. Conversely, a documented failure need not mean that the entire product is unusable; it may define a condition in which a fallback is required or a claim must be narrowed. This distinction is particularly important when the population is heterogeneous and the number of participants is limited. The Institute should report uncertainty plainly rather than laundering it into a binary badge.

The label should also specify what triggers reassessment. A new model, a material change in latency, an interface redesign, a change in supported language, a new data-retention practice, or evidence of a previously unknown harm can all alter access performance. A certification without an expiry date or withdrawal process invites stale claims. A living standard should instead provide a route for correction: user complaints, post-market evidence, an investigation threshold, suspension criteria, and a public correction notice when the certified scope changes.

Independence is not a slogan

The Institute’s independence policy should be stated before certificates are issued. The policy should distinguish the cost of an assessment from the outcome of an assessment. A company may pay for evaluation work only if payment cannot purchase a result.

That requires more than a statement that the label is “not for sale.” At minimum, the Institute should publish fee rules; assessor qualifications; conflict-of-interest disclosures; separation between business development and certification decisions; an appeal process; a complaints process; recertification requirements; and a public registry of issued, suspended, expired, and withdrawn labels. It should state whether membership confers any access, and—if it does—why that access cannot influence an outcome.

The Institute should also publish negative decisions in aggregate and, with appropriate confidentiality protections, explain recurring reasons why products do not qualify. A label earns trust when it is clear that the answer can be no.

Start with a narrower public commitment

The most credible near-term commitment is not “we certify inclusive AI.” It is: we are building a public, evidence-led method for evaluating specified access claims in AI products, starting with the communities and use cases for which we can recruit qualified reviewers and conduct meaningful tests.

That commitment is ambitious enough. It treats accessibility as a property of an interaction between a person, a task, and a system—not as a feature that can be inferred from a press release. It respects the difference between provenance, governance, legal compliance, and observed product performance. It gives companies a disciplined route to improve, while giving users and buyers a way to inspect the basis of a claim.

The Institute should not issue a mark until its criteria, governance, and evidence packet are ready for public scrutiny. A standard that survives disabled-user review, adversarial questions, and failed assessments will be more valuable than a beautiful badge launched before the work is done.

References

  1. C2PA, C2PA and Content Credentials Explainer (specification 2.4).
  2. Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (Jan. 26, 2023).
  3. W3C Web Accessibility Initiative, Understanding SC 1.2.2 Captions (Prerecorded), WCAG 2.1.
  4. W3C Web Accessibility Initiative, Understanding SC 1.2.4 Captions (Live), WCAG 2.1.
  5. W3C Web Accessibility Initiative, Understanding SC 1.2.6 Sign Language (Prerecorded), WCAG 2.1.
  6. Sushant Kafle & Matt Huenerfauth, Evaluating the Usability of Automatically Generated Captions for People who are Deaf or Hard of Hearing, ASSETS ’17.
  7. Korbinian Kuhn et al., Measuring the Accuracy of Automatic Speech Recognition Solutions, ACM Transactions on Accessible Computing.
  8. Larwan Berke, Christopher Caulfield & Matt Huenerfauth, Deaf and Hard-of-Hearing Perspectives on Imperfect Automatic Speech Recognition for Captioning One-on-One Meetings, ASSETS ’17.
  9. Danielle Bragg et al., Sign Language Recognition, Generation, and Translation: An Interdisciplinary Perspective, ASSETS ’19.
  10. Nina Tran, Richard E. Ladner & Danielle Bragg, U.S. Deaf Community Perspectives on Automatic Sign Language Translation, ASSETS ’23.
  11. U.S. Department of Justice, Fact Sheet: New Rule on the Accessibility of Web Content and Mobile Apps Provided by State and Local Governments (2024).
  12. U.S. Department of Justice, Algorithms, Artificial Intelligence, and Disability Discrimination in Hiring (May 12, 2022).
  13. Sanika Moharana et al., “Accessibility People, You Go Work on That Thing of Yours over There”: Addressing Disability Inclusion in AI Product Organizations, AIES 2025.
arrow
Prev Post
Next Post
arrow
We’d love to hear from you

Get in touch to Explore Research collaborations, policy projects, or Accessibility Programs.