A night-shift dispatcher at a regional hospital routes a patient’s automated intake to the emergency room. The hospital's new voice-to-text medical agent boasts a flawless pedigree, logging a 99 percent accuracy rating on its vendor scorecard. But the patient is deaf and speaks with a non-standard cadence. The agent transcribes the prepositions and the filler words perfectly, but it silently drops the word "not" from "did not take insulin." The aggregate score stays mathematically pristine. The patient goes into a coma.
Vendors are currently flooding government and enterprise procurement offices with systems that claim near-perfect transcription and translation. Under federal directives like OMB Memorandum M-24-10, agencies are rushing to certify that these tools meet technical benchmarks before deploying them to the public. But grading high-stakes software on a simple curve is a catastrophic misreading of how risk actually functions in the real world. Accuracy alone is profoundly insufficient for systems that manage rights, health, and employment.
The standard industry metric for a voice model is the Word Error Rate, a blunt calculation that treats every spoken word as equally valuable. A system that misses a critical dosage instruction or hallucinates a Miranda rights warning can still secure a passing grade if it successfully transcribes the rest of the paragraph. This mathematical camouflage hides the errors that actually matter. When a system is evaluated purely on its aggregate hit rate, the software fails perfectly—denying a deaf worker a promotion or a patient a diagnosis while giving the vendor an A-plus.
The Equal Employment Opportunity Commission has already warned that algorithmic screening tools can easily violate civil rights. If an automated HR interviewer silently screens out a candidate whose speech does not match its training data, the company has not purchased an efficiency upgrade. It has purchased a multi-million-dollar liability bomb. We are treating algorithmic access as a compliance nuisance rather than a core architectural requirement.
This is an economic failure before it is a legal one. The American industrial base is staring down a chronic labor shortage. We need Schumpeter, not a sermon. We need tools that creatively destroy old bottlenecks, replacing a scarce human intermediary with abundance. But an AI that cannot accept a text input instead of a spoken command does not just fail a few edge cases. About fifteen percent of American adults report trouble hearing, a number that climbs rapidly as the workforce ages and open-office layouts demand earbuds. When an enterprise software suite cannot handle a deaf worker’s voice, or assumes a signing user is just waving at the camera, it shrinks the labor market. It burns talent. A communication machine that cannot serve the people for whom communication has always been explicitly engineered is a defective product.
The market will not fix this by accident. We must fundamentally rewrite how we evaluate the models we buy. We need transparent model cards that require asymmetric error scoring. A passing grade cannot simply be a raw percentage. A viable model card must explicitly document the specific error type, the civil rights affected by that error, the detectability of the failure in real time, and the presence of a viable human fallback path.
If an automated system cannot confidently parse an input, it must know that it is failing. The software must have the architectural humility to trigger an accessible human fallback path. If a deaf user is struggling with a voice agent, the system should instantly pivot to a text interface or route the session to a human operator, rather than timing out and closing the ticket. This is what detectability means in practice. An error that the system catches and reroutes is an inconvenience. An error the system buries under a high confidence score is a lawsuit.
When federal and state buyers hold a procurement bake-off for new technology, they usually waste time evaluating the wrong features. They judge the cosmetics of a signing avatar or the tone of a voice assistant. A real evaluation must test sign-language understanding in a noisy room, the reliability of visual alerts when the audio fails, and the cascading cost of errors. Federal buyers already live under strict technical standards mapped by the U.S. Access Board. Yet, they routinely sign contracts for generative systems that bypass these protections entirely by pointing to a meaningless accuracy score.
Under the Americans with Disabilities Act effective communication guidance, covered entities have an obligation to provide access that actually works in context. A PDF full of vendor promises does not meet that standard. When the Department of Justice published its digital access rules under Title II, it pointed to strict technical standards like WCAG 2.1. But those rules were built for static websites, not for generative agents that improvise the interface on the fly. Old checklists cannot contain new physics.
This is where a growth-focused conservative and a civil-rights litigator should find immediate common ground. The conservative sees procurement waste. Buying a defective Section 508 system that requires constant human retrofitting costs ten dollars of labor for every one dollar of original design. The advocate sees a machine laundering discrimination. An automated system that silently drops a deaf applicant’s file into the void without triggering a fallback protocol is unacceptable. Both sides should agree that the government should only pay for performance.
But none of this matters if the vendors are allowed to grade their own homework. If the technical fields of a model card are defined entirely by the companies selling the software, they will define away the risk. Real accountability requires deaf technical governance over the evaluation metrics. The people who actually live at the edge of the system’s capabilities must define what a catastrophic error looks like. Without their direct authority over the scorecard, it is just another piece of corporate marketing.
When the Federal Communications Commission mandated hearing-aid compatibility for wireless handsets, it did not ask the industry to try its best. It set an outcome and forced the market to compete. We must do the same for artificial intelligence.
We must update the NIST AI Risk Management Framework and federal acquisition rules under FAR Subpart 39.2 to mandate comprehensive, asymmetrically scored model cards for all high-stakes public systems. These cards must weight error types, affected rights, and detectability, with technical governance led by the disabled users the system is most likely to fail. Let us stop paying for the privilege of being perfectly misunderstood.