Sign Language AI Reaches Your Phone: What Google's SL2T Really Does
💡 On August 12, 2026, Google DeepMind shipped SL2T, the first sign language AI model built into a consumer phone. Deaf and hard of hearing users on Pixel 11 can now sign directly into Gboard or Live Transcribe instead of typing. It is a genuine first, though it covers only American Sign Language today and still makes errors on rare signs and rapid fingerspelling.
- SL2T launched August 12, 2026, on Google Pixel 11, converting American Sign Language to English text in real time.
- Users can sign directly into Gboard (messages, searches, Gemini queries) and Live Transcribe without typing.
- The model tracks 130 body landmarks via MediaPipe on-device and discards raw video immediately, protecting privacy.
- On the FLEURS-ASL benchmark it scored 70 BLEURT, the highest reported for this task, but this is a vendor-reported figure, not independently verified.
- Honest caveat: ASL only at launch, errors persist in low light and with rare signs, and SL2T must not replace certified interpreters in medical, legal, or formal settings.

What Google's SL2T Model Actually Does
For decades, sign language AI existed mainly in research papers and controlled lab demos. On August 12, 2026, Google DeepMind changed that. It shipped sign language AI inside two everyday Android products: Gboard, the keyboard, and Live Transcribe, the accessibility app.
The technology is called SL2T, short for sign-language-to-text. On Pixel 11 phones, a Deaf user can hold their phone up, sign in American Sign Language, and watch their message appear as English text in real time. They can then edit and send it like any other text.
This is the first time this capability exists as a standard phone feature, not a research prototype or specialist device.
How Does It Work, Exactly?
SL2T does not record video and send it to a server. Instead, an on-device layer called MediaPipe Holistic maps 130 key points on the face, hands, and body: positions, angles, and movement, tracked frame by frame. Only those geometric coordinates travel to Google's servers. The raw camera footage is discarded immediately.
The model converts that coordinate stream directly into English text, without passing through an intermediate gloss layer. A gloss is a fixed label for a single sign. Earlier systems relied on large gloss dictionaries, which meant they could not handle signs outside the dictionary and could not represent ASL's spatial grammar or facial markers. Bypassing glosses lets SL2T learn directly from data.
It was trained on more than 100,000 hours of signing video across 50 sign languages, roughly 25% of which was ASL. Google says the multilingual training improved ASL accuracy compared with training on ASL alone.
What Does This Change for Deaf and Hard of Hearing People?
Voice dictation has been a default phone feature for hearing users for more than a decade. Deaf and hard of hearing users have had no equivalent: they have always had to type. SL2T closes part of that gap, at least for ASL signers with a Pixel 11.
The practical use cases are the ones hearing users take for granted: drafting a text, searching the web, querying Gemini, taking notes. In Live Transcribe, a Deaf user can sign a response instead of typing during a back-and-forth conversation. Google's testers reported that signing felt faster and more natural than typing in English.
There are more than 70 million Deaf and hard of hearing people worldwide, communicating in some 200 distinct sign languages. This release covers one language on one phone model. But it demonstrates that consumer-grade accuracy is now reachable, and that matters for what comes next.
Why Language Matters More Than Technology Here
This release is also a statement about how to treat sign languages. Earlier systems often translated ASL into signed English, a word-for-word mapping that ignores ASL's own grammar. SL2T was built from the start to treat ASL as an independent language, with its own spatial constructions, facial markers, and non-linear grammar.
The difference is not just technical. A system that imposes English word order onto ASL implicitly frames ASL as a lesser or derived form of English. Sign language linguists have documented this tension for decades, and treating ASL as a primary language is a governance decision as much as an engineering one.
For anyone working with language, the research on multilingual cognition and brain aging is a useful reminder that language is not just a communication medium but a cognitive architecture. That framing matters when building any system that claims to bridge linguistic worlds.
What Are the Real Limits You Should Know?
Several things SL2T cannot do are as important as what it can.
- Only ASL to English at launch. BSL, LSF, Auslan, and the other major sign languages are not supported.
- Performance drops in low light, at oblique camera angles, and when the signer is not the largest person in frame.
- Errors persist with rare signs, rapid fingerspelling, and ASL's passive constructions and classifier handshapes.
- 60-second clip limit, no memory between clips, and not trained on signers under 18.
- Occasionally generates text when no one is signing (hallucination), though Google says this has been reduced from pre-release builds.
Critically: the advisory committee that co-developed this technology with Google, including the National Association of the Deaf and the World Federation of the Deaf, explicitly scoped SL2T to informal, low-stakes use: messaging, searching, and casual conversation. Medical appointments, legal proceedings, police interactions, and job interviews are not within scope. Certified human interpreters remain legally and ethically required for those.
The benchmark figures cited are vendor-reported. No independent academic evaluation of the deployed model has been published as of August 2026.
What to Watch Next
Google's roadmap includes punctuation and editing commands, multi-signer support, better facial expression tracking, and expanded ASL dialect data including Black ASL, which current models underperform on. The most consequential open question is whether the architecture can transfer to other sign languages without requiring 100,000 hours of new data for each one.
If it can, this becomes a template for closing the voice-dictation gap in 200 sign languages at once. If it cannot, SL2T remains a significant but narrow first step.
FAQ
Does SL2T work for all sign languages, or only American Sign Language?
At launch (August 2026), SL2T supports only American Sign Language to English on Pixel 11. The model was trained on data from 50+ sign languages and Google has stated plans to expand, but no timeline has been given for languages like BSL, LSF, or Auslan.
Can I use SL2T on any Android phone?
No. At launch, SL2T is available only on the Google Pixel 11. Google has indicated additional devices will follow but has not specified which models or when.
Is SL2T accurate enough to replace a human sign language interpreter?
No, and Google's advisory committee explicitly prohibits this use. SL2T is designed for informal settings: messages, searches, casual conversations. It must not be used in medical consultations, legal proceedings, police interactions, job interviews, or any formal setting where certified interpreters are legally or ethically required.
How does SL2T protect my privacy when it uses my camera?
The raw camera video is discarded on-device immediately. MediaPipe converts it into 130 geometric body landmarks before anything leaves the phone. Only those coordinates are processed by Google's servers. The actual video of the signer never travels to Google.
What makes SL2T different from earlier sign language AI systems?
Earlier systems relied on gloss labels, fixed text labels for individual signs, as a middle step. This limited vocabulary and lost ASL's spatial grammar. SL2T generates English directly from movement coordinates, so translation quality scales with training data rather than with hand-built dictionaries. It also treats ASL as an independent language, not signed English.
Source: Google DeepMind Blog, Putting Sign Language AI into Users' Hands (2026)
About the author
Dao Huy (Lucas) is a professional translator working across English, Vietnamese, Chinese, and French, with seven years of experience in technical, legal, and software localization. He follows developments at the intersection of language and technology because, as SL2T shows, the two are rarely separable: the choices a model makes about grammar, word order, and linguistic authority are not just engineering decisions but language policy.
If your business needs English-to-Vietnamese translation, technical documentation, patent and IP localization, or software interface localization, Lucas offers a free quote at daohuy.com.
Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →
