Google has introduced a significant accessibility breakthrough with its new sign-language-to-text (SL2T) model, developed by Google DeepMind. Announced as one of the standout features from the Made by Google event, the technology enables American Sign Language (ASL) users to input text by signing directly into their phone’s camera—eliminating the need to type in many everyday scenarios.
The feature launches first on the Pixel 11 series in Gboard (Google’s keyboard) and Live Transcribe. Users can now sign to perform web searches, draft messages or documents, interact with Gemini, or respond in real-time conversations. In Live Transcribe demos, it supports more natural two-way exchanges between ASL signers and people who do not sign, converting signs into English text on the fly.
Sundar Pichai highlighted the update, noting it was built in partnership with the Deaf community to help ASL signers communicate more seamlessly with their devices and non-signers.How SL2T WorksUnlike traditional voice dictation, sign language translation must interpret complex, simultaneous movements of the hands, body, torso, head, and facial expressions—elements that form full independent languages with their own grammar and vocabulary.The system prioritizes privacy: An on-device model (MediaPipe Holistic) tracks geometric landmarks (key points on the signer’s body) from the camera feed. Only these coordinates are sent to Google’s servers for translation by SL2T; the original video is discarded immediately. The model then converts the sequence of landmarks directly into text, bypassing intermediate “glosses” (simplified sign labels used in earlier research) that often lose important spatial and non-manual details.
This approach supports practical real-world use, including one-handed signing (common when holding a phone) and left-handed signers. It is optimized for low streaming latency and reduced hallucinations when the user is not actively signing.Training and PerformanceSL2T was trained on more than 100,000 hours of data spanning over 50 sign languages, with roughly a quarter of the dataset focused on ASL. Training across languages, dialects, and proficiency levels helped the model capture shared structures.On the FLEURS-ASL benchmark (which evaluates ASL-to-English translation of complex language), SL2T achieved a zero-shot score of 70 BLEURT—described by Google as significantly higher than previously reported results and the strongest performance to date for such models. Additional tests covered fingerspelling, one-handed variants, and assistant-style phrases.
The development involved collaboration with Deaf community members, organizations, and an AI Sign Language Advisory Committee. Google has published an impact report outlining both capabilities and current limitations, such as challenges with uncommon signs, rapid fingerspelling, or certain grammatical constructions when context is limited.Availability and Next StepsThe sign-to-text feature is available at no extra cost on Pixel 11 devices (with preorders underway and wider availability expected shortly after the announcement). Support for additional devices and more sign languages is planned, though no specific timeline has been given. Google frames this as the first consumer deployment of advanced sign language AI, with longer-term goals including expanded language coverage and potential sign language generation.By turning the phone’s camera into a natural input method for ASL, SL2T marks a practical step forward in making smartphones more inclusive. While limited to ASL-to-English at launch and currently exclusive to the latest Pixel hardware, it demonstrates how large-scale multimodal AI can address long-standing communication barriers for Deaf and hard-of-hearing users.