Soniox has officially released Text-to-Speech v2, an update that introduces programmable expressive speech, integrated voice cloning, and low-latency streaming capabilities. According to a report by audioXpress, this launch occurs less than four months after the company introduced its initial text-to-speech model. This move represents a significant shift for Soniox as it expands from a speech recognition specialist into a complete infrastructure platform for real-time voice applications. The new version allows developers to insert audio tags directly into text to alter delivery and vocal behavior. These tags can request responses such as whispering, laughter, hesitation, excitement, tension, sadness, or reassurance within the same generated passage.

The model supports over 60 languages within a single framework, enabling natural language switching mid-sentence without requiring different voices. Furthermore, voice cloning is now fully integrated into the workflow, allowing custom voices to be created from reference recordings up to 20 seconds long while maintaining noise reduction and expressive controls. Improvements also address the pronunciation and rendering of names, specialist terminology, numbers, codes, addresses, and identifiers across languages. The system utilizes character-level timestamps and persistent WebSocket connections to facilitate real-time interaction and interruption handling. The previous model is deprecated and scheduled for removal on August 31, 2026. The new architecture offers regional processing in the United States, European Union, and Japan at approximately $0.70 per generated hour.

For more information, visit audioxpress.com.