Interestana
Home/News/ElevenLabs v4 Speech Model Adds Expression Control, 90 Languages
TechCrunch••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

ElevenLabs v4 Speech Model Adds Expression Control, 90 Languages

ElevenLabs has released its fourth generation speech synthesis model, branded as v4, introducing enhanced capabilities for voice cloning and broader language support. A significant advancement in v4 is its ability to clone a voice with a mere 10-second audio sample, a substantial reduction from previous requirements. This feature allows for more rapid and accessible voice replication, enabling creators and businesses to generate synthetic speech that closely mimics specific vocal characteristics. The model also significantly expands its linguistic reach, now supporting 90 languages, up from its previous iteration. This extensive language coverage aims to make high-quality synthetic speech accessible to a global audience, facilitating content creation and communication across diverse linguistic backgrounds.

Beyond voice cloning and language expansion, ElevenLabs v4 introduces a greater degree of control over expressive nuances in synthesized speech. Users can now fine-tune parameters related to emotion, intonation, and delivery style, allowing for more natural and contextually appropriate vocal performances. This granular control is crucial for applications requiring emotional depth, such as audiobooks, character voiceovers in games and animation, and personalized customer service interactions. The company stated that the model's architecture has been refined to better capture and reproduce subtle vocal inflections that convey emotion and personality.

ElevenLabs, a company specializing in AI-powered voice technology, has been a prominent player in the synthetic media landscape. Their models are designed to generate realistic human speech for a variety of applications, including content creation, accessibility tools, and professional voiceovers. The introduction of v4 represents a strategic move to maintain a competitive edge by offering more sophisticated features and broader utility. The company's focus on ethical AI development is also a key aspect of its public messaging, with efforts to prevent misuse of its voice cloning technology.

The expanded language support to 90 languages positions ElevenLabs to cater to a wider international market. This includes not only major global languages but also a broader range of regional dialects, aiming for greater inclusivity. The improved expression control further enhances the model's versatility, making it suitable for a more diverse set of creative and commercial projects. The company has not yet released specific benchmark data comparing v4's performance against its previous models or competitors in terms of latency, fidelity, or cost, but the emphasis on reduced cloning time and increased language options suggests a focus on accessibility and scalability. The development of v4 underscores the rapid advancements in the field of generative AI for audio, with companies continuously pushing the boundaries of realism and functionality.

Original source — read the full reporting at the publisher:

Read on TechCrunch

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next