3 liens privés
OmniVoice is a state-of-the-art massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it generates high-quality speech with superior inference speed, supporting voice cloning and voice design.
'équipe Next-gen Kaldi de Xiaomi a sorti OmniVoice, un modèle de synthèse vocale qui parle 646 langues et qui est capable de "recopier" une voix à partir d'un extrait de huit secondes minimum. Et tout ça en local et gratuitement !
Open source voice cloning studio with support for multiple TTS engines. Clone any voice, generate natural speech, and compose multi-voice projects — all running locally.
A fast and local neural text-to-speech engine that embeds espeak-ng for phonemization.
Text-to-speech
Kyutai text-to-speech started as an internal tool we used during the development of Moshi. As part of our commitment to open science, we've since open-sourced two text-to-speech models:
Kyutai Pocket TTS, a tiny model with voice cloning, fast enough to run on CPU.
Kyutai TTS 1.6B, a streaming model used in Unmute, great for servers.