
Sesame CSM is an open-source speech generation model for developers and researchers building voice applications on their own hardware. It uses text and audio inputs to generate speech, with support for conversational context and different speakers. It's a model component for applications that need spoken output.
Context matters to its output. CSM can use earlier utterances and their audio as prompts, and the project recommends providing that context for better results. It can also generate a sentence without an audio prompt, in which case it uses a random speaker identity. This makes it relevant to research into spoken dialogue and applications that generate conversations between characters.
The released model is CSM-1B. It runs on a CUDA-compatible GPU and requires access to the CSM-1B and Llama-3.2-1B models on Hugging Face. The Python implementation includes Windows support, and CSM also integrates with Hugging Face Transformers. Its architecture combines a Llama backbone with a smaller audio decoder that produces Mimi audio codes.
The repository and CSM-1B weights use Apache 2.0. Setup also requires access to gated Llama-3.2-1B files with their separate terms. Local model execution is separate from the hosted Hugging Face Space available for testing audio generation. Sesame's interactive voice demo uses a fine-tuned variant of CSM, so the public checkpoint isn't the exact model behind that demo.
Claim this page with an email at sesame.com. Sesame CSM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Sesame CSM?Promote it
Something wrong or outdated on this page?
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
9.1KUpdated 1 year agoApache-2.0
macOS · Windows#Batch processing#Multilingual#ONNX
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.