Favicon of Sesame CSM

Sesame CSM

An open-source speech generation model that uses text and audio context, runs on a CUDA-compatible GPU, and integrates with Hugging Face Transformers.

Screenshot of Sesame CSM website

Sesame CSM is an open-source speech generation model for developers and researchers building voice applications on their own hardware. It uses text and audio inputs to generate speech, with support for conversational context and different speakers. It's a model component for applications that need spoken output.

Context matters to its output. CSM can use earlier utterances and their audio as prompts, and the project recommends providing that context for better results. It can also generate a sentence without an audio prompt, in which case it uses a random speaker identity. This makes it relevant to research into spoken dialogue and applications that generate conversations between characters.

The released model is CSM-1B. It runs on a CUDA-compatible GPU and requires access to the CSM-1B and Llama-3.2-1B models on Hugging Face. The Python implementation includes Windows support, and CSM also integrates with Hugging Face Transformers. Its architecture combines a Llama backbone with a smaller audio decoder that produces Mimi audio codes.

The repository and CSM-1B weights use Apache 2.0. Setup also requires access to gated Llama-3.2-1B files with their separate terms. Local model execution is separate from the hosted Hugging Face Space available for testing audio generation. Sesame's interactive voice demo uses a fine-tuned variant of CSM, so the public checkpoint isn't the exact model behind that demo.

Similar to Sesame CSM