About the tool
Open speech-text foundation model for full-duplex spoken dialogue.
Moshi, from Kyutai, is a speech-to-speech foundation model and dialogue framework that supports full-duplex conversation using the Mimi streaming neural audio codec. Code is released under MIT and Apache-2.0 licenses with CC-BY 4.0 model weights, making it a reference open stack for realtime speech-native agents.
- Pricing
- Open source
- Category
- Voice
- Maker
- Not publicly supplied
- Provenance
- Indexed by Attest from public information
✓ Attested