Verified alternatives ยท Voice

Verified ElevenLabs alternatives

Compare 38 same-category tools using only the descriptions stored in the Attest directory.

Comparison rule

No invented feature matrix.

Each option below is in the same live directory category as ElevenLabs. Its bullets reproduce that tool's stored tagline and description; they do not claim feature parity, superiority or suitability for your workflow.

Same-category directory records

Alternatives to inspect.

Open any name to see its verification record and original outbound links.

Open source

Moshi

  • Open speech-text foundation model for full-duplex spoken dialogue.
  • Moshi, from Kyutai, is a speech-to-speech foundation model and dialogue framework that supports full-duplex conversation using the Mimi streaming neural audio codec. Code is released under MIT and Apache-2.0 licenses with CC-BY 4.0 model weights, making it a reference open stack for realtime speech-native agents.
Freemium

Fish Audio

  • Text-to-speech and voice cloning platform built on the open Fish Speech models.
  • Fish Audio offers hosted TTS and voice cloning with a free tier and API pricing, built on the Fish Speech family of multilingual open models that support 80+ languages and natural-language emotion control. The model code is available on GitHub under a source-available research license.
Open source

Coqui TTS

  • Open-source deep learning toolkit for text-to-speech.
  • Coqui TTS is an MPL-2.0-licensed library for training and running TTS models, including the XTTS voice cloning models, with support for a large number of languages. Although the company behind it shut down, the repository remains publicly available and widely used, with active community forks.
Open source

Kokoro

  • Open-weight 82M-parameter text-to-speech model with fast inference.
  • Kokoro-82M is an Apache-2.0 open-weight TTS model that delivers quality competitive with much larger models at a fraction of the compute cost, making it popular for local and self-hosted speech synthesis. The repository provides an inference library and pretrained voices.
Paid

Inworld

  • Realtime TTS and voice APIs for consumer-facing AI applications.
  • Inworld provides a realtime TTS API with low-latency streaming, voice cloning from short samples, and lipsync timestamps, plus a combined STT-LLM-TTS realtime WebSocket API. It targets voice agents in games, companion apps, and customer support, with pay-as-you-go pricing.
Freemium

Resemble AI

  • Voice cloning, speech synthesis, and deepfake detection for enterprises.
  • Resemble AI offers voice cloning and text-to-speech APIs alongside detection tools for identifying AI-generated audio in real time. It serves enterprise use cases with usage-based pricing, free trials, and on-premise deployment options.
Paid

LMNT

  • Low-latency streaming text-to-speech and voice cloning API.
  • LMNT provides studio-quality voice synthesis and cloning with roughly 150-200ms streaming latency across 31 languages, aimed at realtime agents and interactive applications. Pricing is pay-as-you-go with volume scaling and enterprise plans.
Freemium

Rime

  • Text-to-speech models built for high-volume, realtime business conversations.
  • Rime develops TTS voices designed for conversational realism in applications like customer service calls, with low-latency streaming and on-prem deployment options for enterprises. It offers trial access and tiered API pricing.
Freemium

Cartesia

  • Realtime speech models purpose-built for voice agents.
  • Cartesia offers Sonic (low-latency streaming text-to-speech), Ink (speech-to-text), and Line (a voice agent platform), built on state space model research for fast realtime inference. It positions its API for latency-critical voice agent workloads, with a free tier and usage-based plans.
Paid

NVIDIA Riva

  • GPU-accelerated speech AI microservices for self-hosted deployment.
  • Riva is NVIDIA's speech AI platform providing ASR, TTS, and translation models (now under the Nemotron speech umbrella) that deploy as GPU-accelerated microservices on-premises or in the cloud. It targets enterprises that need low-latency, self-managed speech infrastructure, licensed through NVIDIA AI Enterprise with free evaluation access.
Open source

whisper.cpp

  • Dependency-free C/C++ port of Whisper for local and edge inference.
  • whisper.cpp runs OpenAI's Whisper speech recognition in plain C/C++ with support for Apple Silicon, x86, and multiple GPU backends. Its small footprint makes it a common choice for on-device and embedded transcription. MIT-licensed.
Open source

faster-whisper

  • Optimized Whisper inference using CTranslate2.
  • faster-whisper reimplements Whisper transcription on the CTranslate2 inference engine, delivering up to 4x faster performance with lower memory use than the reference implementation. It is MIT-licensed and widely used for self-hosted transcription in realtime pipelines.
Open source

OpenAI Whisper

  • Open-source general-purpose speech recognition model.
  • Whisper is an MIT-licensed model for multilingual transcription, translation, and language identification, trained on a large corpus of diverse audio. It is a de facto baseline for open speech recognition and underpins many derivative runtimes used in voice products.
Paid

Soniox

  • Realtime multilingual speech recognition and translation APIs.
  • Soniox provides speech-to-text, translation, and speech understanding APIs covering 60+ languages, with realtime streaming suited to conversational applications. Pricing is pay-as-you-go.
Freemium

Speechmatics

  • Multilingual speech-to-text APIs for realtime and batch transcription.
  • Speechmatics offers speech recognition with strong multilingual and accent coverage, realtime streaming, and speaker diarization, and integrates with voice agent frameworks such as Vapi and Pipecat. Pricing is usage-based with a free credit tier and enterprise plans.
Freemium

Gladia

  • Realtime and batch speech-to-text API for voice platforms.
  • Gladia provides transcription and audio intelligence APIs with realtime streaming, multilingual support, and features aimed at call platforms and voice agent builders. New accounts receive free transcription credits, with consumption-based pricing beyond that.
Freemium

AssemblyAI

  • Speech recognition and audio intelligence APIs for voice applications.
  • AssemblyAI offers streaming and asynchronous speech-to-text along with audio intelligence features such as speaker diarization, sentiment, and summarization. Its Universal-Streaming models target voice agent use cases, and pricing is usage-based with a free tier.
Freemium

Deepgram

  • Speech-to-text, text-to-speech, and voice agent APIs in one platform.
  • Deepgram provides realtime and batch speech recognition, the Aura text-to-speech family, and a Voice Agent API that combines them for conversational applications. It is oriented toward developers building latency-sensitive voice products, with free credits for new accounts and usage-based pricing.
Open source

Silero VAD

  • Pre-trained open-source voice activity detector for speech pipelines.
  • Silero VAD is an MIT-licensed voice activity detection model that identifies speech segments across many languages and noise conditions with a small footprint. It is a common building block for endpointing and turn-taking in realtime voice agent stacks.
Freemium

Krisp

  • Noise cancellation and speech enhancement SDKs for calls and voice agents.
  • Krisp offers on-device noise cancellation, voice isolation, accent conversion, and turn-taking models, available as consumer apps and as SDKs for embedding in voice AI pipelines. Its developer SDKs are used to improve audio quality and endpointing accuracy in production voice agents.
Paid

OpenAI Realtime API

  • Speech-to-speech API for low-latency voice agents on OpenAI's realtime models.
  • The Realtime API provides WebRTC and WebSocket sessions to OpenAI's gpt-realtime models, which listen, reason, speak, and call tools without a separate transcription step. It is commonly used as the model layer inside voice agent stacks and is billed on usage.
Paid

Telnyx Voice AI

  • Carrier-grade voice AI agents on Telnyx's own telephony network.
  • Telnyx offers voice AI agents that run on its private telephony backbone, combining STT, LLM inference, and TTS in one platform with sub-200ms latency targets and support for 80+ languages. Pricing is an all-in per-minute rate that includes speech and orchestration.
Paid

Twilio ConversationRelay

  • Twilio Voice feature that connects phone calls to any LLM over WebSockets.
  • ConversationRelay handles the realtime speech-to-text, text-to-speech, and interruption management for a Twilio phone call and streams the conversation to a developer-controlled WebSocket, where any LLM can drive the dialog. It lets teams build phone voice agents on Twilio's telephony network with a choice of speech providers.
Open source

Jambonz

  • Self-hostable open-source voice platform for building conversational AI on telephony.
  • Jambonz is a CPaaS-style voice platform that connects SIP telephony with pluggable speech recognition, TTS, and LLM components, letting teams build custom voice AI applications without vendor lock-in. It can be self-hosted or run in the cloud, with flat-rate licensing for supported deployments.
Open source

TEN Framework

  • Open-source framework for realtime multimodal conversational AI agents.
  • TEN is an Apache-2.0-licensed framework for building voice agents with realtime multimodal capabilities, including components for VAD, turn detection, and avatar integration. It has an active community and is designed to work with a range of STT, LLM, and TTS providers.
Open source

Vocode

  • Open-source Python library for building voice-based LLM applications.
  • Vocode provides abstractions for streaming conversations that combine speech recognition, language models, and speech synthesis, including support for phone calls and realtime interruption handling. The MIT-licensed core library remains available on GitHub and is community-maintained.
Freemium

Daily

  • Realtime audio/video infrastructure and hosting for voice AI agents.
  • Daily provides WebRTC infrastructure APIs and SDKs for realtime audio and video, along with Pipecat Cloud for deploying voice agents built on the Pipecat framework. Pricing is usage-based with free and paid tiers.
Open source

LiveKit Agents

  • Open-source framework and cloud platform for realtime voice AI agents.
  • LiveKit Agents is an Apache-2.0 framework for building voice and multimodal agents in Python or Node.js on top of LiveKit's WebRTC infrastructure, with turn detection, interruption handling, telephony, and MCP tool support. LiveKit Cloud provides managed global media transport, agent deployment, and observability.
Freemium

Ultravox

  • Speech-native multimodal model and realtime platform for voice agents.
  • Ultravox is a multimodal LLM that processes audio directly without a separate speech recognition stage, reducing latency for conversational agents. The model weights are MIT-licensed on GitHub, and Ultravox Realtime offers a hosted platform with a free tier, per-minute pricing, and enterprise plans.
Open source

Bolna

  • Open-source framework and hosted platform for LLM-based voice agents.
  • Bolna is an end-to-end orchestration framework that composes ASR, LLM, and TTS providers over websockets to run voice conversations, with telephony support via Twilio and Plivo. The core framework is MIT-licensed, and the company offers a hosted platform with pay-as-you-go call pricing and multilingual support including Indian languages.
Freemium

Voiceflow

  • Collaborative platform for designing, building, and deploying conversational AI agents.
  • Voiceflow provides a visual builder and developer tooling for creating AI agents across voice, web chat, WhatsApp, and SMS channels. Teams use it to prototype conversation flows, integrate knowledge bases and APIs, and deploy agents to production with analytics.
Paid

PolyAI

  • Enterprise voice assistants for customer service call handling.
  • PolyAI builds and operates voice AI agents for large contact-center deployments in industries like hospitality, banking, and utilities. Its platform is built on proprietary dialog models trained on enterprise conversations and is sold on an enterprise contract basis.
Freemium

Hume AI

  • Voice AI APIs focused on emotional expression, including the EVI speech-to-speech interface.
  • Hume AI builds expression-aware voice technology, best known for its Empathic Voice Interface (EVI), a speech-to-speech API that detects vocal emotion in real time and modulates its responses accordingly. The company also offers expression measurement APIs and evaluation tooling for voice AI systems.
Paid

Millis AI

  • Low-latency voice agent platform with no-code and API options.
  • Millis AI provides tooling for creating LLM-powered voice agents with low response latency, aimed at developers and teams embedding voice into applications and phone workflows. Pricing is usage-based per minute of agent conversation.
Freemium

Synthflow

  • No-code and API-based platform for deploying AI phone agents.
  • Synthflow offers a platform for building voice agents for customer service, appointment booking, and lead qualification, with in-house telephony and integrations to 200+ CRM and contact-center tools. It uses a build-evaluate-launch workflow with real-time monitoring and pay-as-you-go call pricing.
Paid

Bland

  • Enterprise voice AI for automating phone calls at scale.
  • Bland provides infrastructure for building and running AI phone agents, with an emphasis on self-hosted models, per-minute pricing, and compliance requirements in regulated industries. It supports inbound and outbound calling with monitoring and enterprise deployment options.
Freemium

Retell AI

  • Platform for building AI voice agents that handle inbound and outbound phone calls.
  • Retell AI lets businesses build, test, deploy, and monitor conversational phone agents for use cases like call centers, appointment scheduling, and lead qualification. It includes agent-building tools, telephony integration, and analytics, with usage-based pricing and free starter credits.
Freemium

Vapi

  • Developer platform for building and deploying phone and web voice AI agents.
  • Vapi provides an orchestration layer for voice agents, handling telephony, streaming, interruption handling, and monitoring while letting developers plug in their choice of LLM, transcription, and speech synthesis providers. It targets teams shipping production voice agents at scale, with APIs, SDKs, and enterprise features.