Voice agents 5

Voice AI platform for transcription, voice agents, and speech processing
Smallest AI is a voice AI platform offering speech-to-text, text-to-speech, speech-to-speech, voice cloning, and real-time voice agent technologies. Its Pulse speech-to-text models provide accurate transcription across 38+ languages, global accents, and dialects with latency as low as 64 milliseconds. The platform also supports speaker diarization, sentiment and emotion recognition, language identification, voice agent orchestration, telephony, knowledge bases, and enterprise deployment.

Open-weights 8B AI text-to-speech model for expressive English speech.
Miso One is an open-weights, 8B-parameter text-to-speech (TTS) system developed by Miso Labs. It is designed specifically for producing highly realistic, expressive, and emotionally varied English conversational speech, making it ideal for voice-agent research and developer workflows. Built on a Sesame-style conversational speech model (CSM) architecture with Mimi audio codes, it features a highly optimized inference capability boasting a published low latency of 110 ms. In addition to text-to-speech generation, the model supports voice continuation and one-shot voice cloning from audio context with clear consent boundaries.

Ultra-low-latency voice AI APIs for speech generation, transcription, translation, and cloning
Gradium is a voice AI platform for developers that provides ultra-low-latency text-to-speech, speech-to-text, speech-to-speech translation, live translation, voice cloning, and on-device text-to-speech through a unified API. It is designed for building real-time voice agents and conversational applications, with expressive speech generation, accurate transcription, multilingual support, speaker cloning, bidirectional WebSocket streaming, scalable concurrency, and deployment options including cloud, dedicated instances, self-hosted, and on-premises infrastructure.

Multilingual voice AI routing and infrastructure platform
Speko is a voice AI routing and infrastructure platform that provides a unified API for speech-to-text, language models, and text-to-speech. It benchmarks speech and language models across multiple languages, routes each session to suitable-performing models, and supports provider-direct or managed voice workflows through integrations with LiveKit, Pipecat, OpenAPI, AsyncAPI, and MCP.

Uberduck is an AI platform for voice-over, text-to-speech, voice cloning, and AI music generation.
Uberduck provides users with tools to create voice-over audio with over 5,000 expressive voices, custom voice clones, APIs to build audio applications, and AI-generated raps. It also offers features like text to speech, voice conversion, and AI music generation. There is a case study to demonstrate how it can be used to create personalized media and a waitlist to join the upcoming Uberbots platform.