Speech-to-speech 5

Voice AI agents and audio models for business workflows
Boson AI is a voice AI platform for business-critical workflows. It provides real-time speech-to-speech agents, text-to-speech, speech-to-text, and avatar generation through its Higgs Realtime and Higgs Audio products. The platform supports low-latency conversations, interruption handling, tool calling, domain-specific training, more than 100 languages, and integration with existing production systems. Its API is compatible with the OpenAI Realtime API and is designed for customer support, sales, AI receptionists, and other business applications.

Voice AI platform for transcription, voice agents, and speech processing
Smallest AI is a voice AI platform offering speech-to-text, text-to-speech, speech-to-speech, voice cloning, and real-time voice agent technologies. Its Pulse speech-to-text models provide accurate transcription across 38+ languages, global accents, and dialects with latency as low as 64 milliseconds. The platform also supports speaker diarization, sentiment and emotion recognition, language identification, voice agent orchestration, telephony, knowledge bases, and enterprise deployment.

AI voice generator with realistic text-to-speech and speech-to-speech capabilities.
Respeecher Voice Marketplace is an AI voice generator platform that offers realistic text-to-speech and speech-to-speech capabilities. It provides a range of AI voice solutions for creative and professional projects, including film and TV production, game development, advertising, and more. The platform is trusted by industry leaders and offers high-quality AI voices, including celebrity voices, with a focus on ethical use and legal compliance.

Voice cloning and sound design app for cloning, mimicking, and designing voices.
Echo Voice AI is a revolutionary voice cloning and sound design app that empowers users to clone voices, mimic celebrity voices, clone their own voices, design entirely new voices, and transform their voice with Speech to Speech technology.

Ultra-low-latency voice AI APIs for speech generation, transcription, translation, and cloning
Gradium is a voice AI platform for developers that provides ultra-low-latency text-to-speech, speech-to-text, speech-to-speech translation, live translation, voice cloning, and on-device text-to-speech through a unified API. It is designed for building real-time voice agents and conversational applications, with expressive speech generation, accurate transcription, multilingual support, speaker cloning, bidirectional WebSocket streaming, scalable concurrency, and deployment options including cloud, dedicated instances, self-hosted, and on-premises infrastructure.