AI Speech Recognition 148

HaloVoice is an AI-powered real-time voice translator for gaming, streaming, voice chat, and online meetings, featuring voice cloning and bilingual subtitles.
HaloVoice is an AI-powered real-time voice translator designed for gaming, live streaming, voice chat, online meetings, and multilingual communication. It translates spoken conversations across languages in real time, helping people communicate naturally without constantly switching to a text translator.
Real-Time Voice Translation
HaloVoice provides live voice translation while you speak. Instead of typing text into a translator, users can speak naturally and send translated speech directly to other people in real time. Bilingual subtitles display both sides of the conversation, making multilingual communication easier to follow.
With AI voice cloning, HaloVoice can preserve characteristics of the speaker’s voice in translated speech, creating a more natural and personal experience.
Real-Time Voice Translator for Gaming
HaloVoice works as a real-time voice translator for gaming, helping players communicate with international teammates while they play. Translated audio can be sent directly into voice chat through the HaloVoice Virtual Microphone, allowing other players to hear the translated voice in real time.
HaloVoice can be used as a gaming voice translator across multiplayer games and gaming communities, including use cases such as a Valorant voice translator and Minecraft voice translator.
For players using Discord, HaloVoice provides real-time voice translation for Discord. Simply select the HaloVoice Virtual Microphone as your microphone input in Discord to send your translated voice directly into the voice chat.
Real-Time Voice Translator for Streaming
HaloVoice provides real-time voice translation for streaming, allowing streamers to speak in one language while reaching audiences who speak another.
For creators using OBS, HaloVoice provides real-time voice translation for OBS through the HaloVoice Virtual Microphone. Streamers can select HaloVoice as their microphone input in OBS and send translated speech directly into their stream.
HaloVoice can also support multilingual live streaming workflows on platforms such as Twitch, helping creators communicate with audiences across languages without interrupting the broadcast.
Real-Time Voice Translation for Online Meetings
HaloVoice also provides real-time voice translation for Zoom, Slack, Microsoft Teams, and other communication platforms. Users can select the HaloVoice Virtual Microphone as their microphone input so other participants hear the translated voice directly during conversations.
This makes HaloVoice useful for multilingual online meetings, remote teams, international collaboration, and everyday cross-language voice communication.
HaloVoice supports practical language-pair use cases including Spanish to English voice translation, English to Korean voice translation, and other multilingual conversations.
Key Features
• Real-time voice translation
• AI voice translator for multilingual conversations
• Real-time voice translator for gaming
• Real-time voice translation for streaming
• Real-time voice translation for Discord and voice chat
• Real-time voice translation for OBS and live streaming
• Real-time voice translation for Zoom, Slack, and Microsoft Teams
• AI voice cloning for translated speech
• HaloVoice Virtual Microphone
• Bilingual subtitles
• Spanish to English voice translation
• English to Korean voice translation
• Support for Windows and macOS
HaloVoice is built for gamers, streamers, creators, remote teams, and anyone who wants to communicate across languages while keeping conversations fast, natural, and voice-first.

AI-powered platform enhancing global communication through noise cancellation and accent translation.
Sanas uses AI to enhance global communication by offering noise cancellation and accent translation. It allows users to control how they sound while retaining their unique voice, breaking down linguistic barriers and bridging communication across cultures worldwide. Sanas provides a real-time speech understanding platform with features like Accent Translation and Noise Cancellation, operating in over 200 territories. It serves contact centers and enterprises, enabling cost performance improvements and better customer satisfaction.

AI-powered work and life planning app with voice dictation and auto-scheduling.
Voiset is an AI-powered work and life planning app that uses AI to plan your day. It optimizes workflow with AI and voice dictation, offering features like auto-scheduling, speech recognition, smart notes, task management, and a ChatGPT plugin. Voiset caters to personal and team use, helping self-employed individuals, students, startups, freelancers, and accounting agencies manage tasks, schedules, and projects effectively.

24/7 AI speech therapist providing real-time feedback, personalized exercises, and fluency tracking.
Disertus Labs offers Milo, an AI Speech Therapist accessible 24/7 via iMessage, Web, and WhatsApp. Milo allows users to practice speech therapy anytime, anywhere, receiving instant feedback, personalized exercises, and comprehensive progress tracking. It utilizes real-time conversation, live transcription (detecting filler words and repetitions), session summaries (with speech metrics and tips), and an extensive library of 47 guided exercises across categories like breathing, pronunciation, and fluency to help users gain confidence and clarity in their speaking.

Voice-first AI calendar assistant for busy families.
KIN is a Voice-first AI-powered Calendar Assistant designed for busy families. It simplifies scheduling by offering smart reminders and seamless coordination. KIN functions as a shared family calendar app that aims to make managing family life a delight. It allows users to create and edit events effortlessly using voice commands or by typing, set up recurring events by speaking, modify existing plans with voice, and even create events from photos or screenshots.

TikTok influencer database with AI-powered search and analytics for precise targeting.
topYappers is a comprehensive TikTok influencer database that provides access to over 14 million verified creators. It offers advanced AI-powered search and filtering capabilities, allowing users to precisely target influencers based on niche, engagement quality, audience overlap, and recent product promotions. The platform analyzes millions of videos to match brands with creators who already promote similar products, detecting specific product mentions, brand names, and visual appearances. topYappers also offers API access for seamless integration with existing tools and workflows, along with advanced analytics to track creator performance.

Travel dating app for train travelers with AI, AR, and gamification features.
Raily is a travel dating app tailored for train travelers, integrating AI, AR, and gamification to enhance the user experience. It connects people on the move, offering features like Visual Matchmaking, Music Taste Matching, AI Companions, and a Travel Concierge. Raily aims to revolutionize dating for globetrotters, providing opportunities for friendship, romance, travel buddies, or business connections.

AI-powered live captioning software with real-time subtitles in 90+ languages.
Akkadu is an AI-powered live captioning software that provides real-time subtitles in over 90 languages. It's designed to help users understand videos, webinars, video conferences, and live streams in their own language. Akkadu is compatible with various platforms like Zoom, Teams, YouTube Live, Netflix, and more.

Perso Dubbing is an AI video dubbing platform that translates, dubs, and lip-syncs videos into 99+ languages. AI voice cloning preserves each speaker's tone and emotion, and multi-speaker detection handles up to 10 speakers per video. It reduces localization costs by up to 98% compared to traditional dubbing studios. Developed by ESTsoft and trusted by 450,000+ users.
Perso Dubbing is an AI-powered video dubbing and translation platform that localizes content into 99+ languages in minutes, with speech recognition in 100+ languages. Teams upload a video, select target languages, and receive a studio-quality dubbed version — complete with lip-sync and voice cloning that preserves the original speaker's tone, accent, and emotion.
Key capabilities:
• AI Voice Cloning — Matches the original speaker's voice, accent, and emotional tone across all dubbed tracks
• AI Lip Sync — Aligns translated audio with on-screen mouth movements for natural viewing
• Speech-to-Text — Speech recognition in 100+ languages
• Audio Separation — Splits voice and background tracks
• Auto Subtitle Generation — Creates and exports subtitles automatically
• Real-Time Script Editor — Review and refine translations before final export
• Multi-Speaker Support — Detects and dubs up to 10 speakers in a single workflow
Built for marketing teams, e-learning creators, enterprise L&D departments, and media publishers expanding into global markets. Enterprise plans include API access, advanced security controls, and dedicated support. Developed by ESTsoft (est. 1993, KOSDAQ: 047560) — ISO/IEC 27001 and KISA ISMS certified.

AI dubbing and voiceover tool for media and entertainment with cost-effective localization.
Dubformer is an AI dubbing and voiceover tool for media and entertainment, offering monetization-ready dubbing that sounds like real humans. It provides cost-effective AI localization solutions, including best-in-class AI translation and dubbing, unmatched control over AI dubbing, and seamless workflow integration. Dubformer caters to media companies, localization companies, and businesses aiming to reach international audiences.

Transcribe audio, video, YouTube links, and recordings into accurate, searchable text.
Audio Converter AI makes it easy to transcribe audio to text online without installing software. Convert interviews, meetings, lectures, podcasts, videos, voice recordings, and YouTube content into accurate, searchable, and editable transcripts.
The AI transcription engine supports more than 200 languages and can automatically detect the source language. Speaker recognition helps separate conversations, while timestamps make it easy to find important moments in long recordings. After transcription, you can use AI-generated summaries and smart notes to understand content faster and create reusable materials.
Audio Converter AI supports popular formats including MP3, MP4, M4A, WAV, WEBM, MOV, AIFF, OPUS, FLAC, AVI, MKV, FLV, and 3GPP. Files can be up to 10GB, with multiple tasks supported in the queue. Your content is processed in encrypted environments, and files are automatically removed after processing.

AI-powered language learning and assessment platform with AI tutors and assessments in 60+ languages.
Hallo is an AI-powered language learning platform for speaking. It provides fast, affordable, and accurate AI-driven language assessments across speaking, writing, listening, and reading skills, available in over 60 languages. It also offers AI Language Tutor.

AI-powered tool for accent identification and speech analysis.
Accent Guesser is an AI-powered tool designed for speech analysis, focusing on identifying and analyzing accents. It utilizes deep learning to analyze voice patterns, providing quick and reliable accent analysis. The platform aims to offer insights into users' linguistic backgrounds and enhance communication skills through accent identification and analysis. It is designed with a user-centric interface for ease of use and offers features like global accent recognition and comprehensive data analysis to improve accuracy.

Babbly is an AI-powered tool for early speech therapy and infant development monitoring.
Babbly is an early speech therapy tool that transforms playtime into progress. It uses AI-powered infant speech and brain development monitoring to identify the risk of developmental delays as early as 9 months. Babbly helps parents understand their child’s development by analyzing and monitoring their language progression and recommending activities to accelerate their development. It provides objective data to inform parental intuition and helps parents find out if their child is at risk of speech and language delays, which can be a sign of developmental conditions such as autism.

Kardome offers voice user interface technology for clear voice command input in any environment.
Kardome’s voice user interface technology clusters speech signals based on location, giving clear real-time voice command input and audio output in any environment. Kardome’s AI technology offers an all-in-one solution for manufacturers and OEMs looking to improve their existing speech recognition systems. Kardome’s break through technology improves voice recognition accuracy in challenging soundscapes, transforming voice UI from a cloud-dependent experience to a secure, real-time, and customizable user experience driven by neural network technology that is deployable to any smart device.

Speech recognition and translation software for real-time typing, transcription, and subtitle generation.
SpeechPulse is a speech recognition and translation software that uses your computer’s microphone for real-time speech recognition. It can type into your favorite apps, including text editors, web browsers, and office applications. It can also transcribe audio/video files and generate subtitles. It supports offline speech recognition for ultimate privacy and transcription in 99 languages, including English translation.

AI solutions for audio analysis and speech emotion recognition, enabling empathetic AI interactions.
audEERING provides advanced AI solutions for audio analysis and speech emotion recognition. Their technology transforms industries by enabling machines to understand and respond to human vocal expression, creating empathetic AI interactions. They offer products like devAIce®, devAIce® XR, and AI SoundLab, catering to various use cases such as market research, automotive, robotics, healthcare, and extended reality applications.

AI-powered TOEFL Speaking prep with SpeechRater™ for accurate feedback and score prediction.
My Speaking Score is an AI-powered platform designed to help non-native English speakers prepare for the TOEFL Speaking section. It utilizes ETS's SpeechRater™ technology to provide accurate score predictions and actionable feedback on response delivery, language use, and topic development. The platform offers unlimited practice tests, sharable reports, and personalized insights to help users improve their speaking performance and achieve their target TOEFL score.

AI-powered tool to automatically remove profanity from videos.
Bleepify is an AI-powered tool that automatically removes profanity from videos. It uses advanced AI models to detect and censor offensive language from audio and video files with speed and precision. It supports over 40 languages and allows users to edit, review, and download videos effortlessly. Bleepify is designed for podcasters, video creators, and media managers to ensure their content is clean, professional, and audience-ready.

AI-powered English speaking coach for employees, offering personalized feedback and secure language training.
Lucida AI is an AI-powered English speaking coach designed to help employees enhance their communication skills. It offers personalized feedback on pronunciation, grammar, vocabulary, and fluency through real-time conversations with Lucy, an AI coach. Lucida AI prioritizes privacy with end-to-end encryption and can be tailored to company regulations. It provides comprehensive language training at an affordable price, ensuring every team member can benefit from advanced AI-driven coaching.

Accurate speech-to-text API and speech recognition service with various features and language support.
Rev AI is a speech-to-text API and speech recognition service that offers accurate transcription at 0.3¢/min. It provides asynchronous and streaming APIs, human transcription services, and insights like topic extraction and sentiment analysis. Rev AI supports multiple languages and offers features like language identification and forced alignment.

Ello is an AI reading coach for kids in Kindergarten to 3rd Grade.
Ello is building the most natural AI teacher to maximize the learning potential of all children, regardless of resources. Its first product is the world’s most advanced reading coach, powered by proprietary speech recognition and generative AI. Ello is your child’s read-along companion who listens, teaches, and transforms them into an enthusiastic reader. For Kindergarten to 3rd Grade.

Conversation Experience Platform with Generative AI and Speech Recognition.
Seasalt.ai provides a Conversation Experience Platform with Generative AI and Speech Recognition. It aims to help businesses capture, generate, and understand all text and voice conversations with customers, enabling natural, personalized, and actionable interactions.

Online teleprompter with voice-activated scrolling and collaboration features.
Speakflow is an online teleprompter that allows users to write and save scripts, collaborate with their team, and includes voice-activated scrolling. It works across various platforms including Windows, Mac, iOS, and Android, with no downloads required. It aims to reduce production time, help deliver better presentations, and enable video recording directly in the browser. It is also compatible with physical teleprompter hardware.
Hot Articles
How to Build a Multimodal AI Knowledge Base With Gemini Embedding 2
NASA 和 IBM 开源月球模型
Introducing the Australian Youth Safety Blueprint
GLM-5.3-FlashX 上线,智谱把国产卡上的推理速度顶到 200 token/s
Google’s Gemini is the latest AI model to hack other companies
How to Choose Hardware for Running Local LLMs, and Know Exactly When It Beats the Claude API
Latest Articles
Google’s Gemini is the latest AI model to hack other companies
How to Choose Hardware for Running Local LLMs, and Know Exactly When It Beats the Claude API
How to Build a Multimodal AI Knowledge Base With Gemini Embedding 2
Announcing Grok-1.5
Google DeepMind launches institute to widen the AGI debate
Google’s new ‘CC’ is an AI agent that helps families run their households
Hot Tags