Speech recognition 40

Telegram bot that transcribes voice and video notes to text in multiple languages.
EchoScribe is a Telegram bot that automatically transcribes voice notes and video notes into plain text. It utilizes world's best speech recognition software to transcribe any voice/video note coming its way. It understands talking and mumbling in English, Spanish, German, Italian, Chinese, & 52 other languages. It is powered by Fruition.

Multilingual Speech-to-Text API with high accuracy in 14 languages.
SpeechFlow is a multilingual Speech-to-Text API that offers state-of-the-art accuracy in 14 languages. It converts sound to text, speech to text, and audio to text with high accuracy. SpeechFlow supports both cloud and on-prem deployment.

AI-powered online subtitles editor for social media videos.
Subtitles.Love is an AI-powered online subtitles editor that allows users to add subtitles to their social media videos to increase audience interaction. It offers features like automatic speech recognition, resizing, and styling for various social media platforms. The platform supports multiple languages and video formats, aiming to simplify and speed up the process of creating subtitled videos.
RudeCaptcha uses AI to verify humanity by detecting offensive content.
RudeCaptcha is a new way to prove you're human by using AI to check if you're human. It leverages the fact that AI bots infesting the internet aren't allowed to be offensive. Users verify their humanity by swearing at the camera or copying a rude gesture shown on the screen.

Video localization tool using AI for translation and voiceover.
Langswap is a video localization tool that helps content creators reach global audiences. It uses speech recognition and voice cloning technologies to translate and voiceover videos in minutes, without the need for voice actors. Langswap allows users to translate videos without re-recording, making the same voice speak in another language. It saves time and money by automating the dubbing process.

Real-time multilingual voice translation and transcription for international meetings, supports Microsoft Teams.
CSC Voice AI offers real-time multilingual voice translation and transcription features. It provides high-accuracy speech recognition for 24 other languages, and generates detailed reports for meetings. It helps break down language barriers in international meetings by supporting multilingual voice translation and transcription using Azure AI technology. It is compatible with Microsoft Teams.
Real-time voice translation app breaking down language barriers with AI.
Dialects is a cutting-edge real voice translation app designed to break down language barriers and facilitate seamless communication between people who speak different languages or dialects. Dialects leverages advanced speech recognition and machine translation technologies to convert spoken words from one language or dialect into another, enabling real-time, voice-based conversations. Dialects supports a wide range of languages and dialects, including but not limited to English, Spanish, French, Chinese, Arabic, and many more. Our goal is to continuously expand the language and dialect offerings.

AI-powered tool for secure patient-doctor sessions with transcription, analysis, and translation.
MediScoper is an AI-powered tool designed to enhance patient-doctor sessions by providing real-time audio transcription and analysis. It offers diagnostic insights, automated session reports aligned with SOAP standards, and translation in over 60 languages. MediScoper aims to reduce administrative burdens, improve patient care, and ensure data security through anonymous data processing and state-of-the-art security measures.

Open-source AI organization focusing on engineering implementation of AI models.
RapidAI is an open-source organization dedicated to bridging the gap between AI models in academia and practical engineering applications. It focuses on the engineering implementation of AI technologies, including computer vision, natural language processing, and speech, without training models but applying them effectively. RapidAI aims to provide simple, effective, and ready-to-use solutions to lower the barrier to AI adoption.

Automatic transcription service for audio and video files, focusing on speed and accuracy.
Good Tape is an automatic transcription service that makes it easy for journalists (and others) to turn audio recordings into text, regardless of language or sound quality. We save you time and effort so you can focus on what really matters.

Platform for on-device speech AI, enabling speech recognition and wake word detection.
Wavify is a one-stop-shop for voice AI, providing a platform for on-device speech AI. Software engineers can embed features like speech recognition and wake word detection into any software. It offers SOTA models and a cross-platform inference engine, optimized for speed and privacy. Wavify supports multiple languages and runs on various platforms, including Linux, Mac, Windows, iOS, Android, Web, Raspberry Pi, and embedded systems.

AI-powered audio transcription service offering fast, accurate, and affordable transcriptions in multiple languages.
transcribethis.io is an AI-powered audio transcription service that delivers accurate and precise transcriptions, allowing users to focus on important tasks. It offers a faster and cheaper alternative to manual transcription, with enterprise-level AI trained on millions of hours of audio. The service supports nearly 60 languages and provides options for transcribing interviews, conference calls, podcasts, and lectures.

AI-powered language learning app with 3D lessons and speech recognition.
Langony is an AI-powered language learning app that features interactive 3D lessons, speech recognition, and a voice assistant to help users boost their language skills. It supports learning English, Spanish, German, French, Russian, and Italian. Langony aims to make language learning fun and effective, offering engaging lessons and a unique storyline in each lesson.

AI voice assistant for scientific labs, enabling hands-free lab interactions.
Ascenscia is an AI voice assistant for scientific labs with cutting-edge voice technology that understands scientific terms with up to 97% accuracy. It integrates with laboratory software to enable hands-free interactions, allowing scientists to speak to their lab's data and automate, optimize, and accelerate their workflows. Ascenscia aims to improve data accessibility, data capturing, inventory management, and other tasks within the lab environment.
AI note-taking app that summarizes voice notes and generates content.
Speakpen.cc is an AI Note Taking App that summarizes your voice notes and helps you generate content. It transforms scattered thoughts into persuasive, organized articles with ease, streamlining your thinking process to create structured and clear written expressions. The app accurately records thoughts, insights, and to-do lists using speech recognition and natural language processing, and provides tailored tips, reminders, and content suggestions by analyzing your notes.
Nutrition app using AI to estimate meal macros from descriptions, no calorie counting.
NutritionBuddy is an app designed to help users improve their eating habits without the hassle of traditional calorie tracking. It uses speech recognition and artificial intelligence to transform simple meal descriptions into macronutrient tracking records, providing insights into eating habits without manual calorie counting.
Pronunciation assessment API with voice AI model.
SpeechEvalPro is a platform offering pronunciation assessment and scoring API solutions. It utilizes an independently researched and developed educational voice AI model, integrating voice evaluation, speech recognition, and other core technologies to provide high-quality, multi-dimensional Chinese and English pronunciation evaluation APIs. It helps customers create intelligent learning products for human-computer interaction.

Real-time STT/TTS solution using AI-focused Sense Theory for nuanced speech processing.
Speech Intellect is the first STT/TTS solution that works in real-time by totally using a new AI-focused mathematical theory — "Sense Theory". It looks at the sense of each word pronounced by the client. It offers speech-to-text, text-to-speech, and combining solutions, leveraging a sense-to-sense algorithm to reproduce text with intonation and tonality. The platform emphasizes security with Amorphous Encryption and provides flexibility in shaping work scenarios for various business needs.

AI-powered web app for audio transcription, translation, and summarization in 143 languages.
DenoLyrics is a web application built with an AI model that supports 143 languages, no matter if the audio speed is fast or slow. It converts audio to text using artificial intelligence. It allows users to transcribe audio files, create captions, summarize text, and translate languages. Users can also share folders with friends for collaboration.
Mobile app for real-time video translation with in-video closed captions.
TalkVisions is a mobile application that eliminates language barriers by offering in-video closed captioning translations. It uses advanced speech recognition technology to transcribe and translate spoken words into a chosen language, displaying the text as subtitles in videos. TalkVisions allows users to capture and translate spoken language on the fly, making it a powerful tool for communication and learning.

AI-powered video search and analytics platform for efficient video content management.
QuickSight is an AI-powered video search and analytics platform that enables users to quickly search for objects, actions, conversations, or text within their video libraries at scale using natural language. It provides tools for video search, collaborative video review, and AI video generation.

Real-time AI interview assistant providing AI-powered answers and interview support.
ParakeetAI is a real-time AI interview assistant designed to help users excel in job interviews. It uses AI, specifically GPT-4.1, to provide accurate and helpful answers to interview questions. The tool offers features like real-time speech recognition, fast transcription, resume integration, and multilingual support. It works with various video calling platforms and provides post-interview analysis and recommendations.

Converts audio/video to text, summaries, and insights quickly and accurately.
Transcript LOL is a service that converts audio and video into text, summaries, and more in seconds. It offers features like speaker recognition, high accuracy, and the ability to download in multiple formats. It's used for course content, extracting key points from meetings or interviews, and creating social media posts.

AI-powered app to improve English pronunciation and speaking skills with personalized feedback.
ELSA (English Language Speech Assistant) Speak is an AI-powered app designed to improve English pronunciation and speaking skills. It uses voice recognition technology developed with data from people speaking English with various accents to provide instant, detailed feedback on pronunciation, fluency, intonation, grammar, and vocabulary. ELSA offers personalized lessons, interactive games, and real-world conversations to help learners speak English more confidently and clearly.