Speech Recognition 13

AI-powered language technology services for translation and speech recognition in 100+ languages. Lingvanex offers award-winning language technology services, including machine translation and speech recognition. It provides AI-powered tools to translate text, documents, audio, and images into 100+ languages. Lingvanex offers on-premise machine translation, translation APIs, and SDKs for various platforms, including iOS, Android, Mac OS, and Windows. It also provides translators for PC, Slack, browsers, and mobile devices. Lingvanex focuses on secure communication, business intelligence, customer support, regulatory compliance, and forensic & e-discovery solutions. They cater to various industries, including defense and security, education, finance, government, legal, healthcare, manufacturing, marketing, media, retail, software, and travel.
AssemblyAI: AI models for speech-to-text transcription and voice data insights. AssemblyAI provides State-of-the-Art AI models for automatic speech recognition (ASR), natural language processing (NLP), and AI speech-to-text. It enables users to transcribe speech to text and extract insights from voice data. The platform offers speech-to-text, streaming speech-to-text, and speech understanding capabilities, catering to startups and enterprises for reliable source-truth data that powers world-class products.
Reka is an agentic multimodal AI platform for visual understanding and data insights. Reka is an AI research and product company that develops multimodal, modular intelligence solutions. Its platform, Reka Vision, specializes in agentic visual understanding and search across video, image, audio, and text, transforming raw unstructured data into deep insights and actions. Reka delivers complete AI solutions, from visual intelligence platforms for video editing and search to state-of-the-art web agents for researching complex questions, all powered by novel multimodal transformers built from scratch.
Accurate speech-to-text API and speech recognition service with various features and language support. Rev AI is a speech-to-text API and speech recognition service that offers accurate transcription at 0.3¢/min. It provides asynchronous and streaming APIs, human transcription services, and insights like topic extraction and sentiment analysis. Rev AI supports multiple languages and offers features like language identification and forced alignment.
Kardome offers voice user interface technology for clear voice command input in any environment. Kardome’s voice user interface technology clusters speech signals based on location, giving clear real-time voice command input and audio output in any environment. Kardome’s AI technology offers an all-in-one solution for manufacturers and OEMs looking to improve their existing speech recognition systems. Kardome’s break through technology improves voice recognition accuracy in challenging soundscapes, transforming voice UI from a cloud-dependent experience to a secure, real-time, and customizable user experience driven by neural network technology that is deployable to any smart device.
A Google AI-powered learning app providing answers and explanations for homework questions. Socratic is a learning app powered by Google AI that helps students get unstuck in their academic problems across subjects like Science, Math, Literature, and Social Studies. It provides answers, math solvers, explanations, and videos by taking a photo of a homework question. Socratic uses text and speech recognition to surface the most relevant learning resources, including visual explanations of important concepts.
Botjet is a conversational AI platform for building sophisticated chatbot solutions. Botjet is a conversational AI platform that provides the features and capabilities needed to build sophisticated chatbot solutions. It focuses on enabling businesses to drive deeper engagement through CUI-enabled digital touchpoints, making conversational AI adoption simple, lasting, and affordable. Botjet offers technologies like a conversation engine, deep learning, speech recognition, and speech synthesis to create human-like dialog flows for both speech and text conversations.
AI models for African dialects, bridging language barriers and enriching digital experiences. Neoform AI provides AI models for African dialects, aiming to make AI opportunities equally accessible. It bridges language barriers and enriches the digital experience for millions by offering voice assistant and translation services, multilingual customer support, localized navigation systems, transcription and captioning, public announcements and info dissemination, and localized content creation.
Privacy-focused macOS app for instant on-device dictation and AI-powered text processing. Spoke is a native macOS dictation application that provides instant voice-to-text transcription directly into any text field. It runs a high-performance local speech recognition model via CoreML, ensuring that audio never leaves the device for maximum privacy. Users can hold a customizable keyboard shortcut to speak and see their words appear instantly at the cursor. The app also features 'AI Skills,' allowing users to process transcriptions on the fly for tasks like translation, grammar correction, or summarization using their own API keys from providers like OpenAI, Anthropic, or local models via Ollama.
AI-powered voice agents for automated call handling. Callab AI revolutionizes call handling with AI-powered automation, providing AI-voice agents that help businesses automate and enhance customer interactions. It is designed to sound natural, act intelligently, and work at scale, handling support tickets, sales calls, appointment scheduling, and cold calling. Callab AI offers solutions for healthcare, real estate, debt collection, hospitality, retail & consumers, and contact centers.
Opensource voice-to-text macOS app with local AI for privacy and offline use. VoiceInk is an opensource voice-to-text app for macOS that transcribes what you say to text almost instantly with near-perfect accuracy. It uses local AI models to transcribe your speech to text, enabling offline functionality and ensuring data privacy. All data is stored locally, with optional AI enhancement.
Conversational AI APIs for creating interactive characters and speech-enabled applications. Convai offers Conversational AI APIs for Speech Recognition, Language Understanding, generation, and Text to Speech, enabling the design of games, speech-enabled applications, conversation-based Characters, and Speech-based games. It provides a service for games, metaverse, xr, and more, to bring characters to life with real-time perception and action abilities.
Platform for building low-latency voice AI agents with ASR, TTS, and LLM models. Hathora Models provides a platform for building voice agents on open-source or closed models with zero DevOps. It offers low-latency ASR (Automatic Speech Recognition), TTS (Text-to-Speech), and LLM (Large Language Model) models that run in 14 regions for ultra-low latency. Users can start instantly on shared endpoints and upgrade to dedicated infrastructure for privacy, compliance, or VPC requirements. The platform allows users to explore, test, and deploy production-ready models, bring their own models or custom containers, and utilize a "Chain tool" for interactive voice AI pipelines.