Speech-to-Text 10

Deepgram is a Voice AI platform offering STT, TTS, and voice agent APIs for developers. Deepgram is a Voice AI platform that provides APIs for speech-to-text, text-to-speech, and voice agent functionalities. It enables developers to build voice AI products and features with real-time, accurate, and scalable solutions. Deepgram's platform is trusted by top enterprises and startups for various use cases, including contact centers, medical transcription, and conversational AI.
AssemblyAI: AI models for speech-to-text transcription and voice data insights. AssemblyAI provides State-of-the-Art AI models for automatic speech recognition (ASR), natural language processing (NLP), and AI speech-to-text. It enables users to transcribe speech to text and extract insights from voice data. The platform offers speech-to-text, streaming speech-to-text, and speech understanding capabilities, catering to startups and enterprises for reliable source-truth data that powers world-class products.
Deep Learning-based spelling and grammar auto-corrector supporting 30+ languages with speech-to-text. NeuroSpell is a spelling and grammar auto-corrector based on Deep Learning, available in more than 30 languages. It includes a dictaphone (Speech-to-Text) for all languages and can be trained for specific in-domain vocabulary, phrasing, and error corrections. Other features include Human-in-the-loop charge optimization, Text-stream improvement/enrichment, Writing Aid, Proofreading RPA, Customer-workflow inputs enrichment, Speech-to-Text enhancement, and OCR error correction.
AI voice dictation app 4x faster than typing, converting messy speech into clear text. Genspark Speakly is an AI voice dictation application designed to convert spoken language into clear, polished messages, emails, and writings. It is marketed as being 4x faster than typing. The app integrates advanced AI features like Auto-Edits (which remove filler words, fix typos, and format text) and Custom Instructions (allowing users to define how their voice should be transformed, such as translation, CLI commands, or professional rewrites). It works across more than 100 applications and supports over 100 languages, making it a versatile productivity tool.
HIPAA-compliant AI agent for healthcare, using OpenAI with data security. CompliantChatGPT is an AI Agent designed to assist with healthcare-related tasks while ensuring patient data remains safe, secure, and HIPAA Compliant. It leverages OpenAI's GPT models with added security measures like PHI tokenization to maintain compliance.
Unifies speech recognition across 1,600+ languages using AI and LLM-enhanced decoders. Omnilingual ASR is an advanced automatic speech recognition technology that unifies speech recognition across a vast number of languages, scaling from dozens to over 1,600 natively and extending to 5,000+ via few-shot prompts. It achieves this by combining wav2vec-style self-supervision, LLM-enhanced decoders, and balanced multilingual corpora to learn language-agnostic acoustic patterns. This website serves as a comprehensive knowledge base, detailing its research breakthroughs, current technologies, datasets, implementation strategies, and deployment guidance for achieving omnilingual reach in a single model.
AI-powered video dubbing and translation service for creating multilingual videos. DubWiz is an AI-powered video dubbing service that allows users to translate and dub videos into multiple languages directly in their browser. It utilizes AI technologies like Speech-to-Text, Neural Machine Translation, and Neural Text-to-Speech to provide a user-friendly experience for creating multilingual videos.
Best AI-powered automated profanity word detection and censoring. Bleeper is the ultimate AI-powered censorship suite designed for YouTubers, TV stations, and media agencies to ensure broadcast compliance and protect revenue. Using studio-grade vocal isolation, the engine automatically detects and bleeps profanity while keeping background music perfectly intact, making it ideal for music channels and high-volume post-production. Beyond simple detection, Bleeper features smart song recognition with official lyrics alignment for 100% accuracy in music-heavy content. Available as a mobile-responsive web app for creators or a high-performance, unlimited Desktop version for professional studios, it saves hours of manual timeline scrubbing while maintaining full data privacy. Our upcoming roadmap includes advanced visual content moderation for alcohol and NSFW detection, positioning Bleeper as the comprehensive safety standard for the modern media industry.
AI scribe for healthcare to generate clinical notes and billing codes. Claio is an AI scribe designed for healthcare professionals. It transcribes patient visits and instantly generates structured clinical notes, which can then be copied and pasted into an Electronic Health Record (EHR) system. Claio also provides accurate ICD-10 and CPT code suggestions to simplify medical billing, reduce denials, and ensure faster payments. The platform aims to reduce documentation time, simplify daily tasks, and support better patient outcomes by transforming conversations into ready-to-use clinical documentation. It is HIPAA-compliant and developed in partnership with healthcare professionals to fit real workflows.
All-in-one audio AI platform for transcription, text-to-speech, dubbing, and captioning. SIREN is an all-in-one audio AI platform designed to provide solutions for audio transcription, audio pen, text-to-speech, video dubbing, and live stream captioning. It leverages cutting-edge GPU-empowered technologies to transform thoughts into text, generate audio from text, and make content understandable internationally.