ASR 8

AssemblyAI: AI models for speech-to-text transcription and voice data insights.
AssemblyAI provides State-of-the-Art AI models for automatic speech recognition (ASR), natural language processing (NLP), and AI speech-to-text. It enables users to transcribe speech to text and extract insights from voice data. The platform offers speech-to-text, streaming speech-to-text, and speech understanding capabilities, catering to startups and enterprises for reliable source-truth data that powers world-class products.

Gladia is a production-ready Speech-to-Text API for teams shipping voice products—high accuracy, multilingual, real-time + async, and add-ons.
Gladia is a speech-to-text platform built for production, turning raw audio into structured outputs that power real workflows like meeting summaries, CRM enrichment, contact center QA, and real-time voice assistants. With support for 100+ languages and the ability to handle messy real-world audio—overlapping speakers, accents, code-switching, domain-specific terminology—Gladia is designed for the complexity of actual conversations, not clean studio recordings.

Accurate speech-to-text API and speech recognition service with various features and language support.
Rev AI is a speech-to-text API and speech recognition service that offers accurate transcription at 0.3¢/min. It provides asynchronous and streaming APIs, human transcription services, and insights like topic extraction and sentiment analysis. Rev AI supports multiple languages and offers features like language identification and forced alignment.
Unifies speech recognition across 1,600+ languages using AI and LLM-enhanced decoders.
Omnilingual ASR is an advanced automatic speech recognition technology that unifies speech recognition across a vast number of languages, scaling from dozens to over 1,600 natively and extending to 5,000+ via few-shot prompts. It achieves this by combining wav2vec-style self-supervision, LLM-enhanced decoders, and balanced multilingual corpora to learn language-agnostic acoustic patterns. This website serves as a comprehensive knowledge base, detailing its research breakthroughs, current technologies, datasets, implementation strategies, and deployment guidance for achieving omnilingual reach in a single model.

Multilingual Speech-to-Text API with high accuracy in 14 languages.
SpeechFlow is a multilingual Speech-to-Text API that offers state-of-the-art accuracy in 14 languages. It converts sound to text, speech to text, and audio to text with high accuracy. SpeechFlow supports both cloud and on-prem deployment.

A guide for OpenAI's GPT-4o model, explaining its features and usage.
This is the guide of the GPT-4o, which is new model of OpenAI. GPT-4o can reason across text, audio, and video in real time. The guide provides information on how to use GPT-4o, its features, and its capabilities. It also includes frequently asked questions and support resources.

Platform for building low-latency voice AI agents with ASR, TTS, and LLM models.
Hathora Models provides a platform for building voice agents on open-source or closed models with zero DevOps. It offers low-latency ASR (Automatic Speech Recognition), TTS (Text-to-Speech), and LLM (Large Language Model) models that run in 14 regions for ultra-low latency. Users can start instantly on shared endpoints and upgrade to dedicated infrastructure for privacy, compliance, or VPC requirements. The platform allows users to explore, test, and deploy production-ready models, bring their own models or custom containers, and utilize a "Chain tool" for interactive voice AI pipelines.
Memory assistant for Mac that records and allows semantic search.
MemFlow is a memory assistant for Mac users. It automatically records screenshots and sounds in chronological order, allowing users to find memories through semantic search. It captures everything seen, said, or heard on the Mac, leveraging AI to search and generate information, saving users up to 30 minutes per day.