Speech-to-text 57

Kensho is an AI toolkit for data insights, offering transcription, entity recognition, data linking, and PDF extraction. Kensho is an AI toolkit designed to help users uncover insights in their data. It offers solutions for speech-to-text transcription, entity recognition in text, mapping companies to external databases, and extracting data from PDF documents. Kensho's tools aim to automate workflows, improve data quality, and accelerate research processes.
AI-powered form builder with voice input, multilingual support, and digital signatures. Formcraft is an AI-powered voice-to-form software designed to streamline form creation and improve user experience. It allows users to create forms by describing their needs in plain English or using voice input. Formcraft offers features like multilingual support, speech-to-text responses, digital signatures, and real-time analytics to boost completion rates and save time.
Voice-powered app for easy form filling and creation through speech. SpeechForms is a voice-powered app that allows users to fill out forms by speaking instead of typing. It aims to make form-filling as natural as having a conversation, simplifying the process of creating and sending forms.
AI writing assistant with personalized suggestions and plagiarism checker. PenPilot AI is an AI-powered writing assistant designed to help users effortlessly elevate their academic papers, blogs, and more. It offers personalized suggestions, seamless speech-to-text functionality, and a built-in plagiarism checker, enabling users to write smarter, faster, and perfect their content with ease. PenPilot AI writes like the user, making it easy to create natural, high-quality content that resonates with their audience.
A digital journaling app that replicates a physical journal with voice recording and more. Journalizr is a journaling app designed to replicate the experience of a physical journal in the digital space. It allows users to record their thoughts in various ways, including voice recording with transcription, writing, doodling, and image storage. The app aims to provide a spontaneous and creative journaling experience, offering a simple and accessible way to document thoughts and ideas without usage limits.
Online AI tools for vocal removal, stem splitting, and audio transcription EZAudio is a browser-based AI audio toolkit for removing vocals, splitting music into stems, and converting audio or video recordings into text. It supports karaoke and backing-track creation, music production, meeting transcription, interviews, lessons, voice notes, and content workflows without requiring desktop software. Files are encrypted during transfer, automatically deleted within 24 hours, and not used for AI training.
AI-powered assistant for X content ideation, creation, and scheduling. Postel is an AI-powered personal brand assistant for X (formerly Twitter) that streamlines content ideation, creation, and scheduling. It helps creators and businesses generate engaging posts and repurpose content to grow their brand. Postel analyzes thousands of viral posts to create high-performing content that matches the user's unique voice. It aims to provide a seamless content creation flow, turning raw thoughts into personalized, algorithm-loved posts without low-effort AI content.
AI-powered transcription and meeting minutes service with real-time transcription and translation. Notta is a high-precision transcription service equipped with the latest AI speech recognition engine. It features real-time transcription and translation, and can quickly transcribe audio files up to 5 hours long at a time. It allows for easy audio conversion and editing on PC.
Distributed GPU cloud offering compute, storage, and deployment solutions at lower costs. SaladCloud offers distributed GPU cloud services, including compute, storage, and deployment solutions. It provides access to a network of consumer GPUs for tasks like image generation, voice AI, computer vision, data collection, batch processing, and molecular dynamics. The platform aims to democratize cloud computing by offering lower-cost alternatives to traditional cloud providers, particularly for AI transcription and GPU-intensive workloads.
Automated transcription, translation, and subtitling platform for audio/video. Sonix is an advanced automated transcription, translation, and subtitling platform that converts audio and video files to text quickly, accurately, and affordably. It leverages industry-leading speech-to-text AI algorithms to transcribe various content types like podcasts, interviews, speeches, meetings, and films. Beyond transcription, Sonix offers automated translation, AI analysis tools (summaries, topic detection), automated subtitling, and features for sharing, collaboration, organization, and integration with popular workflows.
Gladia is a production-ready Speech-to-Text API for teams shipping voice products—high accuracy, multilingual, real-time + async, and add-ons. Gladia is a speech-to-text platform built for production, turning raw audio into structured outputs that power real workflows like meeting summaries, CRM enrichment, contact center QA, and real-time voice assistants. With support for 100+ languages and the ability to handle messy real-world audio—overlapping speakers, accents, code-switching, domain-specific terminology—Gladia is designed for the complexity of actual conversations, not clean studio recordings.
A note-taking app with speech-to-text, supporting 50+ languages and AI summarization. Dictanote is a modern notes app with built-in speech-to-text integration, allowing users to effortlessly voice type their notes in over 50 languages. It offers a seamless experience between keyboard and voice input, enhancing productivity with accurate dictation and transcription. Dictanote also provides features like AudioScribe, a smart AI writing assistant that converts voice notes into clearly summarized text.
Voice AI agents and audio models for business workflows Boson AI is a voice AI platform for business-critical workflows. It provides real-time speech-to-speech agents, text-to-speech, speech-to-text, and avatar generation through its Higgs Realtime and Higgs Audio products. The platform supports low-latency conversations, interruption handling, tool calling, domain-specific training, more than 100 languages, and integration with existing production systems. Its API is compatible with the OpenAI Realtime API and is designed for customer support, sales, AI receptionists, and other business applications.
AI-powered audio and video transcription service with high accuracy and multi-language support. AccurateScribe.ai is an enterprise-grade audio and video transcription service powered by advanced AI technology. It converts audio and video files into accurate text, supporting over 134 languages with 99.8% AI accuracy. Users can transcribe unlimited audio and video, export in multiple formats (PDF, DOCX, TXT, SRT, VTT), and utilize features like speaker recognition and audio enhancement.
AI-powered transcription service converting audio and video to text in 117+ languages. TranscribeToText.AI is an AI-powered transcription service that converts audio and video into text in 117+ languages with high accuracy. It supports YouTube videos, cloud storage (Google Drive, Dropbox), and live meeting transcriptions from Zoom, Google Meet, and Microsoft Teams. It offers unlimited transcription with support for files up to 10 hours long or 5GB each. Transcripts can be saved as DOCX, PDF, TXT, or as SRT/VTT subtitles.
It's simple, we built the most accurate audio and video transcription software and API ever Vatis Tech provides a high-speed audio and video to text converter that generates transcripts in over 50 languages with 98%+ accuracy. The platform is designed for efficiency, capable of transcribing one hour of content in just one minute and has an accuracy higher than Google, Speechmatics, Microsoft and other alternatives. It includes transcription software, speech-to-text APIs, caption generators, and audio intelligence. Vatis Tech serves various industries such as contact centers, broadcasting, medical, legal, media, newsrooms, podcasting, education, government, and defense & security.
Voice AI platform for transcription, voice agents, and speech processing Smallest AI is a voice AI platform offering speech-to-text, text-to-speech, speech-to-speech, voice cloning, and real-time voice agent technologies. Its Pulse speech-to-text models provide accurate transcription across 38+ languages, global accents, and dialects with latency as low as 64 milliseconds. The platform also supports speaker diarization, sentiment and emotion recognition, language identification, voice agent orchestration, telephony, knowledge bases, and enterprise deployment.
An audio recorder app for Apple devices with transcription and curation features. Bangin' Audio Recorder is an application designed for Apple devices that allows users to record, transcribe, and curate audio and voice memos. It offers a fast, intuitive interface and private synchronization across devices. Key features include timestamped speech-to-text transcription, map view, editing, and sharing capabilities. The app aims to streamline the process of capturing sound and developing ideas, making it easier to manage and utilize recordings.
Palabra.ai is a real-time AI speech translation platform for video calls, live events, broadcasting and API integrations, supporting 60+ languages with near-zero latency. Palabra.ai is a real-time AI speech translation platform that delivers seamless interpretation for video calls, live events, broadcasting and custom integrations via API. Supporting over 60 languages with near-zero latency, Palabra.ai covers the full pipeline from automatic speech recognition (ASR) and translation to text-to-speech (TTS). Up to 4× cheaper than human interpreters, Palabra.ai delivers professional-grade accuracy with custom glossaries that keep industry-specific and technical terms translated correctly every time. It also features voice cloning to preserve the speaker's natural tone and identity across languages, producing a human-sounding output rather than a robotic one. Palabra.ai is designed to be simple for everyone involved. No downloads or complex setups required, making multilingual communication accessible to anyone, anywhere. Free trial available.
AI-powered subtitle and transcription service with translation for content creators and businesses. SubEasy is a professional AI subtitle and transcription service that automatically generates accurate translations. It supports 100+ languages, offering high-accuracy transcription, automatic translation, and precise subtitle timing. It is suitable for content creators, businesses, and various application scenarios, helping to improve work efficiency.
AI-powered mobile app converting speech to structured text for various uses. Letterly is a mobile app that uses AI technology to convert speech into clear and well-structured text. It goes beyond simple transcription by enabling users to easily rewrite their speech into structured notes, engaging social posts, meeting summaries, formal emails, and more.
AI-powered voice cloning, text-to-speech, and speech-to-text platform. Voicv is a cutting-edge voice cloning platform that transforms your voice into a digital asset in minutes, supporting multiple languages and zero-shot learning. It offers advanced AI-powered voice cloning, text-to-speech (TTS), and speech-to-text (ASR) services. Users can create, transform, and convert audio with cutting-edge technology, supporting multiple languages and emotions.
On-device voice assistant for conversational AI coding SKI is an on-device voice assistant for AI coding agents such as Claude Code, Cursor, Codex, Gemini CLI, Windsurf, and OpenClaw. It lets developers speak naturally to their coding agent, have the agent build and manage code, and hear its responses aloud. Unlike dictation tools, SKI provides a full two-way voice conversation with local speech-to-text, neural voice synthesis, full-duplex interruption, agent status updates, multi-project support, and optional meeting participation. Voice data, transcripts, and local meeting recordings remain on the user's computer and are not uploaded.
AI audio and video processing platform with tools for transcription, translation, and editing. RecCloud is a leading AI audio and video processing platform that offers a range of tools for content creation and editing. It includes features like AI speech-to-text, AI subtitles, AI text-to-speech, and AI video translation. The platform is designed to be user-friendly and accessible online.