AI Speech-to-Text 350

Voice-first AI macro tracking app for serious fitness enthusiasts. TrakMac is a voice-first macro tracking mobile application designed specifically for serious fitness enthusiasts and strength athletes. Unlike traditional nutrition apps that require manual database searches or barcode scanning, TrakMac allows users to log their meals by simply describing what they ate in plain spoken language. Powered by AI, the app estimates calories, protein, carbohydrates, and fat within seconds, adjusting targets based on the user's specific training profile rather than generic formulas.
AI-powered online subtitles editor for social media videos. Subtitles.Love is an AI-powered online subtitles editor that allows users to add subtitles to their social media videos to increase audience interaction. It offers features like automatic speech recognition, resizing, and styling for various social media platforms. The platform supports multiple languages and video formats, aiming to simplify and speed up the process of creating subtitled videos.
macOS app converting speech to text with ChatGPT, speeding up writing. WhisperWizard is a macOS application that transforms spoken words into written text with the help of ChatGPT. It speeds up writing workflows by allowing users to speak instead of type, capturing ideas instantly and accessing old recordings. It also offers custom ChatGPT prompts to edit recordings and create templates for routine tasks.
Pay-as-you-go audio/video transcription service with AI content generation features. Transcriptmate is an online audio and video transcription service that offers 'Pay-As-You-Go' transcription without requiring registrations, subscriptions, or monthly commitments. Users pay per file, not per minute, and receive high-quality transcriptions along with additional features like summaries, articles, and social media posts generated from their audio/video content. It supports files up to 3 hours long and delivers transcriptions in csv, srt, and txt formats via email within 2 hours.
Multilingual Speech-to-Text API with high accuracy in 14 languages. SpeechFlow is a multilingual Speech-to-Text API that offers state-of-the-art accuracy in 14 languages. It converts sound to text, speech to text, and audio to text with high accuracy. SpeechFlow supports both cloud and on-prem deployment.
Platform for on-device speech AI, enabling speech recognition and wake word detection. Wavify is a one-stop-shop for voice AI, providing a platform for on-device speech AI. Software engineers can embed features like speech recognition and wake word detection into any software. It offers SOTA models and a cross-platform inference engine, optimized for speed and privacy. Wavify supports multiple languages and runs on various platforms, including Linux, Mac, Windows, iOS, Android, Web, Raspberry Pi, and embedded systems.
Unifies speech recognition across 1,600+ languages using AI and LLM-enhanced decoders. Omnilingual ASR is an advanced automatic speech recognition technology that unifies speech recognition across a vast number of languages, scaling from dozens to over 1,600 natively and extending to 5,000+ via few-shot prompts. It achieves this by combining wav2vec-style self-supervision, LLM-enhanced decoders, and balanced multilingual corpora to learn language-agnostic acoustic patterns. This website serves as a comprehensive knowledge base, detailing its research breakthroughs, current technologies, datasets, implementation strategies, and deployment guidance for achieving omnilingual reach in a single model.
AI voice input agent that turns speech into polished, structured text 网易叭哥说 is an AI-native desktop voice input agent developed by NetEase Youdao. It converts spoken language into clear, accurate, and structured text rather than merely transcribing speech. The tool understands conversational intent, removes filler words, corrects self-revisions, adds punctuation, organizes paragraphs, supports voice translation, and enables voice input across applications. It is available as a free download for Mac and Windows.
AI medical scribe that converts patient conversations into clinical notes, saving time and reducing burnout. Sunoh.ai is an AI medical scribe designed to save physicians time and reduce burnout. It listens to patient-provider conversations and converts them into clinical notes, integrating with EHR systems to streamline documentation. Trusted by over 80,000 physicians, Sunoh.ai aims to make clinical documentation faster, more accurate, and more efficient.
AI-powered voice notes app for creating content from audio in 90+ languages. NoteGen is an AI-powered voice notes app that instantly turns your ideas into effective content. It supports 90+ languages and allows you to effortlessly create journals, notes, scripts, posts, call summaries, and more in one click. You can record or upload audio for note-taking, call summarizing, journaling, creating posts, content scripts, and more.
AI platform to summarize, transcribe, and convert audio to notes. OneAudio is an AI-powered platform that summarizes, transcribes, and converts audio into clean, structured notes. It allows users to share, edit, bookmark, and manage their original audios, transcripts, and summaries. The platform is used for creating notes, emails, articles, messages, and more.
AI-powered voice memos app for iOS that transcribes and summarizes recordings. Scribe Notes is an AI-powered voice memos app for iOS that transcribes and summarizes voice recordings. It uses Whisper and GPT-4o to convert spoken ideas into organized notes, which can be shared or received as email summaries. The app offers both free and premium features, including unlimited notes, longer recording times, and custom instructions for AI summaries.
Veterinary AI scribe for automated medical record creation. ScribVet is a Veterinary AI Scribe that helps veterinarians create detailed medical records by recording themselves during exams. It simplifies practice management by automatically generating SOAP notes, client communications, and other documents, saving time and improving work-life balance.
AI note-taking assistant for generating and organizing notes from various sources. EasyNoteAI is a powerful AI note-taking assistant that helps users efficiently organize their notes. It generates notes from audio, online videos, and PDFs, and automatically creates AI note outlines, summaries, and Q&A from the notes. It supports hundreds of languages and offers features like live transcription, flashcard generation, quizzes, mind maps, and summary reports.
AI-powered meeting note-taking and transcription tool. Minutes AI automates meeting audio notes by instantly creating formatted notes and transcriptions from live audio, uploaded audio files, or imported YouTube links. Users can chat with their audio to extract key insights, list action items, and more. It's designed to be reliable, simple, private, and powerful, helping users never take notes manually again.
AI text-to-speech and voice cloning platform with 600+ voices in 142 languages. Verbatik is an AI-powered text-to-speech and voice cloning platform that converts written text into natural-sounding speech. It offers over 600 realistic voices across 142 languages and accents. Verbatik allows users to clone voices and customize audio for marketing and more. It generates natural voices in 100+ languages, perfect for videos, podcasts, and e-learning. The platform also provides tools for script writing, avatar AI, and a sound studio for enhancing audio projects.
AI Discord bot for voice transcription, meeting notes & DnD. DiscMeet is a Discord bot designed to transform voice calls into actionable meeting notes for both professional teams and Dungeons & Dragons (DnD) campaigns. It offers real-time AI transcription in over 100 languages, including speaker identification and timestamps. The bot provides intelligent organization tools such as automatic conversation threading, team insights with analytics on communication patterns, and comprehensive campaign management features for DnD, including character tracking and AI-powered session summaries. DiscMeet aims to enhance productivity and collaboration within Discord communities by making voice conversations more productive and organized.
AI-powered audio transcription tool for easy note management. Dictaphone is an AI-powered tool that allows users to easily transcribe audio files and manage them as notes. It supports live transcriptions and audio file uploads, providing accurate results in seconds using AI.
Extracts and displays YouTube video transcripts for easy access and review. YouTube Transcript Generator extracts and displays the complete transcripts from any YouTube video. It allows you to quickly access, read, and save video content without watching the entire video, making it easier to find specific information or review content at your own pace.
Free web-based tool to transcribe and summarize MP3 files using AI. WebWhisper is a FREE web-based alternative for MacWhisper that allows you to transcribe and summarize MP3 files effortlessly. It utilizes advanced AI models like GPT-3.5, GPT-4, and Claude to get accurate transcriptions and concise summaries.
AI-powered tool that converts MP3 audio files into accurate text. MP3 to Text is an advanced AI-powered online transcription tool that converts MP3 audio files and other audio formats into accurate written text. Supporting over 90 languages and regional dialects, it offers fast processing, speaker recognition, and multi-format export options, making it an efficient solution for transforming spoken words into accessible documents without requiring software installation.
Automatic transcription software converting audio and video to text with high accuracy and multi-language support. Audiotype is an automatic transcription software that converts video and audio files into editable text transcripts. It supports 36+ languages and boasts 80-95% accuracy. It offers features like MP4 to text, MP3 to text, WAV to text conversion, audio and video transcription, YouTube video transcription, podcast transcription, interview transcription, call recording transcription, meeting recording transcription, Zoom meeting transcription, Microsoft Teams transcription, SRT subtitles, VTT subtitles, and closed captions. It also provides transcription software for journalists, students, businesses, and developers, along with API access, a live demo, and a free trial.
AI-powered voice collaboration platform for transcription, summarization, and actionable insights. Vocol is an all-in-one voice collaboration platform powered by AI, designed to boost work efficiency by turning voice and data into actionable insights. It transcribes and summarizes meetings, supports multilingual transcription (Chinese, Japanese, and English), and integrates with tools like Teams.
Audio & video transcription tool with AI meeting summary and action items generation. Mictoo is an audio and video transcription tool that converts audio to text automatically. It allows users to record audio or upload files to get real-time transcription. Mictoo also uses GPT Open AI to generate meeting summaries, action items, and follow-ups that can be shared with colleagues. It helps users take meeting notes easier, freeing up their minds to engage positively in meetings and enhance productivity.