Voice Generation & Conversion 3251
All Categories
AI Celebrity Voice Generator 34
AI Dubbing 132
AI Podcast 136
AI Podcast Clip Generator 25
AI Podcast Editing 18
AI Recording 52
AI Speech Recognition 156
AI Speech Synthesis 110
AI Speech-to-Text 388
AI Text-to-Speech 415
AI Transcriber 179
AI Transcription 435
AI Voice Assistants 200
AI Voice Changer 58
AI Voice Cloning 227
AI Voice Enhancer 30
AI Voice Generator 368
AI Voice Over 150
Audio To Text AI 131
Tiktok AI Voice Generator 7

Free online text-to-speech tool with 200+ voices and 70+ languages.
Luvvoice is a free online text-to-speech (TTS) tool that turns your text into natural-sounding speech. It offers speech synthesis services and supports multiple languages, with over 200 voices and 70 languages available. Users can convert text to speech online without word limits, listen online, and download files in MP3 format. It also supports file to speech conversion from PDF and TXT formats.

Text-to-speech tool that synthesizes natural speech from short voice samples.
Fish Speech is a text-to-speech (TTS) tool developed by the creators of So-VITS-SVC and Bert-VITS2. It can synthesize natural and fluent speech from just 15 seconds of any voice, maintaining the given timbre, style, and accent. Fish Audio is a platform for audio generation, offering various voice models for users to discover and use.

Text-to-speech solution with AI voices for personal, commercial, and educational purposes.
NaturalReader is a text-to-speech solution designed for personal, commercial, and educational use. It offers a free online platform, mobile apps, and commercial licenses, utilizing AI voices to read text aloud. It supports multiple languages and provides features like voice cloning and content awareness to enhance the listening experience.

Text-to-speech app for listening to digital content on any device.
Speechify is a leading text-to-speech app available on Chrome, iOS, Android, and Mac. It allows users to listen to documents, articles, PDFs, emails, and more. Speechify offers AI voice cloning, AI dubbing, and AI video generation. It is used by millions to hear the internet on any device.

Gladia is a production-ready Speech-to-Text API for teams shipping voice products—high accuracy, multilingual, real-time + async, and add-ons.
Gladia is a speech-to-text platform built for production, turning raw audio into structured outputs that power real workflows like meeting summaries, CRM enrichment, contact center QA, and real-time voice assistants. With support for 100+ languages and the ability to handle messy real-world audio—overlapping speakers, accents, code-switching, domain-specific terminology—Gladia is designed for the complexity of actual conversations, not clean studio recordings.

Free online video recording, editing, and AI-powered multimedia service platform.
RecCloud is a free multifunctional online application dedicated to providing users with comprehensive video recording and editing services. It offers AI tools including Chatvideo, AI speech-to-text, and AI subtitles. RecCloud is an AI video creation platform, offering free multimedia solutions such as AI video chat, AI subtitles, AI speech-to-text, online screen recording, video editing, storage, and sharing.

Converts audio/video to text, summaries, and insights quickly and accurately.
Transcript LOL is a service that converts audio and video into text, summaries, and more in seconds. It offers features like speaker recognition, high accuracy, and the ability to download in multiple formats. It's used for course content, extracting key points from meetings or interviews, and creating social media posts.

AI assistant for audio/video transcription and summaries
Tongyi Tingwu is an Alibaba Cloud AI assistant for work and study that helps users transcribe, organize, translate, and summarize audio and video content.

Automated transcription, translation, and subtitling platform for audio/video.
Sonix is an advanced automated transcription, translation, and subtitling platform that converts audio and video files to text quickly, accurately, and affordably. It leverages industry-leading speech-to-text AI algorithms to transcribe various content types like podcasts, interviews, speeches, meetings, and films. Beyond transcription, Sonix offers automated translation, AI analysis tools (summaries, topic detection), automated subtitling, and features for sharing, collaboration, organization, and integration with popular workflows.
Distributed GPU cloud offering compute, storage, and deployment solutions at lower costs.
SaladCloud offers distributed GPU cloud services, including compute, storage, and deployment solutions. It provides access to a network of consumer GPUs for tasks like image generation, voice AI, computer vision, data collection, batch processing, and molecular dynamics. The platform aims to democratize cloud computing by offering lower-cost alternatives to traditional cloud providers, particularly for AI transcription and GPU-intensive workloads.

AI transcription service for audio and video to text conversion with high accuracy.
Transkriptor is an AI-powered transcription service that converts audio and video files into text with high accuracy. It offers features like meeting recording, translation, subtitle generation, and AI-driven summarization, making it suitable for various use cases, including business meetings, academic research, and content creation.

Free AI transcription tool for audio, video, and conversations, supporting 36+ languages.
Deepgram offers a free transcription tool that converts conversations, audio files, or YouTube videos into text. It supports over 36 languages and dialects, providing accurate and reliable transcripts for students, journalists, podcasters, and professionals. The tool is designed to be simple and efficient, offering a seamless transcription experience without ads or costs. It also provides a Text to Voice API for creating natural-sounding voiceovers.

AI-powered offline voice-to-text app for macOS, supporting 100+ languages.
superwhisper is an AI-powered voice-to-text application for macOS that allows users to dictate emails, send messages, and take notes at speeds up to three times faster than typing. It operates completely offline, ensuring privacy and security as data never leaves the user's device. superwhisper supports over 100 languages and offers features like literal punctuation control in its Pro version.

AssemblyAI: AI models for speech-to-text transcription and voice data insights.
AssemblyAI provides State-of-the-Art AI models for automatic speech recognition (ASR), natural language processing (NLP), and AI speech-to-text. It enables users to transcribe speech to text and extract insights from voice data. The platform offers speech-to-text, streaming speech-to-text, and speech understanding capabilities, catering to startups and enterprises for reliable source-truth data that powers world-class products.

A high-speed video converter, compressor, and editor with AI-enhanced features.
Wondershare UniConverter 16 enables you to experience an ultra-high-speed video converter and compressor, designed to process 4K/8K HDR files. It delivers a high-speed video converter and compressor boasting 20+ functions, making it ideal for handling 4K, 8K, and HDR files for all video enthusiasts and educators. It also offers AI-enhanced features like speech-to-text, video enhancement, and background removal.

UniScribe is an AI-powered platform for audio and video transcription, summarization, and mind map generation.
UniScribe is a platform for transcribing videos and audios. It converts media files to text with high accuracy in multiple languages. It also creates summaries, mind maps, and key questions, and lets you export the text in different formats. UniScribe lets you upload audio and video files or paste YouTube Links, quickly turning them into text with AI.

Deepgram is a Voice AI platform offering STT, TTS, and voice agent APIs for developers.
Deepgram is a Voice AI platform that provides APIs for speech-to-text, text-to-speech, and voice agent functionalities. It enables developers to build voice AI products and features with real-time, accurate, and scalable solutions. Deepgram's platform is trusted by top enterprises and startups for various use cases, including contact centers, medical transcription, and conversational AI.

Browser-based private AI speech-to-text transcription
Whisper Web is a browser-based AI speech recognition tool powered by OpenAI Whisper. It transcribes audio in 100+ languages locally in your browser using WebGPU and WebAssembly, so no data leaves your device in Free mode. It also offers an Unlimited cloud plan for longer files and batch uploads.

AI-powered media management assistant with transcription, video editing, and asset management tools.
Clipto.AI is an AI-powered media management assistant that offers tools for transcription, video editing, and digital asset management. It provides accurate AI transcription in multiple languages, a YouTube downloader, and features like smart asset search and light video cutting. The platform emphasizes privacy and security by running AI algorithms on-device, ensuring data never leaves the user's computer unless explicitly chosen.

AI-powered transcription and meeting minutes service with real-time transcription and translation.
Notta is a high-precision transcription service equipped with the latest AI speech recognition engine. It features real-time transcription and translation, and can quickly transcribe audio files up to 5 hours long at a time. It allows for easy audio conversion and editing on PC.

AI transcription service converting audio and video to text in 98+ languages.
TurboScribe is an AI transcription service that converts audio and video files to accurate text in 98+ languages. It offers unlimited transcription with near-perfect accuracy, exporting in various formats like PDF, DOCX, SRT, and TXT. TurboScribe is powered by Whisper and provides features like speaker recognition and built-in translation.

Rev is a voice platform for transcription, captions, and subtitles using AI and human services.
Rev is a voice platform that provides speech-to-text services, including AI and human transcription, captions, and subtitles. It caters to various industries, offering solutions for legal, research, healthcare, newsrooms, education, and financial services. Rev emphasizes accuracy, security, and tailored summaries, leveraging AI-powered tools and expert human transcribers to deliver high-quality transcripts and insights.

All-in-one marketing and design platform with templates and AI tools.
PosterMyWall is an all-in-one marketing and design platform designed for busy individuals, small businesses, bands, churches, and restaurants. It provides users with over 3 million free design and email templates to easily customize flyers, posters, videos, menus, and social media posts. The platform features powerful AI-driven tools, such as an AI design generator, AI background remover, AI subtitles, AI voiceover generator, AI image generator, and an AI writer for captions and marketing text. Additionally, PosterMyWall integrates marketing tools like a Content Planner, social media scheduling across multiple platforms, trackable email campaigns, and shareable event pages to help users promote their business and events seamlessly.

AI-powered IP phone system for voice conversation analytics and sales performance improvement.
MiiTel is an AI-powered IP phone system designed to enhance voice conversation analytics. It functions as a “Smart PBX,” aiming to increase sales, decrease onboarding time, and enable remote work. The platform offers features such as automatic transcription, conversation and data analysis, CRM integration, and real-time monitoring and coaching.
Hot Articles
How to Build a Multimodal AI Knowledge Base With Gemini Embedding 2
NASA 和 IBM 开源月球模型
Google’s Gemini is the latest AI model to hack other companies
Introducing the Australian Youth Safety Blueprint
GLM-5.3-FlashX 上线,智谱把国产卡上的推理速度顶到 200 token/s
How to Choose Hardware for Running Local LLMs, and Know Exactly When It Beats the Claude API
Latest Articles
Google’s Gemini is the latest AI model to hack other companies
How to Choose Hardware for Running Local LLMs, and Know Exactly When It Beats the Claude API
How to Build a Multimodal AI Knowledge Base With Gemini Embedding 2
Announcing Grok-1.5
Google DeepMind launches institute to widen the AGI debate
Google’s new ‘CC’ is an AI agent that helps families run their households
Hot Tags