Text-to-Speech 24

Readio converts PDFs to audiobooks with a clean and intuitive layout. Readio is designed with a clean and intuitive layout, making it easy for users to convert their PDF files to audiobooks with just a few taps. It allows users to read and listen to all their documents in their native language.
Geleza is an AI-powered platform offering educational and content creation tools for students and creators. Geleza is a comprehensive platform for students, businesses, and creators, offering a unified solution with various AI tools and features. It provides an AI Homework Helper, Interactive PDFs, Math Solutions, Image Creation, Text-to-Speech, Smart Coding, OCR, and Dynamic Question Generation.
AI text-to-speech generator, faster and cheaper ElevenLabs alternative. WavFlow is an AI text-to-speech generator that empowers creators, businesses, and developers to convert text into natural-sounding speech. It offers a faster and cheaper alternative to ElevenLabs, allowing users to transform text into speech with AI in just a few clicks. WavFlow does not require a subscription, and credits do not expire.
Conversational text-to-speech model for natural, expressive dialogue. ChatTTS is a cutting-edge conversational text-to-speech (TTS) model designed for dialogue scenarios such as chatbots and virtual assistants. It transforms text into dynamic, natural-sounding speech, supporting both English and Chinese. The model is trained on extensive data (100,000+ hours for the full version, 40,000 hours for the open-source version) to deliver expressive speech with fine-grained control over prosodic features like laughter, pauses, and interjections.
Automated PA announcements and international name pronunciation for airports, hospitals, and resorts. EasyAnnounce is an automation platform specialized in Public Address (PA) announcements and international name pronunciation. It is designed for environments like airports, hospitals, and resorts where clear communication is vital. The platform uses a purpose-built name pronunciation model and a text-to-speech (TTS) pipeline to generate natural-sounding audio calls in English and other major languages. It offers both a secure web application for manual use and a REST API for developers to integrate accurate name pronunciation and translation into their own voice agent workflows or existing PA systems.
All-in-one audio AI platform for transcription, text-to-speech, dubbing, and captioning. SIREN is an all-in-one audio AI platform designed to provide solutions for audio transcription, audio pen, text-to-speech, video dubbing, and live stream captioning. It leverages cutting-edge GPU-empowered technologies to transform thoughts into text, generate audio from text, and make content understandable internationally.
Platform for building low-latency voice AI agents with ASR, TTS, and LLM models. Hathora Models provides a platform for building voice agents on open-source or closed models with zero DevOps. It offers low-latency ASR (Automatic Speech Recognition), TTS (Text-to-Speech), and LLM (Large Language Model) models that run in 14 regions for ultra-low latency. Users can start instantly on shared endpoints and upgrade to dedicated infrastructure for privacy, compliance, or VPC requirements. The platform allows users to explore, test, and deploy production-ready models, bring their own models or custom containers, and utilize a "Chain tool" for interactive voice AI pipelines.
AI voice cloning tool for instant, realistic, and downloadable audio generation. Voiceley is an AI voice cloning service designed to generate instant, realistic audio quickly. Users can clone their own voice by uploading a clean sample or generate speech using voices from the existing library. The system allows users to type text, generate audio output in seconds, and download the resulting clips for reuse anywhere.
Xiaomi's universal smart platform for multimodal AI, agentic tasks, and voice synthesis. Xiaomi MiMo is a universal smart platform and a suite of advanced large-scale AI models developed by Xiaomi. It is designed to function as a 'New Brain,' bridging the gap between complex algorithms and human intuition. The platform encompasses several specialized models, including MiMo-V2-Pro for top-tier agentic capabilities, MiMo-V2-Omni for multimodal perception (seeing, hearing, and acting), and MiMo-V2-TTS for high-quality speech synthesis. MiMo focuses on the core principles of prediction and compression to understand language, perceive the physical world, and act as a lasting companion in human-machine collaboration.
Voice AI platform for voice morphing, cloning, and content creation. Altered Studio is a Voice AI content creation platform that provides exclusive access to Speech-To-Speech Voice Morphing and integrates various Voice AI technologies into a single user-friendly application for media production. It allows users to change their voice to curated AI voices or custom voices, create professional voice performances, clone voices, clean voice recordings, and utilize text-to-speech features.
Free online flashcard maker with spaced repetition, dictionaries, and text-to-speech. Memozora is a free online flashcard maker that utilizes spaced repetition to optimize learning. It allows users to create their own flashcards, and the platform schedules quizzes based on individual performance. Memozora features multi-language dictionaries (GPT-4 powered) and AI text-to-speech functionality. It works on both PC and mobile devices and supports various quiz formats.
AI-powered flashcard maker and study tool, a Quizlet alternative. NoteKnight is an AI-powered flashcard maker and study tool, positioned as a Quizlet alternative. It offers features like AI-assisted flashcard generation (AutoScribe), image support, explanations, text-to-speech, smart study modes, and mobile-friendly access. It aims to provide a smarter learning experience with a user-friendly UI and powerful AI-enhanced study tools.
MiniMax is an AI company offering text, speech, and video generation models via API. MiniMax is a leading global technology company and one of the pioneers of large language models (LLMs) in Asia. They offer a range of AI models and capabilities, including text, speech, and video generation, through their API platform. Their mission is to build a world where intelligence thrives with everyone.
AI-powered video dubbing and translation service for creating multilingual videos. DubWiz is an AI-powered video dubbing service that allows users to translate and dub videos into multiple languages directly in their browser. It utilizes AI technologies like Speech-to-Text, Neural Machine Translation, and Neural Text-to-Speech to provide a user-friendly experience for creating multilingual videos.
AI models for African dialects, bridging language barriers and enriching digital experiences. Neoform AI provides AI models for African dialects, aiming to make AI opportunities equally accessible. It bridges language barriers and enriches the digital experience for millions by offering voice assistant and translation services, multilingual customer support, localized navigation systems, transcription and captioning, public announcements and info dissemination, and localized content creation.
A better UI for ChatGPT and other AI models with enhanced features and deployment options. TypingMind is a user interface for interacting with AI models like ChatGPT, Gemini, and Claude. It enhances the standard chat experience with features such as chat history search, folders, integrations, a prompt library, and the ability to use your own API keys. It supports various deployment options, including running locally on your browser, as a macOS app, or self-hosting.
AI-powered audio library for audiobooks, podcasts, and blogs. Azalea Labs is an advanced audio platform that utilizes voice AI models to bring public domain books, podcasts, and blogs to life. It serves as a modern digital library, allowing users to listen to a vast collection of stories and knowledge with features optimized for spoken word audio. The platform is currently available on iOS and is expanding its catalog by collaborating with authors and publishers to digitize and distribute their works. It aims to be the world's largest audio library, offering a boundless collection of stories with integrated tools like transcripts and karaoke-style follow-along text.
Classic Microsoft SAM Text-to-Speech voice in your browser. Microsoft SAM Text-to-Speech is a modern JavaScript implementation of the iconic voice synthesizer from Windows XP, originally part of the Microsoft Speech API (SAPI). This website brings the classic Microsoft SAM voice directly to your browser, allowing users to generate speech with its distinctive robotic voice without any downloads or server processing. It aims to preserve the authentic nostalgic charm of the original while adding modern conveniences like browser-based functionality and customizable parameters.
AI platform to create talking e-cards and videos from photos. Virbo is an AI-powered platform that allows users to create talking e-cards and videos from photos. It transforms portraits into dynamic talking avatars with natural human voices in over 100 languages. Virbo offers AI spokesperson video generation online, supporting features like talking photos, URL to video conversion, PPT to video conversion, video translation, AI video generation, AI montage maker, and AI clip generation.
AI tool to animate photos with speech and lifelike expressions. AI Talking Photo Generator is a technology that uses artificial intelligence to animate still photos, making them appear to speak naturally. It transforms photos into lifelike talking animations by analyzing facial features and creating realistic lip movements and facial expressions that synchronize with audio input, delivering natural and expressive results.
Free online AI tool for voice and language transformation. Voice Changer is a free online AI tool that transforms voices using artificial intelligence technology. It offers a rich library of over 100 AI voices and supports more than 20 different languages, allowing users to easily change their voice or language. It is perfect for creating engaging multilingual audio content, providing natural and realistic voice effects for various applications such as content creation, localization, education, animation, marketing, and development.
Deepgram is a Voice AI platform offering STT, TTS, and voice agent APIs for developers. Deepgram is a Voice AI platform that provides APIs for speech-to-text, text-to-speech, and voice agent functionalities. It enables developers to build voice AI products and features with real-time, accurate, and scalable solutions. Deepgram's platform is trusted by top enterprises and startups for various use cases, including contact centers, medical transcription, and conversational AI.
A platform for deploying and running machine learning models with a simple API and pay-per-use pricing. Deep Infra offers cost-effective, scalable, easy-to-deploy, and production-ready machine-learning models and infrastructures for deep-learning models. It provides a platform to run top AI models using a simple API, with pay-per-use pricing and low-latency inference. Users can deploy custom LLMs on dedicated GPUs and access various models for text generation, text-to-speech, text-to-image, and automatic speech recognition.
Groq offers fast AI inference through its hardware and software platform for AI applications. Groq is a hardware and software platform that delivers exceptional compute speed, quality, and energy efficiency for AI inference. Groq provides cloud and on-prem solutions at scale for AI applications, offering high-performance AI models and API access for developers. It aims to provide faster inference at a lower cost than competitors.