Eleven Labs generates lifelike speech, voice clones, sound effects and dubbing in over 70 languages, plus tools for transcription and voice-based apps, useful for creators and podcasters.
AI tools/Voice, speech & music
266 tools · 39 free
Voice, speech & music
Synthetic voices, transcription and music.
Adobe Speech Enhancer cleans up recorded speech by reducing noise and balancing sound quality, useful for podcasters and video creators who recorded without professional microphones.
Vocal Remover strips the singer's voice out of a song to create a free karaoke version in seconds, useful for singers or remix creators.
Udio
★Udio lets you create, share and discover music tracks, writing your own lyrics or generating instrumental-only pieces, great for aspiring musicians and music hobbyists.
Descript
★Descript is an audio/video editor that lets you clone your own voice and use it for high-quality text-to-speech. It helps podcasters and video creators make quick edits to real recordings.
Krisp
★Krisp uses AI to remove background noise and echo from calls, and provides call summaries and talk-time insights afterward. It's useful for anyone taking online meetings in noisy environments.
Wispr Flow instantly converts speech into well-formatted text, letting you write much faster than typing. It's useful for people with disabilities or anyone wanting a faster way to write.
Revocalize AI converts any recorded audio into a vocal track that sounds like a chosen singer, with no singing skill needed. It's currently in private beta.
Verbatik converts text into natural-sounding speech in 142 languages with over 600 voices, including emotion controls. It suits marketing, education, and customer service use.
Voice.ai changes your voice to sound like well-known figures in real time, working across games, meetings, and apps like Discord and Zoom. It suits gamers and streamers.
MyVocal.ai clones your voice in under a minute, then lets you use it for text-to-speech or even singing. A free tier is available for people who want to try it out.
MusicLM is a Google Research model that generates high-fidelity music from text descriptions, staying coherent over several minutes. It can also shape the output from a whistled or hummed melody.
Cleanvoice AI removes filler sounds, stutters, mouth noises, and other artifacts from audio recordings, supporting multiple languages and accents. It suits podcasters who want cleaner audio.
Resemble.ai generates lifelike AI voices, offering text-to-speech, your own voice cloning, emotion control, and cross-language dubbing. Developers can integrate it into their own apps via an API.
Speech Studio is a toolset for adding text-to-speech and speech recognition features from Azure into applications, with a no-code project setup.
Synthesizer V is a music tool that generates realistic AI singing voices using neural synthesis, with customizable pitch, tone, and language options.
Riffusion is an AI music generation platform that creates tracks from text prompts, currently in beta.
Musicfy lets users create and discover AI-generated song covers, clone their own voice, and remix tracks using a built-in music library.
Voicemod is a real-time voice-changing app with over 90 voices and effects, designed for streaming, gaming, and video chat.
ai-coustics enhances speech in real time before it reaches voice AI systems, reducing noise and errors to improve clarity for voice applications.
Voicemaker converts text into natural-sounding speech in multiple languages, with adjustable speed, pitch, and emphasis, suited for audiobooks and podcasts.
Soundful is an AI music generator that creates royalty-free tracks based on chosen genres and customizable settings.
Playcast.ai converts articles, books, or PDFs into audio you can listen to like a private podcast. It's useful for busy people, commuters, or those with visual impairments who prefer listening over reading.
Soundraw generates royalty-free instrumental music based on the genre, mood, tempo, and length you choose.
Audiobox generates voices and sound effects from text prompts or voice input, letting you build interactive audio stories.
FakeYou generates text-to-speech audio clips mimicking different characters in multiple languages using deepfake technology.
Dubbing AI changes or clones your voice in under a second, popular with gamers and streamers for games and content.
PlaylistAI builds Spotify and Apple Music playlists from a prompt, an image, a video, or your most-played songs.
Listnr converts text into realistic voiceovers in over 140 languages, used for ads, e-learning and creating podcasts.
Boomy lets you compose and release original songs in seconds, and earn a share of royalties when they're streamed.
Uberduck turns text into speech using thousands of AI voices, and lets you clone your own voice. It suits podcasters and developers building voice apps.
Revoicer turns text into realistic voiceovers in many languages and emotional tones, useful for sales videos, lessons, and podcasts.
Koe Recast transforms your voice into different styles such as narrator, female, or anime.
Mubert generates AI music based on mood, genre, and duration, and also offers developer tools to add adaptive music into apps and games.
Samplette searches YouTube for music samples matching a chosen style or BPM, though some clips may still carry copyright.
Beepbooply converts text into realistic speech in over 80 languages with more than 900 voices, offering both free and paid options.
Magenta Studio is a set of AI-powered music plugins that generate audio, available as standalone apps and Ableton Live add-ons.
AudioBot converts text into speech in multiple languages with over 500 voice options, useful for videos, presentations, and radio shows.
Narration Box generates expressive AI voiceovers in over 140 languages, useful for e-learning, ads, audiobooks and customer support. It lets anyone produce studio-quality narration without recording equipment.
Vocol AI transcribes conversations into accurate text and highlights key insights from meetings. It's useful for teams wanting to capture and share important discussion points efficiently.
Tracksy generates original, royalty-free music tracks in various styles that users can preview before use. It's suited for anyone needing background music without prior music-making experience.
Cassette AI generates original, royalty-free music tracks from a text description of the genre, mood and length you want. It's for anyone creating music without professional experience.
AnthemScore converts audio files into sheet music or guitar tabs, with automatic note detection and easy correction tools. It's for musicians who want to transcribe recordings into written notation.
SpeechEasy converts text into high-quality audio with a choice of natural-sounding voices, working across desktop and mobile. It suits people who prefer listening to content and values user privacy.
Doctor Mix AI Synth is software for the AX73 synthesizer built using ChatGPT-generated code. It's aimed at musicians curious about trying new AI-assisted digital instruments.
A.V. Mapping helps filmmakers and musicians find suitable music for their projects by analyzing video, text and audio with AI. It also offers noise reduction and sound effect design tools.
Sumly.AI delivers concise, human-reviewed summaries of podcasts and videos to your inbox within 24 hours. It helps busy people stay updated on their favorite shows without listening to full episodes.
WhatTheBeat uses AI to explain the meaning behind songs from popular artists like Drake and Taylor Swift. It's for music fans who want deeper insight into their favorite lyrics.
Drumloop AI generates original drum loops based on a chosen genre, which users can tweak and download, useful for musicians looking for quick beat ideas.
Scribewave AI transcribes and captions audio and video files in over 90 languages with high accuracy, useful for students, video creators and legal professionals who need quick written records.
Orb Producer Suite is a set of four AI-powered music plugins that generate unlimited chord progressions, melodies, basslines, and arpeggios, compatible with most DAWs.
GetSound generates AI-created ambient soundscapes for spas, hotels, and resorts, tailored to location and current weather conditions.
Rythmex converts audio and video into text online, offering 30 free minutes for users like journalists, students, and podcasters.
MusicStar.AI offers lyric writing, vocal recording, voice changing, and album cover design in one place for creating full music tracks. It suits musicians and creators who want to produce songs without separate tools.
HearTheWeb converts text into podcast episodes with AI co-hosts in just a few minutes. Users can pick from many host voices and customize the conversation style and branding.
Audie AI converts written books into audiobooks using natural-sounding text-to-speech technology. Authors and publishers can get a finished audiobook ready in 24 hours or less.
SwearAway automatically detects and mutes profanity in audio files, letting users review and edit the transcript before export. It helps podcasters and educators keep their content clean and appropriate.
Riverside Transcriptions converts audio and video files into accurate text in over 100 languages, free and without signup. Users can download transcripts or subtitle files for their own use.
VOCALOID6 is Yamaha's technology for adding lyrics and singing vocals to music compositions. It offers multiple voicebanks and works alongside other music production software.
VoicePen turns audio or video recordings into transcripts and blog posts, helping podcasters, webinar hosts and educators convert spoken content into text in many languages.
Podsift emails you AI-generated summaries of your favorite podcasts whenever a new episode is released, at no cost.
AudioShake uses AI to separate a song into its individual components like drums or vocals, helping create remixes, instrumentals or remaster old recordings.
Moises App lets musicians remove vocals or specific instruments from a song, change tempo and pitch, and detect chords in real time.
RipX pulls a recording apart so you can drop the vocal, lift single instruments, rearrange a track or clean up a flawed take. A 21-day free trial is offered.
Transkribieren quickly and accurately transcribes audio files like mp3 or wav into text, up to a 25MB file size limit.
Voicify lets you create song covers in the voice styles of famous artists like Drake or Kanye West in seconds.
Speechify reads text aloud in a natural voice so you retain more of what you read, up to 9x normal speed, and can even read a photo of a page.
Muzaic.studio composes a soundtrack for your video based on the mood, tempo, and intensity you want, saving you the hassle of hunting for royalty-free music.
This is a plugin for music producers that turns written lyrics into realistic-sounding vocal performances, offering a choice of four different voices.
This is a platform for creating, discovering, and sharing user-made songs, with rankings of popular and trending tracks for music lovers.
This online platform helps musicians separate track stems, create remixes and mashups, and extract key, chords, and midi from songs.
This platform helps people produce studio-quality podcasts quickly using voice cloning, AI-generated scripts, and a library of pre-licensed music.
Deciphr AI summarizes podcast transcripts and generates timestamps automatically. It helps podcasters save time when preparing show notes for their episodes.
WellSaid Labs converts text into natural-sounding voiceovers using AI voice avatars. It helps businesses quickly produce audio for audiobooks, marketing, and customer support.
TuneBlades uses AI to lengthen or shorten songs while preserving the melody and vocals, exporting the result as mp3, wav, or m4a files. It helps musicians quickly adjust track duration for different uses.
Acoust converts text to speech in over 30 languages using more than 100 natural-sounding voices. It suits people who want to create voiceovers or listen to articles instead of reading them.
Loudly generates original, royalty-free music quickly using AI, giving creators freedom to use tracks without copyright concerns. It suits video makers, artists, and small business owners.
Voice-Swap transforms a vocal recording to sound like a well-known singer, helping music producers collaborate remotely without studio time. Commercial use requires permission from the original artists.
EchoFox is a 24/7 transcription assistant that turns audio into accurate text, handling up to 98 languages and multiple speakers.
Vocalist.ai turns ordinary vocal recordings into studio-quality singing and rapping using licensed AI vocal models.
Voxify generates natural-sounding AI voiceovers with emotion in 140+ languages and accents, letting you adjust tone and pacing.
Altered transforms your recorded voice into a different voice for voiceover work. Content creators can adopt other voices or build multi-character performances.
Emergent Drums generates endless royalty-free drum sounds. Music producers get fresh drum samples quickly without worrying about copyright issues.
ScoreCloud turns a played or sung tune into written sheet music. Musicians and teachers can notate a melody just by singing or whistling it into their phone.
Audioatlas is a tool that helps you find the right song from a global database of over 200 million tracks.
Stable Audio generates high-quality music just from a text description, allowing commercial use on its paid plans.
AudioPen turns your spoken words into clean, summarized text by removing filler and repetition, useful for notes, messages, or posts.
Dadabots is a project that uses deep learning AI to generate and stream math rock and black metal music continuously, 24/7.
MusicFX lets users sign in with a Google account to create music and various sounds with calming or adventurous vibes.
LALAL.AI separates vocals and instruments out of a single song. Music producers can pull out just the vocal track or individual instrument stems in high quality.
ACE Studio generates synthetic singing voices with AI instead of hiring a singer. Music producers and game developers can get a vocal track in different styles quickly.
DeepZen converts text into emotionally rich, lifelike speech for audiobooks, ads, and games. Publishers and content producers can get a voice that's hard to tell apart from a human narrator.
Gotalk.ai converts written text into natural-sounding speech for YouTube videos, podcasts, and phone greetings. Users can adjust tone, pitch, and pace to fit what they need.
Podcastle offers studio-quality recording and AI editing for podcasts, including transcription and noise reduction. Podcasters can record with remote guests and edit quickly.
Speech To Note converts spoken words into written text in real time. It helps people capture lectures, meetings, or personal notes without typing by hand.
Cyanite.ai listens to millions of songs and automatically tags them, making it fast to find the right music for any use case.
Peech is an iOS app that converts written text into spoken audio, helping people with reading difficulties or anyone who prefers listening to articles, books and other content in multiple languages.
Replica Studios offers a library of 40+ AI voices for games, films and other creative projects, mimicking real voice actors' delivery quickly and affordably.
Dolby On is an app for recording and livestreaming audio and video with Dolby-quality sound straight from your phone. Musicians and podcasters can reduce noise and enhance sound before sharing it.
An app that turns written descriptions into original songs for videos, podcasts, or music projects. Useful for creators who want a soundtrack without hiring a musician.
Sofiya converts text into natural-sounding speech in over 135 languages, with a studio for enhancing audio. It suits developers of chatbots, voice assistants, and educational content who need quality speech from text.
eMastered is an online AI mastering engine that improves the sound quality of music tracks, using reference matching and advanced settings. It suits musicians who want to master their songs without going to a studio.
SteosVoice creates high-quality synthetic voices for content, video dubbing, and podcasts, while letting users license their own voice for income. It suits content creators who want natural voices and a way to earn from them.
Vid2Txt is a Mac and Windows app that converts video and audio into text, srt, and vtt files offline. It suits journalists, students, and hearing-impaired people who need quick transcripts.
VoicV is a cloud platform for managing voice-over and dubbing work on media projects, with script handling and syncing features. It suits voice professionals and video creators who want better collaboration and faster production.
Beatoven.ai composes royalty-free, mood-matched music for videos and podcasts based on genre and mood choices. It suits video and podcast creators who need background music without copyright concerns.
Deepgram is an API service that converts speech to text quickly and accurately, used in customer service, healthcare, and conversation analytics. It helps businesses extract useful insights from voice data.
EVITA is an AI singing coach that generates personalized vocal exercises, analyzes songs and characters, and helps write new lyrics. It suits singers and performers looking to develop their voice.
Vozart AI turns text prompts into fully produced songs in minutes without needing musical skills. It suits content creators, songwriters and educators who need custom music quickly with usage rights included.
Castmagic automatically turns podcast audio files into transcripts, show notes and social content, cutting down post-production work. It suits podcast creators wanting to save time.
SoundVerse is an AI music-creation platform that turns a user's ideas into finished tracks and allows collaboration in an online studio. It suits musicians and producers of any level.
TwoShot offers a large sample library plus AI-generated sounds created from simple descriptions, integrating directly into music production software. It suits musicians looking for samples quickly.
VoiceLine is a voice-messaging app with automatic transcription that speeds up communication and reduces the need for long meetings. It suits teams looking to save work hours.
An AI tool that generates copyright-free music, up to ten minutes at a time. It suits content creators who need background music for videos or ads without licensing costs.
A tool that turns text, images, or video into quality audio across different music styles. Good for content creators who want sound without musical training.
A web service that automatically cleans up podcast and video audio, balancing levels and cutting noise. Good for podcasters who want clean sound without audio engineering skills.
A tool for creating podcasts, audiobooks, and other audio content with AI voices, from idea to finished file. Good for creators who want audio content without a studio.
A tool that turns your recorded voice into well-organized written content, guiding you with prompts along the way. Good for writers who think better out loud than on the page.
YuE is an open-source AI model that turns song lyrics into complete tracks with vocals and instruments, supporting multiple languages and styles. It suits songwriters who want to quickly produce music.
A DAW plugin that finds sounds musically compatible with your project, from your own library and an online one. Good for music producers wanting fresh sounds without endless searching.
A text editor you control entirely by voice, letting you write and edit documents without touching a keyboard. Good for people with hand mobility limits or anyone multitasking hands-free.
A service that delivers podcast episode summaries to your inbox, in both text and audio form. Good for busy people who want to keep up with favorite shows without listening to every episode.
iListen turns articles or web pages into short audio podcasts you can listen to instead of reading. It's for people who want to learn on the go.
Letterly turns your spoken words into clean, well-structured text and can rewrite it in different styles, in many languages. It's for anyone who prefers talking over typing.
Podsqueeze generates show notes, timestamps, tweets, and blog posts from your podcast recording with one click. It saves podcasters time on promotion work.
StockmusicGPT generates original, royalty-free music for videos, podcasts, or social content based on the genre and mood you pick. It's for content creators who need a soundtrack.
Epidemic Sound's Soundmatch analyzes your video and recommends matching music, then lets you tweak the track's mood and instrumentation. It's for video creators looking for the right soundtrack.
Outtloud converts documents and text into natural audio you can listen to at up to 4x speed, even while driving. It's for busy people who'd rather listen than read.
Podcast Marketing AI generates promo content for your podcast episodes, like descriptions, social posts, and quote cards, in minutes. It's for podcasters who want to market their show.
Podfy AI generates transcripts, show notes, timestamps, and tweets from your audio file or YouTube channel with one click. It's for podcasters wanting quick promo content.
SpeechLab automatically dubs and translates audio content, helping publishers and creators reach global audiences. It generates subtitles, translated dubbed audio and video, and editable transcripts using voices similar to the original speakers.
Cockatoo converts audio or video into text in over 90 languages, quickly and accurately. It's useful for anyone needing transcripts of meetings, interviews, or videos.
CrystalSound removes unwanted background noise during video calls, leaving only the speaker's voice. It's especially useful for customer service centers and video conferencing users.
Gemelo creates interactive AI Twins that can represent people, engage with customers, and clone voices at no extra cost. It helps businesses build voice assistants that feel personal and human-like.
Sonura lets you generate music, beats, and sound effects just by describing what you want in text. It helps producers and content creators get original music without hiring musicians.
UnlimitedSFX generates custom sound effects with AI, from an elevator bell to fireworks, for games, videos, and other media. It's for game developers and video creators.
VoiceDrop creates personalized voicemail messages using cloned voice technology and delivers them to thousands of contacts without making a call. It helps sales teams reach many people while keeping a personal touch.
NaturalReader converts text into natural-sounding speech in over 25 languages, helping people with reading difficulties or anyone who prefers listening. It's for students and accessibility needs.
Staccato helps musicians and lyricists generate MIDI ideas and song lyrics when they're stuck, while teaching music structure along the way. It's for musicians and songwriters.
Voicestars lets you remake a popular song using an AI voice model of your choice. It is built for music fans who want to hear their favorite songs sung by a different voice.
Vowen turns speech into polished text entirely offline on your computer, keeping your voice data private. It suits anyone who wants fast, secure hands-free dictation.
Neutone Morpho is a real-time plugin that reshapes any audio input into new sound textures using AI models. It suits musicians and sound designers looking to experiment with fresh sonic effects.
PodPilot generates a podcast series just from your organization's website and the topics you want covered. It can then publish straight to Spotify and other platforms.
Sesame offers an AI voice assistant (Maya or Miles) that talks with human-like emotion, timing, and context awareness. It suits businesses wanting more natural-feeling customer service or personal assistance.
TalkText cleans up your spoken words into polished text, cutting filler sounds like ums and ers. It can also restyle existing written text into a different tone.
Jellypod turns your daily email subscriptions into a short audio podcast you can listen to, even offline, instead of reading through many emails yourself.
TalkNotes turns recorded voice notes into organized text in over 50 languages. It suits people who want to capture ideas quickly without typing.
Smallest.ai delivers studio-quality, ultra-fast AI voice generation and real-time conversational voice agents in over 30 languages. It suits businesses wanting fast, affordable voice-based customer interactions.
LyricStudio helps songwriters write lyrics by offering rhyme and phrase suggestions based on their style, with real-time collaboration features. It's useful for musicians who want support finishing a song.
WhisperTranscribe quickly and accurately transcribes audio into text in more than 55 languages. It also helps generate subtitles and repurpose audio content into other formats.
Voicenotes transcribes your spoken words into searchable text and provides content analysis and organization tools for easy note-taking.
Operator converts text messages into voice phone calls, useful for people who prefer texting or face hearing or speech difficulties. It also includes translation to help bridge language gaps during calls.
Studio Lite helps video creators find songs automatically synced to their video's length from a catalog of over two million tracks. It's available as a standalone app or an Adobe Premiere Pro plugin.
Gladia provides a speech-to-text service for developers building meeting assistants, customer support tools and voice agents, accurately capturing names, numbers and emails across languages and accents.
Voicera adds a read-aloud voice to your articles or blog posts with one click, supporting over ten languages without slowing down your site.
LyRuno separates dialogue, music, and sound effects from mixed audio in films and other content. Film editors and audio professionals use it to quickly extract clean audio for further work.
XspaceGPT converts Twitter Spaces into MP3 audio, text, summaries, and mind maps. Content creators and researchers use it to quickly grasp what was discussed without listening to the whole recording.
Allinpod.ai helps podcast creators find and recommend content that fits their listeners, and gives insight into how to boost engagement.
MIDI Agent is an AI plugin for music production software that generates and edits MIDI patterns, helping composers overcome creative blocks.
Nonoisy automatically cleans up audio, removing background noise and leveling volume so your recordings sound better without a big cost.
PodcastAI offers time-saving AI tools for podcast and YouTube creators, covering production help and everyday assistant tasks.
Amical is an open-source app for dictating instead of typing, transcribing meetings, and taking voice notes. It works in more than 50 languages and can run on your device instead of only in the cloud.
HappySRT is an open-source web app that transcribes audio and video, translates it into other languages, and produces short summaries. Creators use it to get timestamped transcripts ready for subtitles.
PocketPod generates a podcast tailored to whatever topic interests you, even an obscure one, or turns a PDF into audio. It saves you the time of hunting for something worth listening to.
PodManager.AI turns raw recordings into publish-ready podcast episodes, handling transcription, noise removal, chapter titles, and even translation into other languages, free for life.
Audimee lets musicians transform their vocals into professional-sounding voices or even instruments using AI, instead of hiring multiple singers.
Mumble Note is an iOS app that turns spoken words into organized notes, generating summaries, to-do items, and translations into over 40 languages.
PodExtra AI helps podcasters prepare for interviews by researching the guest, suggesting questions, and summarizing content. It saves prep time and helps interviews turn out sharper.
Sonofa turns written content, from blog posts to academic papers, into a conversational podcast in any language. Students and professionals use it to listen to material instead of reading while multitasking.
AudioNotes.ai turns voice recordings into clear text notes, letting you choose the language and summary length. It suits students, journalists and workers who need to keep records of meetings or lectures.
Cloudonix merges voice and data into one communication channel, using AI to improve call quality, routing and customer engagement in real time. Sales teams and remote call centers use it without overhauling their existing infrastructure.
Fish Audio generates human-sounding AI voices from text, with multi-language support and emotional control, drawing on over 200,000 voice models. Creators and developers use it for narration, ads and voice assistants, including quick voice cloning from a short sample.
HookSounds offers royalty-free music, sound effects and intros/outros, plus an AI Studio that automatically creates a soundtrack matching your uploaded video. It suits video creators and businesses needing legally cleared music.
AI Phone provides real-time call transcripts and summaries, keyword alerts, and a second phone number for privacy. It helps users manage work and personal calls more easily.
Lazybird generates human-like AI voice overs for videos, podcasts, audiobooks and other content. It helps creators get natural-sounding narration without hiring professional voice actors.
Vscoped transcribes audio and video content in multiple languages and produces branded, customizable subtitles. It helps content creators get accurate and editable captions.
PlainScribe transcribes audio and video, translates into more than 50 languages, and summarizes the key points. You can download transcripts as CSV or SRT/VTT.
Podnotes automatically generates transcripts, summaries and a chat interface for your podcasts and videos, in over 19 languages. It suits podcast creators and content writers.
Sluqe turns your voice notes and meeting recordings into searchable text with key takeaways and action items. Useful for professionals and students who want quick summaries of what was said.
Dumai transcribes voice recordings into searchable, organized text and generates concise summaries, helping users track lectures or meetings without relistening to everything.
Retellio turns customer call recordings into short audio summaries, helping business leaders extract key insights without listening to full calls.
Updaytr lets you call an AI agent to share updates or ideas by voice, then automatically turns the conversation into a structured, formatted report delivered by email on your schedule.
Maastr masters your tracks online, boosting loudness, clarity and balance without flattening the sound, tuned especially for guitar-driven genres like punk and metal. It suits musicians who want studio-quality results affordably.
Producer is an AI music studio you chat with to compose full songs, generate vocals and produce music videos. You can also build your own custom instruments and remix tracks in your own style.
Heynds turns your speech into polished text with AI, translating into over 100 languages - helping you write by talking instead of typing.
Wavve AI turns your spoken words into neatly structured text like meeting notes or emails, so you don't have to type them out yourself.
Vowise converts your speech into accurate real-time text, learning your vocabulary and auto-organizing notes so you can work hands-free.
FlowSpeech turns text, images or PDFs into natural-sounding speech with proper tone and emotion, useful for podcasts, audiobooks and accessibility needs.
Rounded is an AI voice agent platform that automates business phone calls, handling appointments and customer support with natural-sounding conversation. It helps companies provide round-the-clock call coverage while reducing staff workload.
Audjust AI is a web-based audio editor that shortens or extends tracks while keeping their musical structure intact, and can generate music from text or images. It helps creators quickly produce ready-to-use, royalty-free music and loops.
ChordGen is a free browser tool that turns simple text prompts like "melancholic jazz" into playable chord progressions you can edit and export as MIDI. It helps songwriters and learners overcome writer's block without needing music theory.
FastScribe transcribes audio and video with high accuracy, labels different speakers, detects over 99 languages, and delivers results fast - a one-hour file can be done in about five minutes.
Koolio turns a prompt, document or recording into a finished podcast or ad, automatically handling transcription, voice cloning, filler-word removal and studio-quality audio polish.
VIXSOUND is an AI assistant inside Ableton Live that generates melodies, chords and drums from voice or text commands. It speeds up music production while keeping files private and giving producers full creative control.
FluidVoice is a Mac app that turns your speech into text offline, then cleans up the wording automatically. It suits anyone who prefers dictating over typing and wants full privacy.
Boson AI builds voice systems for customer support and sales that can hold real-time conversations in over 100 languages and handle interruptions naturally. It also offers text-to-speech and speech-to-text tools.
Cadence is a screen-recording app that cleans up your voice with one click, removing filler words and background noise, and can even shift your accent. It's built for teachers and tutorial creators.
Dograh is open-source software for building AI voice agents that handle phone calls, which you can host on your own servers. Companies can pick their own speech and telephony tools and support over 70 languages.
Liso turns text you highlight on any webpage or article into audio you can listen to later. A Chrome extension sends the text straight to your listening library.
LongScribe turns long videos or live broadcasts from YouTube, TikTok, and other sites into clean text, subtitles, and Word files. It processes in the background even without a browser tab open.
ScriberGPT turns uploaded audio or video, even a YouTube link, into a full transcript with speaker labels, timestamps, and punctuation. It supports over 90 languages and exports SRT, DOCX, and TXT files.
TranscriptAPI.com pulls fast, timestamped, multilingual transcripts from any YouTube video through a simple API, ready for captions or feeding into an AI assistant.
Viora is a Mac app where you hold one shortcut, speak naturally, and it writes, edits selected text, or answers a question inside any app. It remembers your preferences so its output matches your style.
Vocallab AI generates natural-sounding speech from text or clones a person's voice from a short audio sample, offering over 300 professional voices in multiple languages. It also exports captions with words highlighted karaoke-style.
CastReader reads webpages, PDFs, and Kindle books aloud in natural voices while highlighting each word, helping busy readers or people who find reading text difficult.
Gulab Music writes a song, frames the shots, and syncs the vocals into a finished music video, starting from just a short description of your idea.
MeetStream is an API that lets developers plug in bots that join and record Zoom, Google Meet, and Teams calls, then deliver transcripts labeled by speaker.
Monologue is a Mac and iOS app that turns your speech into clean text inside any app, learning your vocabulary and context so little editing is needed. It suits writers and professionals who'd rather dictate than type.
OnCallClerk answers business calls around the clock with a natural-sounding AI voice - booking appointments, capturing lead details, filtering spam, and writing a summary of every call. It helps businesses stop missing calls and cut receptionist costs.
OpenTypeless is a free, open-source voice dictation app for Windows, Mac, and Linux that recognizes 99 languages and cleans up grammar based on which app you're typing into.
Podsuite turns a single uploaded podcast episode into full transcripts, show notes, short highlight clips, and blog and social posts. It saves podcasters hours of manual work after every episode.
Sabato AI sets up voice agents that answer customer phone calls for online stores, handling order status, returns, and complaints, plugged straight into the store's systems.
Soniox turns text into natural-sounding speech in over 60 languages and can clone a person's voice from just a short recording.
Sparrow-2 is a real-time AI model that listens to conversation and decides when to listen, wait, or speak, even in noisy rooms with multiple speakers.
Stellar sends out AI voice agents to make phone calls that qualify leads, book appointments, and confirm events in natural conversation, improving themselves from every call. It helps businesses cut response times without hiring extra staff.
Tonecraft VO turns recorded audio or video into speaker-separated scripts, then generates fresh AI voice-overs so you can fix dialogue without re-recording it.
Typecast API gives developers access to over 700 AI voices with emotional tone that adjusts to context, suited for apps, conversational agents, and audio content.
videotranscript.ai turns YouTube videos or audio files into searchable, timestamped text, complete with chapters and short summaries.
Vocal Slice transcribes your recording right on your computer, then lets you cut a precise audio clip just by selecting words in the text, with nothing sent online.
Voiceover QA compares AI-narrated audio against the original script to catch missing or changed words, giving you a timestamped list of fixes before you publish.
Whisperstream types out your speech directly into any Windows app at the press of a key, processing everything on your own machine with no audio sent online, for a one-time fee.
Willow Voice turns your spoken words into clean, correctly formatted text on your computer or phone, so you dictate with one hotkey instead of typing.
Wisprkey types out what you say directly into any Mac app, running entirely on the device with no internet, recognizing 31 languages and reading text back aloud too.
AI Jingle Maker turns a written script into a finished radio jingle by pairing an AI voice with a chosen music style, ready to download as an MP3.
This app generates full songs, beats, and lyrics from a prompt, letting you pick genre, instruments, mood, and tempo. It also removes vocals and splits tracks into stems.
AI Songmaker creates custom songs and melodies from your descriptions, making music-making easy for both professionals and hobbyists.
All Voice Lab generates AI speech, clones a real voice, changes voices in a recording, and dubs video into new languages, supporting 33 languages with control over tone, speed, and pitch.
Audeus is a Chrome extension that reads Google Docs, Gmail, webpages, and PDFs aloud, letting you pick the voice and reading speed.
AudioCleaner AI cleans up audio in videos or recordings right in the browser, stripping out background noise, echo, and filler words. It helps content creators get clear sound without installing any software.
AudioScribe transcribes audio and video recordings into text, letting you search the transcript, get summaries, and identify each speaker.
Covers.ai can swap a song's singer, lyrics, genre, or language using AI voices. You record or upload vocals, pick an AI voice, and create covers, duets, mashups, or text-to-speech clips for short-form content.
DiffRhythm generates a full song, vocals and instruments included, from lyrics and a style description. It can produce a track nearly five minutes long in just seconds, which suits musicians testing ideas fast.
This mobile-first music app turns a text description, chosen genre, or your own lyrics into a full song with vocals and instruments. It suits anyone wanting to make a song without a studio.
This free tool converts written text into natural-sounding AI speech. You can pick a language and voice style, adjust pitch and speed, and download the resulting audio.
This AI voice platform offers over 500 voices in more than 100 languages, and can write scripts, add subtitles, and assemble simple videos. It suits creators who need narration voiceovers.
InspireMusic is a framework for generating long-form, high-fidelity music, songs, and audio from text or sample audio prompts. It uses advanced super-resolution techniques to keep the sound coherent even over longer pieces.
This is a self-hosted speech-to-text API that companies can run on their own systems instead of relying on outside servers. It supports many languages and reduces background noise.
Mureka generates songs, lyrics, and background tracks from prompts, letting you pick genre, mood, and tempo, then download the full track or individual stems.
This audio-enhancement tool serves listeners, artists, and producers, using algorithms to optimize sound quality for different needs. It suits people wanting cleaner audio in their music work.
Music Remover splits an uploaded audio or video file into vocal, background music, and other tracks, so you can download each part on its own.
NowVoice converts written scripts into AI-generated speech using over 400 voices across 13 languages, for use in videos, ads, and tutorials.
PodLM creates podcast scripts and narrated audio from a URL, pasted text, a document, or a chosen topic. It supports multiple AI speakers, voice choices, and background music, no recording gear needed.
RambleFix turns spoken thoughts into clean, organized text. You talk into the microphone, and the AI arranges your words into readable sentences, handy for anyone who thinks faster out loud than on a keyboard.
Retell AI is for building voice agents that take and make customer calls. A team picks which model, which voice and which telephone provider to put behind them.
This Google app generates full songs with vocals from a text prompt, and also supports remixing, stem splitting, and music-video creation. It suits musicians wanting to try ideas quickly.
This speech-to-text API converts recorded audio or video into text, supporting 100 languages and files up to 24 hours long. It can also separate different speakers.
Speechmatics converts speech to text and text to speech in more than 55 languages, recognizing multiple speakers with very low delay.
SpotScribe transcribes Spotify, Apple Podcasts, YouTube, or TikTok episodes into searchable text, adding a summary and letting you ask questions about the episode.
antiplaylist is an index of more than 700 music genres, from experimental to underground to world music. Each genre gets AI-generated descriptions and artwork alongside human-picked tracks, with no account needed to browse.
Voice Enhancer cleans up speech in uploaded audio or video, stripping out background noise, echo, and hesitations, with a before-and-after comparison before you download.
This was a browser-based music studio for recording, arranging, and mixing without installing software, with an AI assistant that added chord progressions and split stems. It is currently offline.
Wondera generates full songs from text prompts, including vocals, music videos, and licenses for commercial use.
No tools found. Try another word.