ace-step | ace-step | — | Generate music with lyrics from text using ACE-Step |
ace-step-prompt-to-audio | ace-step | — | Generate music from a simple prompt using ACE-Step |
seed-audio-1.0 | bytedance | Audio | Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or… |
music-generator | cassetteai | — | CassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds. |
sound-effects-generator | cassetteai | — | Create stunningly realistic sound effects in seconds - CassetteAI’s Sound Effects Model generates high-quality SFX up to 30 seconds long in… |
elevenlabs-music | elevenlabs | — | Generate high quality, realistic music with fine controls using Elevenlabs Music! |
elevenlabs-sound-effects-v2 | elevenlabs | — | Generate sound effects using ElevenLabs advanced sound effects model. |
elevenlabs-text-to-dialogue-eleven-v3 | elevenlabs | — | Generate realistic audio dialogues using Eleven-v3 from ElevenLabs. |
elevenlabs-tts-eleven-v3 | elevenlabs | — | Generate text-to-speech audio using Eleven-v3 from ElevenLabs. |
elevenlabs-tts-multilingual-v2 | elevenlabs | — | Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2. |
gemini-tts | google | — | Use Gemini TTS Models to convert your prompts to real audio. |
lyria-3-clip | google | — | 1049k ctx, 66k out |
lyria-3-pro | google | — | 1049k ctx, 66k out |
lyria3 | google | — | Lyria 3 is most recent music model from Google |
ltx-2.3-quality-text-to-audio | lightricks | — | Text to Audio high-quality using LTX-2.3 |
ltx-2.3-quality-text-to-audio-lora | lightricks | — | Text to Audio high-quality using LTX-2.3 with Lora |
minimax-music-v1.5 | minimax | — | Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical… |
minimax-music-v2 | minimax | — | Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse… |
minimax-music-v2.5 | minimax | — | MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description. |
minimax-music-v2.6 | minimax | — | MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description. |
music-3 | minimax | — | MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long |
sfx1.6-text-to-audio | mirelo-ai | — | Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes. |
csm-1b | sesame | — | CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. |
v1.1-text-to-music | sonilo | — | Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact… |
v1.1-text-to-sound-effects | sonilo | — | Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact… |
stable-audio | stability-ai | — | Open source text-to-audio model. |
stable-audio-25-text-to-audio | stability-ai | — | Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI |
stable-audio-3-medium-base-text-to-audio | stability-ai | — | Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes… |
stable-audio-3-medium-text-to-audio | stability-ai | — | Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text… |
stable-audio-3-small-music-base-text-to-audio | stability-ai | — | Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes… |
stable-audio-3-small-music-text-to-audio | stability-ai | — | Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes… |
stable-audio-3-small-sfx-base-text-to-audio | stability-ai | — | Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as… |
stable-audio-3-small-sfx-text-to-audio | stability-ai | — | Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts… |
suno-v4 | suno | — | |
suno-v4-5-all | suno | — | |
suno-v4.5 | suno | — | |
suno-v4.5-plus | suno | — | |
suno-v5 | suno | — | |
suno-v5.5 | suno | — | |