> ## Documentation Index
> Fetch the complete documentation index at: https://docs.infery.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Music models

> Every music model on Infery, grouped by what it takes as input.

Called through `POST /v1/music/generations`.

**41 models.** Grouped by input, because that is the choice you make first — a model that
turns an image into a video is not interchangeable with one that starts from a prompt.

Prices are not listed here: they change, and a stale price is worse than none. See the
[live catalogue](https://infery.ai/models) for current rates, and each model's own page there for its full
parameter schema.

## Text → music (39)

| Model                                           | Owner        | Also accepts | Notes                                                                                                                                        |
| ----------------------------------------------- | ------------ | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `ace-step`                                      | ace-step     | —            | Generate music with lyrics from text using ACE-Step                                                                                          |
| `ace-step-prompt-to-audio`                      | ace-step     | —            | Generate music from a simple prompt using ACE-Step                                                                                           |
| `seed-audio-1.0`                                | bytedance    | Audio        | Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or…   |
| `music-generator`                               | cassetteai   | —            | CassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds.                            |
| `sound-effects-generator`                       | cassetteai   | —            | Create stunningly realistic sound effects in seconds - CassetteAI's Sound Effects Model generates high-quality SFX up to 30 seconds long in… |
| `elevenlabs-music`                              | elevenlabs   | —            | Generate high quality, realistic music with fine controls using Elevenlabs Music!                                                            |
| `elevenlabs-sound-effects-v2`                   | elevenlabs   | —            | Generate sound effects using ElevenLabs advanced sound effects model.                                                                        |
| `elevenlabs-text-to-dialogue-eleven-v3`         | elevenlabs   | —            | Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.                                                                          |
| `elevenlabs-tts-eleven-v3`                      | elevenlabs   | —            | Generate text-to-speech audio using Eleven-v3 from ElevenLabs.                                                                               |
| `elevenlabs-tts-multilingual-v2`                | elevenlabs   | —            | Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.                                                             |
| `gemini-tts`                                    | google       | —            | Use Gemini TTS Models to convert your prompts to real audio.                                                                                 |
| `lyria-3-clip`                                  | google       | —            | 1049k ctx, 66k out                                                                                                                           |
| `lyria-3-pro`                                   | google       | —            | 1049k ctx, 66k out                                                                                                                           |
| `lyria3`                                        | google       | —            | Lyria 3 is most recent music model from Google                                                                                               |
| `ltx-2.3-quality-text-to-audio`                 | lightricks   | —            | Text to Audio high-quality using LTX-2.3                                                                                                     |
| `ltx-2.3-quality-text-to-audio-lora`            | lightricks   | —            | Text to Audio high-quality using LTX-2.3 with Lora                                                                                           |
| `minimax-music-v1.5`                            | minimax      | —            | Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical…    |
| `minimax-music-v2`                              | minimax      | —            | Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse…  |
| `minimax-music-v2.5`                            | minimax      | —            | MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.        |
| `minimax-music-v2.6`                            | minimax      | —            | MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.        |
| `music-3`                                       | minimax      | —            | MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long                             |
| `sfx1.6-text-to-audio`                          | mirelo-ai    | —            | Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.                           |
| `csm-1b`                                        | sesame       | —            | CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.        |
| `v1.1-text-to-music`                            | sonilo       | —            | Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact…     |
| `v1.1-text-to-sound-effects`                    | sonilo       | —            | Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact…    |
| `stable-audio`                                  | stability-ai | —            | Open source text-to-audio model.                                                                                                             |
| `stable-audio-25-text-to-audio`                 | stability-ai | —            | Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI                                                        |
| `stable-audio-3-medium-base-text-to-audio`      | stability-ai | —            | Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes…       |
| `stable-audio-3-medium-text-to-audio`           | stability-ai | —            | Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text…  |
| `stable-audio-3-small-music-base-text-to-audio` | stability-ai | —            | Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes…     |
| `stable-audio-3-small-music-text-to-audio`      | stability-ai | —            | Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes…  |
| `stable-audio-3-small-sfx-base-text-to-audio`   | stability-ai | —            | Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as…  |
| `stable-audio-3-small-sfx-text-to-audio`        | stability-ai | —            | Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts…      |
| `suno-v4`                                       | suno         | —            |                                                                                                                                              |
| `suno-v4-5-all`                                 | suno         | —            |                                                                                                                                              |
| `suno-v4.5`                                     | suno         | —            |                                                                                                                                              |
| `suno-v4.5-plus`                                | suno         | —            |                                                                                                                                              |
| `suno-v5`                                       | suno         | —            |                                                                                                                                              |
| `suno-v5.5`                                     | suno         | —            |                                                                                                                                              |

## Audio → music (2)

| Model           | Owner   | Also accepts | Notes                                                                                                                                     |
| --------------- | ------- | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------- |
| `minimax-music` | minimax | —            | Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical… |
| `zonos`         | zonos   | —            | Clone voice of any person and speak anything in their voice using zonos' voice cloning.                                                   |
