Skip to main content
Speech synthesis and audio transformation, through POST /v1/audio/speech and /v1/audio/transformations. 88 models. Grouped by input, because that is the choice you make first — a model that turns an image into a video is not interchangeable with one that starts from a prompt. Prices are not listed here: they change, and a stale price is worse than none. See the live catalogue for current rates, and each model’s own page there for its full parameter schema.

Text → audio (45)

Audio → audio (38)

Video → audio (5)