Skip to main content
Transcription, through POST /v1/audio/transcriptions. 18 models. Grouped by input, because that is the choice you make first — a model that turns an image into a video is not interchangeable with one that starts from a prompt. Prices are not listed here: they change, and a stale price is worse than none. See the live catalogue for current rates, and each model’s own page there for its full parameter schema.

Audio → text (18)