POST /v1/audio/transcriptions.
18 models. Grouped by input, because that is the choice you make first — a model that
turns an image into a video is not interchangeable with one that starts from a prompt.
Prices are not listed here: they change, and a stale price is worse than none. See the
live catalogue for current rates, and each model’s own page there for its full
parameter schema.
Models
Speech-to-text models
Every speech-to-text model on Infery, grouped by what it takes as input.
Transcription, through