Skip to main content
Models that read an image or video as input and answer in text, through POST /v1/chat/completions. 13 models. Grouped by input, because that is the choice you make first — a model that turns an image into a video is not interchangeable with one that starts from a prompt. Prices are not listed here: they change, and a stale price is worse than none. See the live catalogue for current rates, and each model’s own page there for its full parameter schema.

Image → text (9)

Video → text (4)