Vision
Vision lets supported models answer questions about images and PDFs. Use it for image understanding, screenshots, diagrams, scanned pages, and document summaries.
Mixlayer supports vision input through both Chat Completions and Responses. Use a model documented as supporting vision; text-only models return invalid_media when media is included.