44. Multimodal and Vision Models
Understand when and why to use vision-capable AI models.
By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.
The lesson
Vision models like GPT-5.5, Claude 3, and Gemini can process images alongside text, enabling UI mockup analysis, chart interpretation, and photo understanding.
Use vision models when the image contains essential information: diagrams, schematics, UI designs, data visualizations, or photographs relevant to the task.
Images are tokenized and count significantly toward context limits and cost - a single high-res image can be thousands of tokens.
For technical schematics or detailed diagrams, use frontier vision models with high-resolution capabilities.
Vision models do not replace careful inspection - they assist interpretation but may miss subtle details.
A biomedical technician photographs a device fault panel and asks a vision model to read the error codes, treating the reading as a hint before checking the service manual.
A quality inspector at a manufacturer uploads photos of surface defects and asks a vision model to group them by pattern, then confirms each group by eye.
A site engineer compares progress photos across weeks with a vision model, then walks the site to verify what the model reported.
A mechanic photographs a diagnostic screen and asks a vision model to interpret the readings, checking the result against the workshop manual before advising the customer.
An agronomist uploads leaf photos and asks a vision model to describe visible symptoms, then confirms the finding with a physical field sample.
An architect uploads a facade concept and a site photo, asking a vision model to describe how the two relate, while treating it as one input to careful inspection.
A robotics engineer photographs a wiring diagram and asks a vision model to list the labelled terminals, then verifies each one against the original drawing.
A content manager at an online store asks a vision model to describe product photos for catalogue text, checking colours and details against the item itself.
A claims assessor uploads damage photos and asks a vision model to list visible features, then compares that list with the policy wording before deciding.
An asset inspector photographs equipment nameplates and asks a vision model to read the ratings, then confirms the figures against the maintenance register.
Check yourself
Question 1: When should you use a vision-capable model like GPT-5.5 or Claude Opus 4.8 with image input?
- Only for text summarization
- When analyzing UI mockups, diagrams, charts, or photos relevant to the task — correct
- For faster text-only output
- When you have no GPU
Answer: When analyzing UI mockups, diagrams, charts, or photos relevant to the task
Vision models like GPT-5.5 and Claude Opus 4.8 excel at tasks requiring understanding of images, diagrams, charts, or UI elements.
Question 2: What is a cost consideration when using vision models?
- Images are always free to process
- Image tokens count toward context and cost significantly — correct
- Vision models cannot see text
- Only small images work
Answer: Image tokens count toward context and cost significantly
Images are tokenized and can significantly increase context usage and cost.
Question 3: For analyzing a scanned technical schematic, which model choice makes sense?
- Text-only embedding model
- Vision-capable frontier model with high-resolution support — correct
- A 1B parameter text model
- A music generation model
Answer: Vision-capable frontier model with high-resolution support
Technical schematics need vision-capable frontier models that can read detailed diagrams.
← Previous lesson · All 91 lessons · Next lesson →
The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.