44. Multimodal and Vision Models

Understand when and why to use vision-capable AI models.

By Jacques Botte, founder of Toptronic®. Last updated 19 September 2026.

The lesson

Vision models like GPT-5.5, Claude 3, and Gemini can process images alongside text, enabling UI mockup analysis, chart interpretation, and photo understanding.

Use vision models when the image contains essential information: diagrams, schematics, UI designs, data visualizations, or photographs relevant to the task.

Images are tokenized and count significantly toward context limits and cost - a single high-res image can be thousands of tokens.

For technical schematics or detailed diagrams, use frontier vision models with high-resolution capabilities.

Vision models do not replace careful inspection - they assist interpretation but may miss subtle details.

A biomedical technician photographs a device fault panel and asks a vision model to read the error codes, treating the reading as a hint before checking the service manual.

A quality inspector at a manufacturer uploads photos of surface defects and asks a vision model to group them by pattern, then confirms each group by eye.

A site engineer compares progress photos across weeks with a vision model, then walks the site to verify what the model reported.

A mechanic photographs a diagnostic screen and asks a vision model to interpret the readings, checking the result against the workshop manual before advising the customer.

An agronomist uploads leaf photos and asks a vision model to describe visible symptoms, then confirms the finding with a physical field sample.

An architect uploads a facade concept and a site photo, asking a vision model to describe how the two relate, while treating it as one input to careful inspection.

A robotics engineer photographs a wiring diagram and asks a vision model to list the labelled terminals, then verifies each one against the original drawing.

A content manager at an online store asks a vision model to describe product photos for catalogue text, checking colours and details against the item itself.

A claims assessor uploads damage photos and asks a vision model to list visible features, then compares that list with the policy wording before deciding.

An asset inspector photographs equipment nameplates and asks a vision model to read the ratings, then confirms the figures against the maintenance register.

Check yourself

Question 1: When should you use a vision-capable model like GPT-5.5 or Claude Opus 4.8 with image input?
  1. Only for text summarization
  2. When analyzing UI mockups, diagrams, charts, or photos relevant to the task — correct
  3. For faster text-only output
  4. When you have no GPU

Answer: When analyzing UI mockups, diagrams, charts, or photos relevant to the task

Vision models like GPT-5.5 and Claude Opus 4.8 excel at tasks requiring understanding of images, diagrams, charts, or UI elements.

Question 2: What is a cost consideration when using vision models?
  1. Images are always free to process
  2. Image tokens count toward context and cost significantly — correct
  3. Vision models cannot see text
  4. Only small images work

Answer: Image tokens count toward context and cost significantly

Images are tokenized and can significantly increase context usage and cost.

Question 3: For analyzing a scanned technical schematic, which model choice makes sense?
  1. Text-only embedding model
  2. Vision-capable frontier model with high-resolution support — correct
  3. A 1B parameter text model
  4. A music generation model

Answer: Vision-capable frontier model with high-resolution support

Technical schematics need vision-capable frontier models that can read detailed diagrams.

← Previous lesson · All 91 lessons · Next lesson →

The full course — 91 lessons and 273 quiz questions — ships inside the app. Get TPEE to study it offline.