Skip to course content
Free computer vision course

Computer Vision and Multimodal AI

Module 10

Multimodal Prompts and Vision-Language Model Limits

By the end of this module you can prompt a vision-language model effectively and state its real limits: confident descriptions of things that are not there, text it cannot read, and detail it invents to fill a gap.

Units

  1. Unit 10.00: Asking a vision-language model a checkable question
  2. Unit 10.01: What it cannot see, and will answer anyway
  3. Unit 10.02: Counting, reading small text, and spatial relations
  4. Unit 10.03: Grounding an answer in the image region
  5. Unit 10.04: Deciding when the model should refuse

Module work