Multimodal Prompts and Vision-Language Model Limits
By the end of this module you can prompt a vision-language model effectively and state its real limits: confident descriptions of things that are not there, text it cannot read, and detail it invents to fill a gap.
Units
- Unit 10.00: Asking a vision-language model a checkable question
- Unit 10.01: What it cannot see, and will answer anyway
- Unit 10.02: Counting, reading small text, and spatial relations
- Unit 10.03: Grounding an answer in the image region
- Unit 10.04: Deciding when the model should refuse
