AI / A CONCEPT NOTE
Multimodal Models
AI models that understand text, images, audio, and video together
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Multimodal models (GPT-4V, Claude 3, Gemini) can process and reason across multiple data types simultaneously. You can show them a screenshot and ask 'What's the error here?', give them a diagram and ask them to generate Terraform code, or show a video and ask for a summary. The model encodes all modalities into a shared representation space.
02 / FOLLOW THE MECHANISM
How a multimodal model processes an image
Input
the user sends: 'What's wrong with this Dockerfile?' + an image of the Dockerfile.
Vision encoder
the image is split into patches and processed through a vision transformer (ViT) into embeddings.
Projection layer
image embeddings are projected into the same latent space as text embeddings via a connector module.
Language model
the combined text + image embeddings are processed through the transformer — the model 'sees' the image and generates text about it.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
send an image to GPT-4 Vision
curl -X POST -d '{"model":"gpt-4-vision-preview","messages":[{"role":"user","content":[{"type":"text","text":"What's in this image?"},{"type":"image_url","image_url":{"url":"https://example.com/screenshot.png"}}]}]}' https://api.openai.com/v1/chat/completionsrun multimodal model locally with Ollama
ollama run llava 'Describe this image' --image photo.jpg05 / CHECK YOURSELF
Could you explain Multimodal Models to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringTokenizationhow LLMs break text into pieces their neural networks can process