AI / A CONCEPT NOTE

Multimodal Models

AI models that understand text, images, audio, and video together

~80 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Multimodal models (GPT-4V, Claude 3, Gemini) can process and reason across multiple data types simultaneously. You can show them a screenshot and ask 'What's the error here?', give them a diagram and ask them to generate Terraform code, or show a video and ask for a summary. The model encodes all modalities into a shared representation space.

02 / FOLLOW THE MECHANISM

How a multimodal model processes an image

  1. Input

    the user sends: 'What's wrong with this Dockerfile?' + an image of the Dockerfile.

  2. Vision encoder

    the image is split into patches and processed through a vision transformer (ViT) into embeddings.

  3. Projection layer

    image embeddings are projected into the same latent space as text embeddings via a connector module.

  4. Language model

    the combined text + image embeddings are processed through the transformer — the model 'sees' the image and generates text about it.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

send an image to GPT-4 Vision

curl -X POST -d '{"model":"gpt-4-vision-preview","messages":[{"role":"user","content":[{"type":"text","text":"What's in this image?"},{"type":"image_url","image_url":{"url":"https://example.com/screenshot.png"}}]}]}' https://api.openai.com/v1/chat/completions

EXAMPLE 02 · REFERENCE

run multimodal model locally with Ollama

ollama run llava 'Describe this image' --image photo.jpg

Explore command anatomy in the CLI lab