Models

Multimodal AI

Models that understand and generate across multiple input types — text, images, audio, and video.

Multimodal models process combinations of text, images, audio, and video in a single forward pass. GPT-4o, Claude 3.5+, Gemini 2, and open models like Llama 3.2 Vision have collapsed the boundary between what used to be separate NLP and computer-vision pipelines.

For enterprises, multimodal capability opens use cases that were previously multi-model engineering projects: 'here is a screenshot of an error, tell me what the user is trying to do' or 'watch this five-minute video and extract the action items'.

Cost and latency for multimodal inputs are still higher than text-only, but falling fast. Vision inputs typically add 1,000–3,000 tokens per image at current tokenizers.

Key points

  • Single model replaces multi-step OCR + captioning + LLM pipelines
  • Vision + text is production-ready; video and audio-in still maturing
  • Native tool-calling on multimodal inputs unlocks new agent designs
  • Prompt caching materially reduces cost for repeated image contexts

Common use cases

Screen understanding for support agents
Product photo analysis in e-commerce
Meeting recording summarization
Accessibility (image description, live captions)

Frequently asked

Related terms

Applying Multimodal AI in your business?

We help enterprises design, ship, and govern AI systems end-to-end.