OpenAI · LLM

GPT-4o

Omni-model unifying text, vision and audio in a single network with realtime capabilities.

multimodalvoicemultimodal

Capabilities

  • realtime voice
  • vision
  • tool use

Use cases

voice assistants · multimodal apps

Lineage

View the full GPT lineage →

GPT-4o head-to-head comparisons

Go deeper

Sources & verification

Information is compiled from official announcements, documentation and reputable industry coverage.
Last verified Sep 10, 2026

Report an issue →