AI Glossary term

Multimodal

A model that handles more than one kind of input or output - text plus images, audio or video.

Also called: vision, image input

A model that handles more than one kind of input or output - text plus images, audio or video.

Why it mattersScreenshots, whiteboards, PDFs and invoices become usable input.

See also Generative AI

Explore more AI terms

Browse the full glossary for plain-English definitions across models, agents, data, and safety.

Back to all terms