HomeGlossary › Multimodality

Multimodality

A definition from the GenAI for Business glossary.

A multimodal model can process more than one type of data, text, image, audio, video, or code and combine them in output. For example, describing an image, generating a video from text, or analyzing both visuals and language together. This is central to next-generation AI (the current GPT, Claude, and Gemini flagships).

Want the full picture? This term comes from GenAI for Business, a complete free book on generative AI strategy and implementation by Prof. Shubin Yu (HEC Paris).
Token Limit / Context WindowWorkflow Automation