Multimodality
A multimodal model can process more than one type of data, text, image, audio, video, or code and combine them in output. For example, describing an image, generating a video from text, or analyzing both visuals and language together. This is central to next-generation AI (the current GPT, Claude, and Gemini flagships).
Want the full picture? This term comes from GenAI for Business, a complete free book on generative AI strategy and implementation by Prof. Shubin Yu (HEC Paris).