A multimodal model might read a screenshot and answer questions about it, or take an image plus text and produce a video.
Example
Uploading a photo to a chatbot and asking it to write a video prompt based on it.
Why it matters to you
It makes AI tools more flexible and closer to how people communicate.
Related terms
Computer VisionThe part of AI that lets computers understand images and video./knowledge-base/computer-vision/LLM (Large Language Model)An AI model trained on huge amounts of text so it can understand and write language./knowledge-base/llm/Image-to-VideoTurning a still image into a moving video clip with AI./knowledge-base/image-to-video/