AI That Sees and Hears
Beyond Words
AI isn't just about text and chat. Modern AI can look at a photo and tell you what's in it. It can hear a song and name the artist. It can watch a video and sum up what happened. This is called multimodal AI.
- Vision AI: recognises faces, reads text in images, spots objects in photos
- Voice AI: turns speech into text, generates speech from text
- Medical AI: analyses X-rays and scans to help doctors spot problems
- These AIs are already in your life: face unlock, voice assistants, photo search
One Brain, Many Senses
Multimodal is a fancy word with a simple meaning: many modes — words, pictures, and sounds, all understood by one AI. Try these party tricks. Show a multimodal AI a photo of your messy desk and ask for a funny poem about it — it can do that. Hum a tune to a music app and it names the song. Point a camera at a plant and ask "what is this, and is it safe for my cat?" Snap your maths homework and ask it to check your working. The magic is the mixing: it can look at a picture AND talk about it, hear a question AND answer in words. Just like you use eyes and ears together, multimodal AI blends senses into one understanding.
- "Multimodal" = one AI that handles words, images, and sound together
- Photo in, poem out — hummed tune in, song name out
- The clever part is mixing the senses, not just having them
- Family experiment: photograph something odd and ask an AI to describe it
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.