Computer Vision
AI That Sees: From Specialised Models to Multimodal LLMs
Computer vision is the field enabling machines to interpret visual information — images and video. It underpins facial recognition, medical imaging, autonomous vehicles, and industrial inspection. The field is undergoing a structural shift: standalone CV models (YOLO, ResNet, SAM) are increasingly being supplemented or replaced by multimodal LLMs that combine vision with language understanding, enabling richer reasoning over visual content.
- Object detection: identify and locate objects in an image — powers security cameras, inventory systems, autonomous vehicles
- Image classification: what category does this image belong to? — spam filters, content moderation, medical triage
- Semantic segmentation: label every pixel — surgical planning, satellite analysis, autonomous driving
- Video understanding: track objects and actions across frames — surveillance, sports analytics, manufacturing QA
Standalone CV vs. Multimodal LLMs: Choosing the Right Tool
The choice between a dedicated computer vision model and a multimodal LLM depends on what you're optimising for. Standalone CV models win on speed and cost for high-volume, single-task classification. Multimodal LLMs win when you need language reasoning alongside visual understanding — asking "why is this anomalous" rather than "is this anomalous".
- Standalone CV (YOLO, EfficientDet): fast, cheap, ideal for real-time or high-volume single-task pipelines
- Multimodal LLMs (GPT, Gemini, Claude): slower, costlier, but can explain, compare, and reason across visual content
- Hybrid: use standalone CV to detect, then multimodal LLM to interpret and recommend — best of both
- The convergence direction: frontier labs are absorbing specialised CV capabilities into their multimodal models each release cycle
Security-Specific Computer Vision Use Cases
Security applications of computer vision span the physical and digital worlds. The most mature use cases are already in production; the emerging ones are where competitive advantage lies in 2025–2026.
- Mature: badge access analytics, CCTV object detection, licence plate recognition, perimeter monitoring
- Emerging: deepfake image/video detection for BEC and fraud, visual phishing detection (logo spoofing in screenshots), malware UI screenshot analysis
- AI-native: multimodal SOC copilots that can receive a screenshot of a suspicious email or alert dashboard and reason about it directly
- Key question: is your current visual security workflow powered by a standalone CV model, a multimodal LLM, or still human eyes?
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.