Computer Vision
The field of AI that enables machines to interpret and act on visual information from images and video.
Computer vision spans classification, object detection, segmentation, OCR, pose estimation, and video understanding. Modern systems are dominated by transformer architectures (ViT, DETR) and multimodal foundation models that combine vision with language.
Enterprise deployments run the gamut from quality inspection on factory lines to shelf analytics in retail to medical image analysis. Edge deployment on NVIDIA Jetson, Coral, or mobile NPUs is common when latency or bandwidth rules out the cloud.
Vision-language models (GPT-4V, Claude, Gemini) have collapsed the cost of many CV tasks — a screenshot and a prompt now replace a custom-trained classifier for prototypes.
Key points
- Labeled data quality dominates model quality
- Edge inference reduces latency and bandwidth cost
- Vision-language models cover long-tail cases classical CV models miss
- Privacy and consent are first-class concerns in any camera deployment
Common use cases
Frequently asked
Related terms
Extracting machine-readable text and structure from images, scans, and PDFs.
Running AI models on-device or on local hardware instead of in the cloud.
Models that understand and generate across multiple input types — text, images, audio, and video.
Applying Computer Vision in your business?
We help enterprises design, ship, and govern AI systems end-to-end.