What is computer vision? How machines see images, the classic pipeline versus deep learning, and where it's used, from cameras to medical scans

Computer vision is the field of getting machines to extract meaning from images and video: what is in the picture, where it is, and what is happening. How it evolved from hand-engineered features to convolutional networks to vision transformers, the core tasks (classification, detection, segmentation, tracking), where it is genuinely used today, and how it relates to the vision-language models that now sit alongside it.

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Key takeaways

  • Computer vision is the general field; classification, detection, segmentation and tracking are its core tasks, each answering a different question about an image.
  • Deep learning, specifically convolutional neural networks and now vision transformers, replaced decades of hand-engineered feature extraction because it learns the features instead.
  • Classic computer vision techniques, edge detection, feature matching, still power a lot of real-time and resource-constrained systems where a full deep model is overkill.
  • Vision-language models are a newer, adjacent layer that reads and reasons about images in open-ended language; classic computer vision tasks still do the precise, structured work.
  • The field is judged by measured accuracy on real, messy images, not by how a demo looks on a clean sample photo.

What is computer vision?

Computer vision is the field of building systems that extract meaning from images and video: what objects are present, where they are, what is happening, and how things change across frames. It sits at the intersection of image processing, geometry, and, since the early 2010s, deep learning, and it answers questions a plain digital image cannot answer on its own, a photograph is just a grid of numbers until a system interprets it [1].

Also asked as: how does computer vision work · how computer vision works · how to use computer vision · what is computer vision in ai · what is computer vision syndrome · what is computer vision how it helps the ai · what is computer vision in machine learning · what is computer vision engineer · what is computer vision class 6 · what is computer vision in simple words · what is computer vision with example · what is computer vision class 10 · what are computer vision projects · what are computer vision models

Also asked as: what is computer vision · computer vision definition · computer vision explained · what does computer vision do · computer vision meaning · introduction to computer vision · computer vision for beginners

FaceVision, one of the systems I built, does exactly this kind of interpretation for one narrow task, face detection and recognition, entirely in the browser [12]. This page is the wider field that task sits inside, from the classic techniques that came before deep learning to the vision-language models sitting alongside it now [11].

A camera captures light. Computer vision is the decades of work that turns that light into an answer to a question you actually care about. Pranjul Rathour

What are the core tasks in computer vision?

Classification answers "what is the main thing in this image", one label per image. Object detection answers "what is in this image and where", a label plus a bounding box for each object [5]. Segmentation answers "which exact pixels belong to which object", a precise outline rather than a box [6]. Tracking answers "where did this specific object go across frames of video". Each task answers a genuinely different question, and choosing the wrong one for a job, detection when you needed segmentation, is a common source of a system that technically works but does not fit the actual need.

Also asked as: why is deskewing so critical for ocr models when cnns can handle slight rotations in other vision tasks · which learning tasks do brains use to train themselves to see · are modular neural networks more effective than large, monolithic networks at any tasks

Also asked as: computer vision tasks · classification vs detection vs segmentation · what is object detection · what is image segmentation · types of computer vision tasks · computer vision task comparison

My longer page on object detection and YOLO specifically, one of the most widely used detection approaches, is written up separately [9].

How did computer vision work before deep learning?

Classic computer vision relied on hand-engineered features: edge detectors, corner detectors, and descriptors such as SIFT and HOG that a researcher designed to capture useful visual patterns, followed by a classical classifier trained on those extracted features [1]. This required real domain expertise to design well, and performance plateaued because the features themselves, not just the classifier, limited what the system could learn to recognise. It still works, and still runs today, especially where speed and low resource use matter more than squeezing out the last few percent of accuracy.

Also asked as: is computer vision deep learning · is computer vision part of deep learning · difference between computer vision and deep learning · why does face recognition not work on my iphone · why does face recognition not work sometimes · why does face recognition not work · how do i get my face recognition to work · how do face recognition cameras work · how does face recognition work on android · how does face recognition work on samsung · how does face recognition work on phone · how does face recognition work in the dark · how does face recognition work with twins · how does face recognition attendance system work

Also asked as: computer vision before deep learning · classic computer vision techniques · sift hog feature extraction · hand engineered features computer vision · opencv classic techniques · old computer vision methods

What changed with deep learning and CNNs?

Convolutional neural networks learn their own features directly from pixels during training, rather than relying on a human to design them, and AlexNet's 2012 ImageNet result showed this could beat hand-engineered approaches by a wide margin [2]. ResNet then showed very deep networks could be trained reliably using residual connections, pushing accuracy further [3]. This is the shift that made deep learning the default approach to almost all computer vision tasks within a few years: the model discovers what features matter, instead of a person guessing.

Also asked as: how can i implement continual (incremental) learning in a face-recognition model without retraining from scratch · how can we scale up the number of classes for deep learning after training a model · which model is better for incremental learning · do different camera angles affect the performance of the deep learning model · embedding quality of transfer learning model vs contrastive learning model · how to handle images of different sizes that are smaller than the input layer of a deep learning model · is it possible to do attribute, value extraction prediction model in machine learning · how to label overlapping objects for deep learning model training · are there any references or examples of line recognition on a chalkboard with a machine learning model · why are traditional ml models still used over deep neural networks · what are some use cases of few-shot learning · does opencv use machine learning

Also asked as: cnn for computer vision · how do convolutional neural networks work · alexnet resnet explained · deep learning vs classic computer vision · why cnns changed computer vision · history of computer vision deep learning

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

What are Vision Transformers, and how do they relate to CNNs?

A Vision Transformer splits an image into fixed-size patches, treats each patch like a word token, and processes the sequence with the same transformer architecture used for language, rather than the sliding convolutional filters a CNN uses [4]. At sufficient training data scale, Vision Transformers match or beat CNNs on many benchmarks, and they are also the backbone that most vision-language models use to turn an image into tokens a language model can read [11]. CNNs remain competitive, especially with less training data, and are often more efficient for real-time or on-device use.

Also asked as: do vision transformers handle arbitrary sequence lengths the same way as normal transformers · are yolopark transformers model kits · are vision transformers scale invariant like cnns

Also asked as: vision transformer vs cnn · what is a vision transformer · vit explained · how do vision transformers work · cnn vs transformer for images · are vision transformers better than cnn

How does computer vision relate to vision-language models?

Vision-language models sit one layer up: they use a vision component, often a Vision Transformer, to turn an image into tokens, then hand those tokens to a language model that can answer open-ended questions about the image in natural language [11]. Classic computer vision tasks, detection, segmentation, precise counting, remain more accurate and far cheaper for the narrow, structured questions they answer; vision-language models are better for open-ended understanding, reading, and reasoning about an image, at the cost of speed, precision on exact counts, and reliability without validation.

Also asked as: large language models explained · what is a large language model llm · why doesn't clip use a pretrained large language model as the text encoder · difference between computer vision and natural language processing · best computer vision ai models · deploying your ai/ml models: a practical guide from training to production · what are the state-of-the-art models for identifying objects in photos · how can realize the evaluation/validation of unsupervised models through unlabeled data · is there any research on models that make predictions by also taking into account the previous predictions · what models will you suggest to use in industrial anomaly detection and predictive analysis on live streamed data · how do ocr models work · can i train two stacked models end-to-end on different resolutions · how to evaluate sequence to sequence models · how and why do state-of-the-art models in medical segmentation differ from general segmentation models

Also asked as: computer vision vs vision language models · vlm vs computer vision · when to use vlm vs cnn · do vision language models replace computer vision · classic cv vs multimodal ai

My full comparison of classic vision models against multimodal LLMs, with the projects each fits, is on the vision-language models page [11].

Where is computer vision actually used today?

Manufacturing quality inspection, checking parts on a line far faster and more consistently than a human. Medical imaging, flagging regions in a scan for a radiologist's review. Autonomous vehicles and robotics, understanding a scene in real time. Retail, counting stock and tracking shelf state. Agriculture, identifying crop disease from a leaf photo. Security and access, face recognition and liveness detection, done carefully and with real privacy boundaries [12]. Document processing, reading structured fields from scanned forms. Every one of these is a narrow, measured task, not a general "understand everything" system.

Also asked as: what are the metrics to be used for unsupervised monocular depth estimation in computer vision · are information processing rules from gestalt psychology still used in computer vision today · what are the main algorithms used in computer vision · why are rnns used in some computer vision problems · how does ai actually work · what is clip model used for · what on-device ai benchmarks actually feel like

Also asked as: computer vision applications · computer vision use cases · where is computer vision used · computer vision in industry · real world computer vision examples · computer vision in healthcare

What mistakes do beginners make learning computer vision?

Jumping straight to a pretrained model without understanding what task it actually performs, then being confused when it doesn't do a different task well. Testing only on clean, well-lit sample images and being surprised when real photos, blur, bad lighting, occlusion, break the system. Treating accuracy on a benchmark dataset as a guarantee of accuracy on your own images, which it is not. Skipping the classic techniques entirely and never learning why a heavier deep model is sometimes the wrong tool for a fast, resource-constrained job.

Also asked as: how does computer science make · how much computer science make · how much computer science make a year · how much computer science make per hour · can reinforcement learning algorithms be applied to computer vision problems · how to make an ai model come alive · how to use opencv in python, make your hand invisible using opencv magic effect

Also asked as: computer vision beginner mistakes · common cv mistakes · computer vision learning path mistakes · why is my computer vision model not working

Computer vision interview questions

Define computer vision and name its four core tasks with the question each answers. Explain how CNNs changed the field relative to hand-engineered features. Compare CNNs and Vision Transformers. Explain how a vision-language model differs from a classic detector, and when you would choose each. Describe a computer vision project you built and what broke on real images versus clean samples.

Also asked as: computer vision interview questions · cv engineer interview · deep learning vision interview questions · image recognition interview questions

Where should I start?

Run a pretrained classifier and a pretrained detector on ten of your own real photos, not sample images, and note where each one is confidently wrong. That afternoon teaches more about the field's actual limits than a week of reading. For a hands-on session on computer vision for students, from classic techniques to vision-language models, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Szeliski, Computer Vision: Algorithms and Applications (2nd ed., 2022)szeliski.org
  2. Krizhevsky, Sutskever, Hinton, ImageNet Classification with Deep Convolutional Neural Networks (AlexNet, 2012)papers.nips.cc
  3. He et al., Deep Residual Learning for Image Recognition (ResNet, 2015)arxiv.org
  4. Dosovitskiy et al., An Image is Worth 16x16 Words (Vision Transformer, 2020)arxiv.org
  5. Redmon et al., You Only Look Once: Unified, Real-Time Object Detection (YOLO, 2016)arxiv.org
  6. Kirillov et al., Segment Anything (2023)arxiv.org
  7. OpenCV documentationdocs.opencv.org
  8. ImageNet, the dataset that accelerated the fieldimage-net.org
  9. What is object detection (YOLO), Pranjul Rathourpranjulrathour.github.io
  10. How does face recognition work, Pranjul Rathourpranjulrathour.github.io
  11. What are vision-language models, Pranjul Rathourpranjulrathour.github.io
  12. FaceVision source code, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Vision, speech & OCR

All Vision, speech & OCR guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur