How does face recognition work? Detection, embeddings, thresholds and liveness, from building FaceVision

Face detection finds faces; face recognition turns each face into an embedding and compares distances against a threshold; liveness detection stops a photo from passing. How each step works, why systems fail on twins and low light, how accurate they are, the privacy rules, and how I built a browser-only version that never uploads a face.

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Key takeaways

  • Detection answers "where is a face"; recognition answers "whose face"; liveness answers "is this a live person or a photo". They are three separate models.
  • Recognition compares embedding vectors. The match threshold trades false accepts against false rejects, and you must choose it from your own data.
  • Photos, videos and masks defeat naive systems. Liveness checks, active or passive, are mandatory for anything that unlocks or pays.
  • Accuracy varies by lighting, camera, pose and demographic group. Test on your real users before you trust a benchmark number.
  • You can run the whole pipeline in the browser with ONNX Runtime Web and store only embeddings, so no face image ever reaches your server.

How does face recognition work?

Face recognition works in four steps. A detector finds each face in the image and its landmarks. The face is aligned and cropped to a standard size. A recognition network turns the crop into an embedding, a vector of a few hundred numbers in which faces of the same person land close together. Finally the system compares that embedding with stored ones and declares a match when the distance falls under a threshold. Everything else, liveness, enrolment, privacy, is engineering around those four steps.

Also asked as: how does face recognition work · how face recognition works · how facial recognition works · face recognition explained · what is face recognition · face recognition technology · how does facial recognition technology work · how face recognition works step by step

The embedding idea comes from FaceNet, which trained a network so that the Euclidean distance between two face vectors directly measured how likely they were the same person [1]. ArcFace sharpened it with an angular margin loss that spreads identities further apart, and it is the recipe behind most open models today [2].

I built FaceVision in 2026 as a browser-only detection, recognition and liveness system: the camera frames never leave the device, only embeddings are sent to a FastAPI and PostgreSQL backend, and no image is ever stored [13]. This page is what that build taught me, with the papers underneath each step.

The moment you decide never to store a face image, half your privacy problems disappear and the other half become simple. Store the vector, throw away the pixels. Pranjul Rathour, from building FaceVision

What is the difference between face detection and face recognition?

Face detection finds faces: it returns bounding boxes and usually landmarks for eyes, nose and mouth, for every face in a frame, and it does not know or care who they are. Face recognition takes one detected face and identifies or verifies the person. Detection is the first stage of recognition, but a security camera that counts people, or a phone that focuses on a face, needs only detection.

Also asked as: what is face detection · face detection vs face recognition · difference between face detection and face recognition · face detection explained · face recognition vs facial recognition · what is facial recognition

MTCNN was the standard detector for years [3]; RetinaFace improved accuracy on small and turned faces and remains a strong default [4]. Google's MediaPipe ships a fast detector that runs on phones and in browsers [11].

What is a face embedding, and how are faces compared?

A face embedding is a fixed-length vector, often 128 or 512 numbers, produced by a network trained so that the same person's faces cluster together and different people's faces sit apart. Comparison is a distance: cosine similarity or Euclidean distance between two vectors. If the distance is below a threshold, the system calls it a match. The threshold is not part of the model; it is a decision you make from your own error data.

Also asked as: what is face embedding · face recognition embeddings explained · how are faces compared · face similarity score · cosine similarity face recognition · face recognition threshold · face matching algorithm

The embedding is why you never need to store the face image. A vector cannot be turned back into a photograph in any useful way, it is far smaller, and it can be compared in a database index like any other vector. FaceVision stores exactly this: the vector, the user id, the timestamp, nothing visual [13].

How do I choose the match threshold?

Collect pairs of embeddings from your real users and cameras, labelled same-person or different-person, then plot the two error rates as the threshold moves: false accepts, a stranger let in, and false rejects, the right person turned away. Choose the threshold that meets your risk. A door lock wants false accepts near zero and tolerates retries; a class attendance system tolerates a few false accepts and hates queues. Never ship a vendor's default threshold blind.

Also asked as: face recognition threshold · false accept rate vs false reject rate · how accurate is face recognition · face recognition accuracy · far vs frr · equal error rate face recognition · how to improve face recognition accuracy

NIST's ongoing Face Recognition Technology Evaluation publishes false match and false non-match rates for hundreds of algorithms under controlled conditions, and it is the reference for what state-of-the-art accuracy looks like [5]. It also shows how far the same algorithm's numbers move between a mugshot-quality image and a webcam frame, which is why your own data matters more than the leaderboard.

Why does face recognition fail?

For a short list of reasons that repeat everywhere: low light and backlight, motion blur, extreme pose, masks and sunglasses, low-resolution or heavily compressed cameras, faces too small in the frame, enrolment photos that look nothing like daily conditions, and a threshold tuned on a different population. Twins and close relatives are a real limit for some models. Most "the AI is broken" tickets are a camera or lighting problem.

Also asked as: why face recognition not working · face recognition not working · why does face recognition fail · face recognition problems · limitations of face recognition · can face recognition be fooled · can face recognition identify twins · can face recognition distinguish twins · face recognition in low light

On twins: FaceVision, like every embedding system I have tested, sometimes cannot separate identical twins at a threshold that still rejects strangers. For high-stakes use, that is a reason to add a second factor, not a reason to loosen the threshold.

Is face recognition biased, and how accurate is it across groups?

Accuracy differs across demographic groups for most algorithms. NIST's 2019 demographic study measured higher false positive rates for some groups than others, with the size of the gap depending strongly on the algorithm [6]. The gap has narrowed in recent evaluations but has not vanished. The practical response is to test on your own user population, report the error rates per group, and choose a threshold and a fallback that protect the worst-served group, not the average.

Also asked as: is face recognition biased · face recognition bias · facial recognition accuracy by race · face recognition demographic bias · is facial recognition accurate · face recognition fairness

For a college attendance system in Kanpur, the relevant test set is your own students under your own classroom lights, not a benchmark shot in a studio elsewhere.

What is liveness detection, and why do I need it?

Liveness detection, also called presentation attack detection, checks that the face in front of the camera is a live person and not a printed photo, a phone screen showing a photo or video, or a mask. Without it, any recognition system that unlocks or pays can be defeated by a picture from social media. Active liveness asks the user to do something, blink, turn, follow a dot. Passive liveness analyses a single frame or short clip for texture, depth cues and screen artefacts.

Also asked as: what is liveness detection · liveness detection face recognition · active vs passive liveness detection · how does liveness detection work · anti spoofing face recognition · can face recognition be fooled by photo · face liveness check · presentation attack detection

ISO/IEC 30107 defines the framework and terminology for presentation attack detection and the metrics vendors quote [7]. Texture analysis, the observation that a re-photographed face carries colour and frequency artefacts a live face does not, is one of the classic passive approaches [8]. FaceVision uses a passive check on the device before any embedding is computed, so a photo held to the camera is rejected before it can even be compared [13].

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Can I run face recognition in the browser, without a server seeing the face?

Yes. ONNX Runtime Web runs detection, recognition and liveness models in the browser using WebAssembly or WebGPU [10]. The camera frame stays on the device, the models produce an embedding locally, and only the embedding travels to your API for comparison and storage. It is slower than a GPU server for large galleries, but for verification and small identification sets it is fast enough, and it removes the biggest privacy and compliance burden: you never possess a face image.

Also asked as: face recognition in browser · face recognition javascript · onnx runtime web face recognition · on device face recognition · face recognition without server · client side face recognition · what is onnx runtime · face recognition privacy

The architecture of FaceVision, which is the concrete version of this answer:

Models are exported to ONNX from PyTorch or taken from open toolkits such as InsightFace, which publishes ArcFace-family recognition models and RetinaFace detectors [9]. Quantising them shrinks download size and speeds inference on phones, at a small accuracy cost you should measure.

How do I build a face recognition attendance system?

Enrol each person with three to five embeddings taken under the real classroom conditions, on the same kind of camera. At check-in, detect, run liveness, embed and compare against the enrolled set for that class only, which keeps the gallery small and the false-match rate low. Log the score, not the image. Add a fallback for rejections, such as a one-time code, so a bad lighting day does not stop a class. Publish the privacy notice before the first day.

Also asked as: face recognition attendance system · how to make face recognition attendance system · face recognition attendance system python · best face recognition attendance system · face recognition attendance system project · face attendance system in college · attendance system using face recognition

Attendance is the most common student project in this space and the one most often built wrong: images saved to a folder, a global gallery of thousands, no liveness, a threshold copied from a tutorial. Built right, with per-class galleries, on-device liveness and embeddings-only storage, it is also a strong portfolio piece, because you can explain each decision.

Facial data is biometric personal data, and India's Digital Personal Data Protection Act, 2023, sets the framework for processing personal data with notice, consent, purpose limitation and security obligations [12]. Rules and enforcement continue to develop, so treat this as the direction, not legal advice. The engineering answer is independent of jurisdiction: collect consent, tell people what is stored and for how long, store embeddings rather than images, delete on request, and restrict who can run a search.

Also asked as: is face recognition legal in india · face recognition privacy · face recognition data protection · face recognition consent · facial recognition law india · face recognition gdpr · face recognition ethics

I wrote a longer note on what changed in my own practice after FaceVision, and most of it is the list above [14].

Which library or model should I use for face recognition?

For a Python prototype: InsightFace for detection and ArcFace-family recognition, or the face_recognition library for simplicity at lower accuracy [9]. For phones and browsers: MediaPipe for detection [11] and an ONNX-exported ArcFace model in ONNX Runtime Web for recognition [10]. For anything that unlocks or pays: add a liveness model and test it against printed photos and phone screens yourself before you believe its numbers.

Also asked as: best face recognition model · best face recognition library python · best face detection model · face recognition python · opencv face recognition · deepface vs insightface · face recognition api · best face recognition algorithm · mediapipe face detection

What are common face recognition interview questions?

Explain detection versus recognition versus verification versus identification. Explain what an embedding is and why distance measures similarity. Explain how you would set a threshold and what false accept and false reject mean. Explain what liveness detection defends against. Explain how you would design the system so no face image is stored. If you can answer these from a project you built, you are ahead of most candidates.

Also asked as: face recognition interview questions · computer vision interview questions face recognition · biometrics interview questions · face recognition project explanation

Where should I start with a face recognition project?

Build verification first, not identification: enrol yourself and two friends, run detection and embedding in the browser with ONNX Runtime Web, store only embeddings, and plot your own false accept and false reject curve from a hundred attempts. Then add passive liveness and try to fool it with a photo of yourself. That project teaches every concept on this page and is exactly the shape of FaceVision.

Also asked as: face recognition project · face recognition project ideas · face recognition tutorial · learn face recognition · face recognition course · face recognition for beginners · how to start with computer vision

If your college wants a hands-on session on computer vision that respects privacy from the first line of code, that is one of the talks I offer. Email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Schroff, Kalenichenko & Philbin, FaceNet: A Unified Embedding for Face Recognition and Clustering (2015)arxiv.org
  2. Deng et al., ArcFace: Additive Angular Margin Loss for Deep Face Recognition (2019)arxiv.org
  3. Zhang et al., Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks (MTCNN, 2016)arxiv.org
  4. Deng et al., RetinaFace: Single-stage Dense Face Localisation in the Wild (2019)arxiv.org
  5. NIST Face Recognition Technology Evaluation (FRTE), ongoingpages.nist.gov
  6. Grother, Ngan & Hanaoka, Face Recognition Vendor Test Part 3: Demographic Effects (NIST IR 8280, 2019)doi.org
  7. ISO/IEC 30107-1: Biometric presentation attack detection, frameworkiso.org
  8. Boulkenafet, Komulainen & Hadid, Face Spoofing Detection Using Colour Texture Analysis (2016)ieeexplore.ieee.org
  9. InsightFace: open-source face analysis toolkitgithub.com
  10. ONNX Runtime Web documentationonnxruntime.ai
  11. MediaPipe Face Detection, Googleai.google.dev
  12. Digital Personal Data Protection Act, 2023, India (Ministry of Electronics and IT)meity.gov.in
  13. FaceVision source code, Pranjul Rathourgithub.com
  14. Privacy by design for AI apps: what I do differently after FaceVision, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Vision, speech & OCR

All Vision, speech & OCR guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur