What is edge AI and on-device AI? Running models in the browser and on phones with ONNX Runtime, WebGPU and quantization, from building a browser-only face recognition system

Edge AI runs models where the data is, on phones, laptops, browsers and small devices, instead of in the cloud. Why it matters for privacy, latency and cost, how ONNX Runtime Web and WebGPU make browser inference practical, what fits on a device, how to export and quantize a model, and how FaceVision keeps every camera frame on the user's machine.

Presenting BrandHive
Presenting BrandHive

Key takeaways

  • Edge AI moves inference to the device so raw data never leaves it: lower latency, no per-call cost, offline operation and a simple privacy story.
  • ONNX is the interchange format; ONNX Runtime runs the same model on servers, phones and in the browser through WebAssembly or WebGPU.
  • Small vision, speech and embedding models fit comfortably; small language models are possible; large models still need a server.
  • Quantize and prune before shipping to a device, then measure accuracy on the device, not in the training notebook.
  • Design the split: compute on the device, store only what you must on the server, and make that boundary explicit to users.

What is edge AI?

Edge AI is running machine learning inference on the device that has the data, a phone, a laptop, a browser tab, a camera, a sensor board, instead of sending the data to a server. The model is downloaded once and runs locally, so predictions arrive in milliseconds, work offline, cost nothing per call, and the raw input, a face, a voice, a document, never leaves the device. On-device AI, browser AI and TinyML are the same idea at different scales.

Also asked as: what is edge ai · edge ai meaning · edge ai explained · what is on device ai · on device ai vs cloud ai · edge ai vs cloud ai · edge computing ai · what is browser ai · tinyml vs edge ai · ai vs edge ai

I built FaceVision in 2026 as a browser-only face detection, recognition and liveness system: the camera frames are processed in the user's browser with ONNX Runtime Web, only a numeric embedding is sent to the server, and no image is ever stored [11]. That architecture is the concrete version of every argument on this page.

The most effective privacy feature I have ever shipped was not encryption or a policy. It was never having the data in the first place. Pranjul Rathour, from building FaceVision

Why run AI on the device instead of the cloud?

Four reasons, and each alone can justify it. Privacy: raw data stays with the user, which simplifies compliance and trust. Latency: no network round trip, so camera and audio features feel instant. Cost: no per-inference server bill, which matters when usage grows. Availability: it works offline and on bad connections, which in India is not an edge case. The trade-offs are model size, device variability and the work of exporting and testing.

Also asked as: benefits of edge ai · why edge ai · advantages of on device ai · edge ai privacy · edge ai latency · edge ai cost savings · edge ai offline · disadvantages of edge ai · edge ai challenges

What is ONNX, and what is ONNX Runtime?

ONNX is an open file format for trained models that most frameworks can export to, so a model trained in PyTorch can run in a different runtime [1]. ONNX Runtime is Microsoft's inference engine for that format, with builds for servers, mobile and the web; ONNX Runtime Web runs models in the browser using WebAssembly on the CPU or WebGPU on the graphics card [2]. Together they are the most portable path from a training notebook to a device.

Also asked as: what is onnx · what is onnx runtime · onnx runtime web · onnx vs pytorch · how to convert pytorch to onnx · onnx runtime web tutorial · onnx model in browser · onnx runtime mobile · onnx explained

My step-by-step guide to ONNX Runtime Web, with the preprocessing pitfalls, is on the portfolio [13].

What is WebGPU, and does browser AI need it?

WebGPU is the web standard that gives browser code access to the graphics card for general computation, not just rendering [3]. For AI in the browser it means models run several times faster than on WebAssembly alone, which turns real-time camera and audio models from a demo into a product. It does not replace WebAssembly: you ship both and fall back to the CPU path on browsers or devices without WebGPU support. Small models run acceptably on WebAssembly; larger ones need WebGPU.

Also asked as: what is webgpu · webgpu ai · webgpu vs webassembly for ai · webgpu machine learning · webgpu llm · browser gpu inference · webgpu support · webnn

What models can run on a device or in a browser?

Comfortably: face detection and recognition, object detection with MobileNet-class or small YOLO models, image classification, OCR for clean text, speech recognition with small Whisper variants, text embeddings, and classifiers. Possibly: small language models of one to three billion parameters, quantized, on a laptop or a recent phone, with llama.cpp or a browser runtime. Not yet: large language models, high-resolution generative image models, or anything that needs tens of gigabytes.

Also asked as: what models can run on device · small models for edge · llm on device · run llm in browser · on device speech recognition · on device face recognition · edge ai models list · best models for mobile · small language models on phone

MobileNets set the pattern for models designed for devices [8]; Transformers.js brings Hugging Face models to the browser [4]; MediaPipe ships ready-made vision and audio pipelines for web and mobile [7].

How do I make a model small enough for a device?

Quantize to 8-bit or 4-bit integers, which cuts size and speeds inference on most hardware with a small accuracy loss [9]. Prune or distil into a smaller architecture when quantization is not enough. Fix input sizes and remove training-only operations at export. Then measure accuracy and latency on the actual target devices, because a model that passes in a notebook can fail on a low-end phone's memory or a browser's WebAssembly limits.

Also asked as: model optimization for edge · quantization for mobile · model compression edge ai · reduce model size · pruning and distillation · int8 quantization onnx · how to make model run faster on mobile · model size for browser

Presenting BrandHive
Presenting BrandHive

How does FaceVision keep faces on the device?

The browser captures camera frames and runs three ONNX models locally: a face detector, a liveness check, and a recognition network that produces a 512-number embedding. Only that embedding is posted to a FastAPI backend, which stores it in PostgreSQL per user and compares it on verification. No image is transmitted or stored anywhere. If the server is breached, the attacker has vectors that cannot be turned back into faces. The design is on GitHub [11], and the reasoning about privacy is written up separately [12].

Also asked as: facevision architecture · browser face recognition privacy · face recognition without uploading images · embeddings only storage · privacy by design face recognition · on device biometrics · client side inference example

What are the platform options beyond the browser?

On Android, LiteRT, formerly TensorFlow Lite, and ONNX Runtime Mobile [5][2]. On iOS, Core ML, with converters from PyTorch and ONNX [6]. On laptops and small servers, ONNX Runtime and llama.cpp for language models [10]. For ready-made vision and audio, MediaPipe on all of them [7]. The browser is the widest reach with the least installation; native platforms give more performance and background execution.

Also asked as: tensorflow lite vs onnx · core ml vs onnx · edge ai frameworks · on device ai android · on device ai ios · mediapipe tutorial · litert · best framework for edge ai · edge ai hardware

What are common edge AI use cases?

Face and document capture for onboarding, attendance and KYC with privacy by design. Real-time translation and captions. Offline speech recognition for field workers and farmers. Quality inspection on factory cameras. Health and fitness sensing on wearables. Smart-camera detection without a cloud subscription. Browser tools that process a user's data without ever seeing it, which is how 13 of the 15 free tools at tools.scult.in work.

Also asked as: edge ai use cases · edge ai applications · edge ai examples · on device ai examples · edge ai in healthcare · edge ai in agriculture · edge ai in manufacturing · browser ai use cases

Edge AI interview questions

Explain the trade-offs between on-device and cloud inference. Explain what ONNX and ONNX Runtime are for. Describe how you would take a PyTorch model to a browser and what could go wrong in preprocessing. Explain quantization's effect on size and accuracy. Design a privacy-preserving face verification system. Each answer is above, and a browser demo you built is the strongest evidence.

Also asked as: edge ai interview questions · on device ml interview · onnx interview questions · mobile ml interview · embedded ai interview questions

Where should I start?

Export a small image classifier to ONNX, load it with ONNX Runtime Web in a plain HTML page, and classify a webcam frame without a server. Then quantize it and measure the speed difference. That afternoon is the whole field in miniature. For a hands-on session on privacy-first vision systems at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. ONNX: Open Neural Network Exchangeonnx.ai
  2. ONNX Runtime Web documentationonnxruntime.ai
  3. WebGPU specification, W3Cw3.org
  4. Transformers.js: run Hugging Face models in the browserhuggingface.co
  5. TensorFlow Lite / LiteRT documentationai.google.dev
  6. Apple Core ML documentationdeveloper.apple.com
  7. MediaPipe Solutions, Googleai.google.dev
  8. Howard et al., MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications (2017)arxiv.org
  9. Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (2017)arxiv.org
  10. llama.cpp: LLM inference on CPUs and edge devicesgithub.com
  11. FaceVision source code, Pranjul Rathourgithub.com
  12. On-device AI in the browser and privacy, Pranjul Rathourpranjulrathour.scult.in
  13. ONNX Runtime Web browser ML guide, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Vision, speech & OCR

All Vision, speech & OCR guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur