Key takeaways
- Edge AI moves inference to the device so raw data never leaves it: lower latency, no per-call cost, offline operation and a simple privacy story.
- ONNX is the interchange format; ONNX Runtime runs the same model on servers, phones and in the browser through WebAssembly or WebGPU.
- Small vision, speech and embedding models fit comfortably; small language models are possible; large models still need a server.
- Quantize and prune before shipping to a device, then measure accuracy on the device, not in the training notebook.
- Design the split: compute on the device, store only what you must on the server, and make that boundary explicit to users.
What is edge AI?
Edge AI is running machine learning inference on the device that has the data, a phone, a laptop, a browser tab, a camera, a sensor board, instead of sending the data to a server. The model is downloaded once and runs locally, so predictions arrive in milliseconds, work offline, cost nothing per call, and the raw input, a face, a voice, a document, never leaves the device. On-device AI, browser AI and TinyML are the same idea at different scales.
Also asked as: what is edge ai · edge ai meaning · edge ai explained · what is on device ai · on device ai vs cloud ai · edge ai vs cloud ai · edge computing ai · what is browser ai · tinyml vs edge ai · ai vs edge ai
I built FaceVision in 2026 as a browser-only face detection, recognition and liveness system: the camera frames are processed in the user's browser with ONNX Runtime Web, only a numeric embedding is sent to the server, and no image is ever stored [11]. That architecture is the concrete version of every argument on this page.
The most effective privacy feature I have ever shipped was not encryption or a policy. It was never having the data in the first place. Pranjul Rathour, from building FaceVision
Why run AI on the device instead of the cloud?
Four reasons, and each alone can justify it. Privacy: raw data stays with the user, which simplifies compliance and trust. Latency: no network round trip, so camera and audio features feel instant. Cost: no per-inference server bill, which matters when usage grows. Availability: it works offline and on bad connections, which in India is not an edge case. The trade-offs are model size, device variability and the work of exporting and testing.
Also asked as: benefits of edge ai · why edge ai · advantages of on device ai · edge ai privacy · edge ai latency · edge ai cost savings · edge ai offline · disadvantages of edge ai · edge ai challenges
What is ONNX, and what is ONNX Runtime?
ONNX is an open file format for trained models that most frameworks can export to, so a model trained in PyTorch can run in a different runtime [1]. ONNX Runtime is Microsoft's inference engine for that format, with builds for servers, mobile and the web; ONNX Runtime Web runs models in the browser using WebAssembly on the CPU or WebGPU on the graphics card [2]. Together they are the most portable path from a training notebook to a device.
Also asked as: what is onnx · what is onnx runtime · onnx runtime web · onnx vs pytorch · how to convert pytorch to onnx · onnx runtime web tutorial · onnx model in browser · onnx runtime mobile · onnx explained
My step-by-step guide to ONNX Runtime Web, with the preprocessing pitfalls, is on the portfolio [13].
What is WebGPU, and does browser AI need it?
WebGPU is the web standard that gives browser code access to the graphics card for general computation, not just rendering [3]. For AI in the browser it means models run several times faster than on WebAssembly alone, which turns real-time camera and audio models from a demo into a product. It does not replace WebAssembly: you ship both and fall back to the CPU path on browsers or devices without WebGPU support. Small models run acceptably on WebAssembly; larger ones need WebGPU.
Also asked as: what is webgpu · webgpu ai · webgpu vs webassembly for ai · webgpu machine learning · webgpu llm · browser gpu inference · webgpu support · webnn
What models can run on a device or in a browser?
Comfortably: face detection and recognition, object detection with MobileNet-class or small YOLO models, image classification, OCR for clean text, speech recognition with small Whisper variants, text embeddings, and classifiers. Possibly: small language models of one to three billion parameters, quantized, on a laptop or a recent phone, with llama.cpp or a browser runtime. Not yet: large language models, high-resolution generative image models, or anything that needs tens of gigabytes.
Also asked as: what models can run on device · small models for edge · llm on device · run llm in browser · on device speech recognition · on device face recognition · edge ai models list · best models for mobile · small language models on phone
MobileNets set the pattern for models designed for devices [8]; Transformers.js brings Hugging Face models to the browser [4]; MediaPipe ships ready-made vision and audio pipelines for web and mobile [7].
How do I make a model small enough for a device?
Quantize to 8-bit or 4-bit integers, which cuts size and speeds inference on most hardware with a small accuracy loss [9]. Prune or distil into a smaller architecture when quantization is not enough. Fix input sizes and remove training-only operations at export. Then measure accuracy and latency on the actual target devices, because a model that passes in a notebook can fail on a low-end phone's memory or a browser's WebAssembly limits.
Also asked as: model optimization for edge · quantization for mobile · model compression edge ai · reduce model size · pruning and distillation · int8 quantization onnx · how to make model run faster on mobile · model size for browser

How does FaceVision keep faces on the device?
The browser captures camera frames and runs three ONNX models locally: a face detector, a liveness check, and a recognition network that produces a 512-number embedding. Only that embedding is posted to a FastAPI backend, which stores it in PostgreSQL per user and compares it on verification. No image is transmitted or stored anywhere. If the server is breached, the attacker has vectors that cannot be turned back into faces. The design is on GitHub [11], and the reasoning about privacy is written up separately [12].
Also asked as: facevision architecture · browser face recognition privacy · face recognition without uploading images · embeddings only storage · privacy by design face recognition · on device biometrics · client side inference example
What are the platform options beyond the browser?
On Android, LiteRT, formerly TensorFlow Lite, and ONNX Runtime Mobile [5][2]. On iOS, Core ML, with converters from PyTorch and ONNX [6]. On laptops and small servers, ONNX Runtime and llama.cpp for language models [10]. For ready-made vision and audio, MediaPipe on all of them [7]. The browser is the widest reach with the least installation; native platforms give more performance and background execution.
Also asked as: tensorflow lite vs onnx · core ml vs onnx · edge ai frameworks · on device ai android · on device ai ios · mediapipe tutorial · litert · best framework for edge ai · edge ai hardware
What are common edge AI use cases?
Face and document capture for onboarding, attendance and KYC with privacy by design. Real-time translation and captions. Offline speech recognition for field workers and farmers. Quality inspection on factory cameras. Health and fitness sensing on wearables. Smart-camera detection without a cloud subscription. Browser tools that process a user's data without ever seeing it, which is how 13 of the 15 free tools at tools.scult.in work.
Also asked as: edge ai use cases · edge ai applications · edge ai examples · on device ai examples · edge ai in healthcare · edge ai in agriculture · edge ai in manufacturing · browser ai use cases
Edge AI interview questions
Explain the trade-offs between on-device and cloud inference. Explain what ONNX and ONNX Runtime are for. Describe how you would take a PyTorch model to a browser and what could go wrong in preprocessing. Explain quantization's effect on size and accuracy. Design a privacy-preserving face verification system. Each answer is above, and a browser demo you built is the strongest evidence.
Also asked as: edge ai interview questions · on device ml interview · onnx interview questions · mobile ml interview · embedded ai interview questions
Where should I start?
Export a small image classifier to ONNX, load it with ONNX Runtime Web in a plain HTML page, and classify a webcam frame without a server. Then quantize it and measure the speed difference. That afternoon is the whole field in miniature. For a hands-on session on privacy-first vision systems at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.
Sources
- ONNX: Open Neural Network Exchangeonnx.ai
- ONNX Runtime Web documentationonnxruntime.ai
- WebGPU specification, W3Cw3.org
- Transformers.js: run Hugging Face models in the browserhuggingface.co
- TensorFlow Lite / LiteRT documentationai.google.dev
- Apple Core ML documentationdeveloper.apple.com
- MediaPipe Solutions, Googleai.google.dev
- Howard et al., MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications (2017)arxiv.org
- Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (2017)arxiv.org
- llama.cpp: LLM inference on CPUs and edge devicesgithub.com
- FaceVision source code, Pranjul Rathourgithub.com
- On-device AI in the browser and privacy, Pranjul Rathourpranjulrathour.scult.in
- ONNX Runtime Web browser ML guide, Pranjul Rathourpranjulrathour.scult.in




