On-Device and Browser AI: WebGPU, WebNN, Gemini Nano, Apple Foundation Models, and Edge Runtimes
How to design AI features that run locally in the browser or on device, with privacy, latency, model-size, and fallback tradeoffs.
Built for: Frontend teams, product engineers, privacy-sensitive apps, offline-first products, browser extensions, and mobile developers.
key takeaways
- +On-device AI is strongest for privacy, latency, offline use, and high-frequency small tasks.
- +Browser AI requires graceful degradation: capability detection, cloud fallback, and downloadable model budgets.
- +WebGPU is practical today for many local inference paths; WebNN is the standards path for hardware acceleration.
- +Do not ship giant local models unless the user benefit is obvious and the download/storage cost is acceptable.
routing graph - local first, cloud when needed
Where local AI wins
Local AI is not about replacing every cloud model. It is about moving the right tasks close to the user: autocomplete, summarization of local files, smart search, translation, classification, image labeling, privacy-preserving drafts, and offline copilots.
The user experience benefit is immediate: no network round trip, fewer privacy concerns, and features that continue working when connectivity is poor.
- -Best: short text transforms, embeddings, OCR helpers, local search, lightweight vision, autocomplete.
- -Risky: long reasoning tasks, large multimodal analysis, heavy generation, workflows requiring current web data.
- -Hybrid: local prefiltering or extraction, cloud reasoning only when needed.
Runtime choices
The browser AI stack is split across standards, vendor APIs, and JavaScript runtimes. WebGPU gives general-purpose GPU acceleration. WebNN aims to expose neural-network acceleration across hardware. Transformers.js makes Hugging Face model inference practical in JavaScript. Platform APIs such as Gemini Nano in Chrome and Apple Foundation Models on Apple platforms expose built-in models where available.
Capability-first routing
typescript
async function chooseAiRuntime() {
if ("ai" in globalThis && "summarizer" in (globalThis as any).ai) {
return "browser-built-in";
}
if ("gpu" in navigator) {
return "webgpu-transformers";
}
return "cloud-fallback";
}Product constraints
The hard parts are not only model quality. They are package size, warmup time, battery, memory pressure, model caching, privacy messaging, and fallback behavior when a device is old or locked down by enterprise policy.
- -Use model lazy-loading behind explicit user intent.
- -Cache model files with versioned names and clear storage limits.
- -Benchmark low-end devices, not only the developer laptop.
- -Keep a server fallback for unsupported browsers and high-complexity requests.
- -Explain local processing in product copy where privacy is a feature.
Evaluation
Local model evals should include UX metrics. A smaller local model may be objectively weaker but better for a given interaction if it responds instantly and keeps data private.
- -Quality: task accuracy, hallucination rate, formatting success.
- -Performance: cold start, first token, total latency, memory, battery.
- -Compatibility: browser, OS, GPU, enterprise restrictions.
- -Fallback: whether the cloud path produces equivalent outputs when local runtime is unavailable.
Sources and further reading
Transformers.js documentation
Hugging Face JavaScript inference library for browser and Node.js use.
WebGPU specification
W3C specification for GPU computation and rendering in the web platform.
WebNN specification
W3C neural-network API draft for hardware-accelerated web ML.
Chrome built-in AI docs
Chrome documentation for built-in AI APIs and Gemini Nano availability.
Apple Foundation Models framework
Apple documentation for on-device foundation model integration.