Traffic route - WebGPU - Local AI

Can my browser run local AI?

A friendly compatibility desk for visitors asking whether WebLLM, browser AI, LM Studio, Ollama, llama.cpp, or small local models are realistic on their device. Start with a live browser gate, then compare public local AI benchmark rows for Apple M2, RTX 4060, RTX 4090, Snapdragon X Elite, and other hardware.

Browser-firstNo upload neededEstimated fit, not a guaranteeDesigned for guides + Atlas

AI compatibility passport

This is the user-facing hook: answer the question first, then show the table.

Your recommended path

Next action cards

What users need to understand

WebGPU yesOffer tiny browser demos first. Mention model download size before any AI chat.
WebGPU noSuggest LM Studio, Ollama, or regular browser tools. Do not make them feel blocked.
MobileUse Lite mode. Avoid large downloads. Good for reading guides and tiny utilities.
Desktop GPUCan explore Story mode, benchmarks, and Skyling local chat experiments later.

Runtime decision matrix

Help visitors pick a path without reading dozens of forum comments.

WebLLM

Browser AI experiments, tiny demos, local-first curiosity.

WebGPUDownload size

Transformers.js

Embeddings, classification, small browser ML helpers, route memory later.

BrowserUtility ML

LM Studio

Friendly desktop local chat, stronger models, visual controls.

DesktopBeginner-friendly

Ollama / llama.cpp

Developer-friendly local models, CLI workflows, app integration.

DesktopTechnical

Model tier table

Starter compatibility tiers. Use the public benchmark table below for verified llama.cpp benchmark rows.

TierTypical sizeBrowser pathDesktop pathRAMVRAMBest forLast checkedConfidenceSource

Local AI benchmark table

Local AI speed depends heavily on the model, quantization, runtime, backend, driver, OS, cooling, runtime version, and memory bandwidth. This benchmark table collects public reference results so you can compare rough hardware tiers before choosing a local AI setup.

How to read this: These rows use public benchmark sources where available. PP t/s means prompt processing speed for ingesting long prompts. TG t/s means text generation speed users usually feel during replies. Treat every row as a reference point, not a guarantee. 5 public rows shown, full table loads with JavaScript
HardwareRuntime / BackendModel / QuantPP t/sTG t/sTest conditionConfidence / CheckedSource
Apple M2
Apple laptop baseline
llama.cpp
Metal
Llama 2 7B
Q4_0
179.5721.91pp512 / tg128verified
2026-06-11
Source
llama.cpp Apple Silicon M-series benchmark discussion #4167
Apple M2 Pro
Apple laptop midrange
llama.cpp
Metal
Llama 2 7B
Q4_0
294.2437.87pp512 / tg128verified
2026-06-11
Source
llama.cpp Apple Silicon M-series benchmark discussion #4167
Apple M2 Max
Apple desktop or high-end laptop
llama.cpp
Metal
Llama 2 7B
Q4_0
671.3165.95pp512 / tg128verified
2026-06-11
Source
llama.cpp Apple Silicon M-series benchmark discussion #4167
Apple M2 Ultra
Apple desktop workstation
llama.cpp
Metal
Llama 2 7B
Q4_0
1238.4894.27pp512 / tg128verified
2026-06-11
Source
llama.cpp Apple Silicon M-series benchmark discussion #4167
NVIDIA GeForce RTX 3060 12GB
Entry CUDA desktop
llama.cpp
CUDA
Llama 2 7B
Q4_0
2407.6776.92pp512 / tg128, Flash Attention enabledverified
2026-06-11
Source
llama.cpp CUDA scoreboard discussion #15013

Includes public local AI benchmark rows for Apple M2 local AI, RTX 4060 local AI, RTX 4090 local AI, Snapdragon X Elite local AI, WebGPU compatibility planning, llama.cpp benchmark comparison, and tokens per second reference checks.

Community benchmark template

Later, this becomes the data engine that makes the Atlas trustworthy.

Confidence: verifiedVerified data points are benchmarked locally on target hardware under active testing.
Confidence: proposedProposed data points are projected estimates from official model documentation or community cards.
Confidence: starterStarter compatibility rows are initial estimates until verified with specific model builds.
No hype claimsNever promise GPT-level quality from browser models. Always note VRAM/RAM constraints.