Can my browser run local AI?
A friendly compatibility desk for visitors asking whether WebLLM, browser AI, LM Studio, Ollama, llama.cpp, or small local models are realistic on their device. Start with a live browser gate, then compare public local AI benchmark rows for Apple M2, RTX 4060, RTX 4090, Snapdragon X Elite, and other hardware.
AI compatibility passport
This is the user-facing hook: answer the question first, then show the table.
Your recommended path
Next action cards
What users need to understand
Runtime decision matrix
Help visitors pick a path without reading dozens of forum comments.
WebLLM
Browser AI experiments, tiny demos, local-first curiosity.
Transformers.js
Embeddings, classification, small browser ML helpers, route memory later.
LM Studio
Friendly desktop local chat, stronger models, visual controls.
Ollama / llama.cpp
Developer-friendly local models, CLI workflows, app integration.
Model tier table
Starter compatibility tiers. Use the public benchmark table below for verified llama.cpp benchmark rows.
| Tier | Typical size | Browser path | Desktop path | RAM | VRAM | Best for | Last checked | Confidence | Source |
|---|
Local AI benchmark table
Local AI speed depends heavily on the model, quantization, runtime, backend, driver, OS, cooling, runtime version, and memory bandwidth. This benchmark table collects public reference results so you can compare rough hardware tiers before choosing a local AI setup.
| Hardware | Runtime / Backend | Model / Quant | PP t/s | TG t/s | Test condition | Confidence / Checked | Source |
|---|---|---|---|---|---|---|---|
| Apple M2 Apple laptop baseline | llama.cpp Metal | Llama 2 7B Q4_0 | 179.57 | 21.91 | pp512 / tg128 | verified 2026-06-11 | Source llama.cpp Apple Silicon M-series benchmark discussion #4167 |
| Apple M2 Pro Apple laptop midrange | llama.cpp Metal | Llama 2 7B Q4_0 | 294.24 | 37.87 | pp512 / tg128 | verified 2026-06-11 | Source llama.cpp Apple Silicon M-series benchmark discussion #4167 |
| Apple M2 Max Apple desktop or high-end laptop | llama.cpp Metal | Llama 2 7B Q4_0 | 671.31 | 65.95 | pp512 / tg128 | verified 2026-06-11 | Source llama.cpp Apple Silicon M-series benchmark discussion #4167 |
| Apple M2 Ultra Apple desktop workstation | llama.cpp Metal | Llama 2 7B Q4_0 | 1238.48 | 94.27 | pp512 / tg128 | verified 2026-06-11 | Source llama.cpp Apple Silicon M-series benchmark discussion #4167 |
| NVIDIA GeForce RTX 3060 12GB Entry CUDA desktop | llama.cpp CUDA | Llama 2 7B Q4_0 | 2407.67 | 76.92 | pp512 / tg128, Flash Attention enabled | verified 2026-06-11 | Source llama.cpp CUDA scoreboard discussion #15013 |
Includes public local AI benchmark rows for Apple M2 local AI, RTX 4060 local AI, RTX 4090 local AI, Snapdragon X Elite local AI, WebGPU compatibility planning, llama.cpp benchmark comparison, and tokens per second reference checks.
Community benchmark template
Later, this becomes the data engine that makes the Atlas trustworthy.
Related route
Turn the visit into a Bluesky journey.