Chat with a large language model. No server, no account.
Model weights stream to your GPU through WebGPU. Nothing you type leaves this machine.
This browser can't run WebGPU yet
To chat here you need a browser with WebGPU: Chrome or Edge 113+ on a computer with a dedicated or recent integrated GPU. Firefox and Safari are getting there, but today Chrome works best.
Software renderer: your GPU is not reachable
WebGPU is falling back to a CPU renderer (SwiftShader or llvmpipe), so generation will be
many times slower and the larger models can fail while loading. On Linux this is usually the
browser GPU sandbox denying dlopen on the Mesa/Vulkan drivers. Verified fix,
Chrome stable (see "Running on Linux" in the README), relaunch the browser with:
google-chrome --enable-features=WebGPUService --ignore-gpu-blocklist --disable-gpu-sandbox
Loading a model here still works, it will just be slow.
Recent conversations
How it works
The weights stream from Hugging Face straight into a cache in your browser, so a finished download survives reloads. Press Load and they move on to your GPU through WebGPU; from there the whole conversation runs on this machine: prompt processing, token generation, sampling, everything. There is no server side and nothing you type leaves your computer.
Dense models keep every weight resident in GPU memory for the whole conversation. The two MoE cards work differently: only a slice of the experts is active for each token, so experts page in and out of the GPU as needed and the download can be much larger than the memory it asks for.
The VRAM cap is a soft guide, not a switch: cards above it stay selectable, and the real check happens at load, on the actual tensors of the model you picked. If your card has more memory than the default guess, raise the cap under Advanced.
Conversations and model files stay in this browser’s storage. Clearing site data for this page removes both.