webgguf LLM chat, entirely in your browser
Advanced
Engine

Reasoning effort applies to Qwen 3.8 and later, whose chat template defaults to xhigh. We ship medium: on a browser at a handful of tokens per second, xhigh is the difference between an answer and a page that looks stuck. The level is written into the system prompt, so it is visible in the exported prompt and two levels are not comparable runs. With Reasoning set to off there is nothing for it to apply to.

The expert-cache budget is derived from the VRAM cap (cap − allocated − reserve). Both the scheduling and the cache policy are ignored on dense models, which is seven of the nine cards: there is no expert cache to schedule.

Longer contexts cost VRAM for the KV cache and prefill time on the first turn. If the selected context does not fit the VRAM cap, the load stops and shows the count.

Sampling

temperature 0 = greedy decoding: reproducible, comparable across runs, and top-k, top-p and the seed do nothing until you raise it. Above 0 the sampling happens on the CPU from the same logits and the seed is recorded in the export, while the multi-step decode switches off on the cards that ship it, so answers also come out slower.

System prompt

Applied on the first turn of a context (the KV cache keeps it afterwards). "New chat" makes it effective again.

Export

The JSON carries load parameters, sampling with seed, the rendered prompt with token ids, and per-turn metrics. It is not a benchmark reference: no warm-up discard, no replicates, host state undeclared.

Chat with a large language model. No server, no account.

Model weights stream to your GPU through WebGPU. Nothing you type leaves this machine.

How it works

The weights stream from Hugging Face straight into a cache in your browser, so a finished download survives reloads. Press Load and they move on to your GPU through WebGPU; from there the whole conversation runs on this machine: prompt processing, token generation, sampling, everything. There is no server side and nothing you type leaves your computer.

Dense models keep every weight resident in GPU memory for the whole conversation. The two MoE cards work differently: only a slice of the experts is active for each token, so experts page in and out of the GPU as needed and the download can be much larger than the memory it asks for.

The VRAM cap is a soft guide, not a switch: cards above it stay selectable, and the real check happens at load, on the actual tensors of the model you picked. If your card has more memory than the default guess, raise the cap under Advanced.

Conversations and model files stay in this browser’s storage. Clearing site data for this page removes both.