Nanbeige, in the browser
Experiment report · 7 September 2026 · Apple M4 development machine
Custom WGSL is the selected browser engine. It produced about 23 tokens/second on the measured M4 Chrome configuration, while preserving trained two-pass execution. The selected compact artifact is 2,345,711,616 bytes (2.346 GB), calibrated on RunPod. This is the strongest size/quality tradeoff found in this experiment, not a universal best implementation.
Status: Chrome compact-artifact acceptance passed the recorded validation suite in fast and portable FP32 profiles. Safari 26.5 full-model generation and benchmarks are measured on the Q8-embedding alternate. iPhone 15 Pro Max, Pixel 10a and iPad M2 remain untested.
Deployment: Verified on the public Hugging Face Static Space in Chrome 152 / M4: model loading, generation, Stop, identical 128-token rerun after Stop, and JSON export all passed. First public-origin load took 58.4 seconds.
What must stay correct
The original model and technical report describe a Looped Transformer with 3B non-embedding parameters. Its 22 physical layers execute twice using shared weights, with 44 separate logical KV-cache slots. The final RMSNorm applies at both loop boundaries.
The implementation preserves 48 query heads, eight KV heads, head dimension 128, RoPE theta 70,000,000, and untied embedding/output weights. Query width is 6144 despite hidden width 3072; treating it as an ordinary square projection breaks the model. The llama.cpp implementation and MLX port informed this review; neither supplied browser performance measurements for this experiment.
Changing depth was tested, not assumed beneficial. On the small eight-text diagnostic corpus, perplexity was 73.88 at two loops, versus 73,270 at one and 46,168 at three. These results strongly reject a depth slider for this checkpoint; they are not a general claim about recurrent models. The browser supports trained depth two only. Depth results
Engine choice and measured speed
The release comparison used M4 Chrome 152, two prompts of 45/48 tokens, 64 generated tokens, and three repeats per prompt after a separately recorded warmup. Table entries are medians across six runs. TTFT excludes model loading; timings measure worker wall time, not isolated GPU kernels.
| Configuration | Decode tokens/s | TTFT | Weight download |
|---|---|---|---|
| ONNX Runtime Web, release | 18.34 | 417.0 ms | About 3.08 GB |
| Custom WGSL, compact, fast | 22.83 | 565.1 ms | 2.346 GB |
| Custom WGSL, compact, portable FP32 | 22.90 | 564.5 ms | 2.346 GB |
| Safari 26.5, custom WGSL, calibrated Q8 embedding | 22.51 | 813 ms | 2.601 GB |
The compact fast configuration decoded about 24% faster than ONNX, but ONNX prefilling was faster. Portable FP32 was effectively tied with fast WGSL on this machine. These are complete runtime/artifact configurations, not an isolated engine comparison. Safari's six runs used full-size Gram8192 and subgroup-free FP16; they include its first run and are not a controlled Chrome/Safari ranking.
Michionlion's ONNX export was run with its required development ORT operators and GPU-resident caches. Transformers.js supplies tokenization/chat templates; its ONNX route shares ORT rather than constituting an independent compute backend. Native MLX/llama.cpp were source references. WASM-only inference was not implemented or benchmarked and is not a working fallback.
The WGSL engine keeps quantized weights on the GPU, uses tiled prefill and matrix-vector decode, fuses SwiGLU, caches bindings and reduces argmax on-device. Further optimization had mixed results: a small-batch GEMV prefill path worsened median TTFT to 3.41 seconds; a 32-token tile reached 965 ms versus 558 ms baseline. R4 decode was effectively tied; fused norm R4/R8 did not win. The baseline remained selected. Raw events · Grouped timings
Weight quality and compact selection
All candidates use the same 4,080 scored WikiText-2 test tokens at sequence length 256. Calibration uses WikiText train activations from both loop passes. The selected method uses damped within-group covariance, scale search and integer coordinate descent; it is not full GPTQ or AWQ.
| Candidate | GB | Screening perplexity | Source top-1, six prompts |
|---|---|---|---|
| Source weights evaluated in FP32 | — | 41.774 | Reference |
| Initial Q4 group32, Q8 embedding | 2.601 | 48.828 | 6/6 |
| Full-model unweighted MSE search | 2.601 | 51.788 | 4/6 |
| Q4 group64, Q8 embedding | 2.471 | 52.360 | 5/6 |
| Gram2048, Q8 embedding | 2.601 | 44.827 | 5/6 |
| Gram8192, Q8 embedding | 2.601 | 45.237 | 6/6 |
| Gram8192, Q4 embedding with MSE | 2.346 | 44.859 | 6/6 |
The compact choice saves 255,197,184 bytes (9.81%) against the Q8-embedding version. Its mean prompt KL versus source is 0.07180; its screening perplexity remains about 7.4% above source. Gram2048 narrowly won perplexity but lost one next-token match and had higher prompt KL. Embedding-only MSE helped even though full-model MSE hurt.
Repeated candidate selection on this tiny screening set makes it a development set, not an independent final quality evaluation. Six prompts do not establish general instruction, code, multilingual or long-context quality. Quality summary · Selection notes
Selected compact artifact: Q4 embedding, Gram8192 + MSE. Full-size alternate: Q8 embedding, Gram8192.
Correctness evidence and practical limits
The final compact artifact was checked on six prompts in fast and portable FP32 profiles, plus a split-prefill check for each. All 14 runs matched their quantized reference's top-1 and all 104 compared continuation tokens. Across 166,144 logits per case, maximum error was 0.008517 for fast FP16 KV and 0.00010443 for portable FP32; cosine exceeded 0.99999995. These are observed bounds on this suite, not universal error guarantees.
Chunk 64 versus chunk one agreed on top-1 and eight continuation tokens in both profiles. Maximum logit differences were 0.004543 and 0.00007296. Earlier baseline tests also covered subgroup-free FP16. Artifact names and exact input IDs determine reference matching; missing references remain unavailable. Full parity results
A separate comparison against the source model distinguishes quantization quality from implementation correctness: compact fast matched source top-1 on 6/6 prompts, with mean KL 0.07190; ONNX matched 4/6, with mean KL 0.08289. This tiny first-token comparison is not a broad quality benchmark. Source comparisons
At context 512/1024/2048, FP16 KV allocates 88/176/352 MiB; FP32 doubles that. These are calculated allocations, not measured process peaks. Model weights, staging and browser overhead come in addition. Loading consumes one shard at a time, verifies hashes and uses best-effort persistent caching. The compact artifact has 31 weight bundles, largest 80,523,264 bytes; Q8 has 35, largest 127,598,592 bytes. Counts exclude the manifest. Download size does not equal live memory. Smaller context cannot remove the multi-GB weight floor.
One first-load observation of the compact artifact at new URLs took 63.34 seconds. Reloading the same cached artifact after a profile change took 3.30 seconds; the recorded cached ONNX load took 4.83 seconds. These are separate load observations, not generation timings, a cache-eviction experiment, or portable network-speed guarantees. Load metadata
On Safari, both portable and FP16 small-kernel checks passed after fixing adapter reuse in the test harness. The earlier synthetic failure was not evidence that Safari lacks FP16. Safari 26's WebGPU release and Chrome's Android support establish platform capability, not success on these phones. Start physical-device testing at 512 context, probe actual features, and use the direct Space URL; iframe storage and cache eviction differ. WebKit storage policy
Reproduce and continue
Conversion/calibration tools are in source/experiments/quantization; run them on RunPod, never on the Mac. Validation tools and the summarizer reproduce application checks and result tables without loading weights.
RunPod used one RTX A6000 48 GB at $0.53/hour for about 3,774 seconds: estimated compute $0.5556 against the $20 budget. Pod deletion succeeded and a subsequent GET returned HTTP 404; no persistent volumes were provisioned. Container-storage charges may be additional, so this is not a finalized invoice. Budget and deletion evidence
Next: test physical phones/tablets, thermal drift, cached reloads and a fresh quality holdout. Broader throughput claims require longer prompts and matching artifacts across engines.