← Back to AI Product LabSOTA EDGE INFERENCE v1.0
โšก 100% Client-Side WebGPU Computing

COGNITIVE_EDGE: BROWSER_LLM

Quantized SmolLM2-135M-Instruct (q4) running offline on your local graphics processor. Zero server latency, zero cloud costs.

๐Ÿ’Ž
Google DeepMind & ONNX Edge

Browser Model Engine

SmolLM2-135M-Instruct (Quantized to 4-bit)

Model Architecture Specs4-Bit ONNX (q4)

SmolLM2-135M is an ultra-lightweight language model developed by Hugging Face. Designed for zero-latency mobile execution and lightweight text formatting.

Primary CapabilitiesFormatting, Classification, Zero-Latency
Context Window2,048 Tokens

This laboratory demo downloads and executes a neural network **directly inside your web browser**. It utilizes your device's physical graphics card via the new **WebGPU API** for hardware-accelerated token generation.

Model Size~80 MB (cached locally)
Execution Provider~180MB RAM required
Host Compute Costโ‚น0.00 (Zero Server Usage)

๐Ÿ” How this operates under the hood

01. WARMUP

The page spawns a background **Web Worker** thread to prevent UI freezing, keeping the chat interface responsive at 60 FPS.

02. DOWNLOAD

Loads quantized model parameters from the **Hugging Face Hub** (download stream cached in browser storage for instant future loads).

03. COMPILATION

Compiles optimized GPU shaders directly on your graphics processor using the browser's native **WebGPU Execution Provider**.

04. INFERENCE

Token generation runs 100% on your device. Zero characters typed are sent to external servers, guaranteeing **100% privacy**.

๐Ÿ–ฅ๏ธ YOUR HARDWARE PROFILE DIAGNOSTIC

WebGPU:Checking...
Device RAM:Checking...
CPU Cores:Checking...
โš ๏ธ **Note on Download**: The initial startup requires a one-time ~80MB download. Subsequent loads are instant as the model files are cached directly in your browser's Cache Storage.
Requires modern Chrome, Edge, or enabled Safari WebGPU flag.
โšก Theme Adaptive Shift
Switching layouts matching domain reading affinity...