Quantized SmolLM2-135M-Instruct (q4) running offline on your local graphics processor. Zero server latency, zero cloud costs.
๐
Google DeepMind & ONNX Edge
Browser Model Engine
SmolLM2-135M-Instruct (Quantized to 4-bit)
Model Architecture Specs4-Bit ONNX (q4)
SmolLM2-135M is an ultra-lightweight language model developed by Hugging Face. Designed for zero-latency mobile execution and lightweight text formatting.
This laboratory demo downloads and executes a neural network **directly inside your web browser**. It utilizes your device's physical graphics card via the new **WebGPU API** for hardware-accelerated token generation.
Model Size~80 MB (cached locally)
Execution Provider~180MB RAM required
Host Compute Costโน0.00 (Zero Server Usage)
๐ How this operates under the hood
01. WARMUP
The page spawns a background **Web Worker** thread to prevent UI freezing, keeping the chat interface responsive at 60 FPS.
02. DOWNLOAD
Loads quantized model parameters from the **Hugging Face Hub** (download stream cached in browser storage for instant future loads).
03. COMPILATION
Compiles optimized GPU shaders directly on your graphics processor using the browser's native **WebGPU Execution Provider**.
04. INFERENCE
Token generation runs 100% on your device. Zero characters typed are sent to external servers, guaranteeing **100% privacy**.
๐ฅ๏ธ YOUR HARDWARE PROFILE DIAGNOSTIC
WebGPU:Checking...
Device RAM:Checking...
CPU Cores:Checking...
โ๏ธ LOCAL MODEL CACHE & LIFECYCLE MANAGER
Manage model weights stored locally in your browser's Cache Storage. Delete individual models to reclaim disk space, or clear all cached weights.
Initializing cache scanner...
Total AI Cache: Calculating...
โ ๏ธ **Note on Download**: The initial startup requires a one-time ~80MB download. Subsequent loads are instant as the model files are cached directly in your browser's Cache Storage.
Requires modern Chrome, Edge, or enabled Safari WebGPU flag.
Downloading Model Weights...
Connecting...
0%0.00 MB / 80.00 MB
Files Loading Status
โก Real-Time Telemetry HUDLIVE
Backend ProviderChecking...
JS Heap RAM:Calculating...
VRAM Allocation:Estimated ~180MB
Speed--- t/s
Latency (TTFT)--- ms
SmolLM2-135M-Instruct (q4)
โ๏ธ Generation Params
Temperature0.4
Max New Tokens200
โก Advanced Hyperparametersโผ
Top-P (Nucleus)0.90
Top-K40
Repetition Penalty1.10
Customize how the local model behaves. Changes are saved to your browser and persist across sessions.
โ System prompt saved to localStorage
๐ Offline RAG Context
0 docs
Drop local .txt, .md, .json, or .pdfdocuments. Client-side BM25 retrieves relevant sections for local context. Zero server upload.
Model loaded successfully. The local neural network is initialized on your device and ready to accept prompts.
Ready~0 / 200 tokens
๐พ Browser Model Storage & Cache
View local ONNX model weights stored in browser CacheStorage. Pre-cache for 100% offline use or purge to free disk space.
Scanning browser CacheStorage...
Total Stored: 0 MB
โ ๏ธ OOM Risk Safeguard
High-Memory Model Warning
Your device reports limited RAM or lacks WebGPU hardware acceleration. Loading this model (~1GB+) risks triggering a browser tab crash (Out Of Memory).
๐ ๏ธ Register Custom Browser JS Tool
Define a custom JavaScript function that FunctionGemma-270M can trigger locally inside your browser.
VirtualSach.in CLI Shell v2.4.0 [Type help for available commands]
sachin@virtualsach.in:~$
Virtual Sachin
ONLINE ยท Edge AI
SPEAKING
Live Mic
VIRTUAL SACHIN:
I'm the digital twin of Sachin. Ask me about my 18-year cloud infrastructure operations, open-source AI architecture patterns, or production war stories.
๐ Real-Time Escalate
Send a real-time message directly to Sachin's Telegram supergroup. He replies from Telegram, and it streams right here.