Use when Browser-native AI mastery. Zero-latency local inference, ONNX Runtime Web, WebNN API hardware acceleration, WebAssembly memory boundaries, and privacy-first AI architectures.
Installs into .claude/skills of the current project.
Are you the author of Browser Native Ai?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/harmitx7-browser-native-ai-tribunal-kit)
---
name: browser-native-ai
description: "Use when Browser-native AI mastery. Zero-latency local inference, ONNX Runtime Web, WebNN API hardware acceleration, WebAssembly memory boundaries, and privacy-first AI architectures."
version: 5.0.0
last-updated: 2026-09-13
skills:
- generative-ui-expert
- webgpu-performance
- 60fps-animation
tools: Read, Grep, Glob, Bash, Edit, Write
scripts-binding:
- .agent/scripts/lint_runner.js
- .agent/scripts/verify_all.js
---
# Browser-Native AI (Local SLMs)
---
## 🛠️ Technical Architecture & Reference Recipes
---
## 1. Core Principles
- **Privacy by Default:** Data never leaves the browser. This is critical for HIPAA compliance, banking, and private notes apps.
- **Zero-Latency:** Because the model runs in memory, token generation and text embeddings happen instantly.
- **Hardware Acceleration First:** Always attempt to use WebGPU (`executionProviders: ['webgpu']`) or WebNN before falling back to WebAssembly (Wasm).
## 2. ONNX Runtime Web Integration
Use `@huggingface/transformers` (Transformers.js) or `onnxruntime-web` for execution.
```typescript
import { pipeline, env } from '@huggingface/transformers';
// Use WebGPU backend for acceleration
env.backends.onnx.wasm.numThreads = 1;
env.allowLocalModels = false;
// Instantiate an SLM or Embedding model
const extractor = await pipeline('feature-extraction', 'Xenova/all-MiniLM-L6-v2', {
device: 'webgpu', // Fallback to 'wasm' if needed
});
// Run inference entirely offline
const output = await extractor('Hello world', { pooling: 'mean', normalize: true });
console.log(output.data); // Float32Array embedding
```
## 3. Memory & Asset Management
- **Quantization:** Only load `q4` (4-bit quantized) models into the browser to prevent crashing mobile devices. A 7B parameter model is ~4GB quantized, which is too large. Target 0.5B to 1.5B parameter models (e.g., Llama-3.2-1B, Phi-3-mini).
- **Caching:** Cache model weights using the Origin Private File System (OPFS) or Cache API so the user only downloads the 500MB payload once.
- **Web Workers:** AI inference blocks the main thread in Wasm mode. **Always** run inference inside a Web Worker so the UI stays 60fps.
## 4. LLM Traps & Pre-Flight Checks
- **TRAP:** Running inference on the React main thread.
- **FIX:** Move pipeline instantiation and execution to a `worker.js` and communicate via `postMessage`.
- **TRAP:** Failing to handle model download progress.
- **FIX:** Pass a `progress_callback` to the pipeline to show a loading bar (e.g., "Downloading weights 45%").
- **TRAP:** Loading float16 or float32 models.
- **FIX:** Only request ONNX models that are specifically quantized (`_q4f16`) for web.
## Verification Protocol
Before submitting code, ensure:
1. `postMessage` architecture is used for non-blocking inference.
2. WebGPU is requested as the primary execution provider.
3. Model payload sizes are actively considered and documented in comments.