Tiny, calibrated text classifiers that run in the browser. One decision, about 20 MB, offline, no per-call cost, and each model knows when it is unsure.
Runs entirely in your browser. The first run downloads the selected model from Hugging Face; your text never leaves this page.
Below the calibrated threshold a decision should be escalated to a larger model. Models are trained on synthetic data; expect lower accuracy on real text.
task spec (YAML) → collect inputs → teacher labels → train → calibrate → evaluate → export → browser
A small MiniLM sentence encoder is fine-tuned with a classification head, exported to ONNX q8 and run with transformers.js. Every base model that fits the download budget is trained and the best fit is kept. Each prediction is a typed decision with a calibrated confidence and a threshold chosen for a target precision. Exports are checked for parity: the browser must match Python on the test set.
| Task | Base | F1 | MB | WebGPU parity |
|---|---|---|---|---|
| Comment moderation v4 | MiniLM-L6 | 0.934 | 24.32 | 100% |
| Prompt injection v3 | MiniLM-L3 | 0.942 | 18.58 | 100% |
| Sentiment v2 | MiniLM-L6 | 0.880 | 24.32 | 100% |
| Support triage v2 | MiniLM-L6 | 0.939 | 24.32 | 100% |
q8 test F1 on small synthetic holdouts (50–257 cases); real-world accuracy is unmeasured. Chromium WASM and WebGPU predictions agree with Python ONNX on 99.6–100% of test cases. Training report
| Base | Download | F1 | Handled | p95 |
|---|---|---|---|---|
| MiniLM-L3 | 18.6 MB | 0.943 | 88% | 6.8 ms |
| MiniLM-L6 (kept) | 23.7 MB | 0.948 | 96% | 12.0 ms |
“Handled” is the share of test cases answered without escalation; p95 is browser latency.
uv run nodd run examples/comment_moderation.yaml
import { nodd } from "@nodd/browser";
const m = await nodd.load("/models/comment_moderation/v4");
const d = await m.decide("Buy cheap followers at …");
// { label: "spam", confidence: 0.99, … }