Lucid
On-device computer vision for calmer feeds.
A Chrome extension that filters distracting and explicit content on Instagram and TikTok using a vision transformer running entirely in the browser — no cloud calls, no images leaving the machine.
- Engineering + ML
- 2026
- In progress
Overview
Lucid is a Manifest V3 extension built as an end-to-end machine-learning systems case study: data collection methodology, a labelling toolkit, the fine-tuning notebook, the trained model, and a documented architecture for in-browser inference. Two signals feed one decision — a user-curated blocklist matched against the post caption, and a fine-tuned image classifier run locally through ONNX Runtime Web. The architecture generalises to any visual classification task that benefits from staying on-device.
Problem
Feed interfaces give you almost no control over what reaches your eye. Doing something about it means classifying content faster than you can scroll past it — and doing that in the cloud would mean uploading every thumbnail a person looks at, which trades one problem for a worse one.
Approach
Run the model locally and combine it with a signal that costs nothing. A blocklist match against the caption resolves in about a millisecond; the image classifier takes around 700 milliseconds. Stage one fires immediately, stage two refines the verdict afterwards.
Architecture
A per-tab content script picks a platform strategy, observes the DOM, and extracts caption and thumbnail. Both signals meet in a decision combiner. Inference runs in an offscreen document rather than the content script, so a 700 ms forward pass never blocks scrolling. The service worker routes messages; the model is fetched once from a public GitHub Release and cached via the Cache API, backed by an L1/L2 verdict cache.
Key interactions
- Two-stage decision — blocklist ~1 ms, classifier ~700 ms
- Sticky-blur invariant: stage two may upgrade a blur but never removes one
- Inference in an offscreen document, off the scrolling path
- 4-class ViT-Tiny fine-tune — val macro F1 0.83, gore F1 0.93
- ~22 MB model lazy-loaded and cached, keeping the extension ~11 MB
- PlatformStrategy interface — a new platform is one file
Technical decisions
The sticky-blur invariant
Once the blocklist fires, the classifier is allowed to relabel or strengthen the blur but never to remove it. Without that rule, a caption match would blur instantly and a disagreeing model would unblur 700 ms later — showing the user exactly the thing they asked not to see.
Ship the runtime, fetch the model
Bundling a 22 MB classifier into the extension package would make every install pay for it up front. Hosting it as a versioned GitHub Release asset and caching after first use keeps the download at ~11 MB and makes model upgrades independent of extension releases.
Offscreen document over a worker
MV3 service workers are terminated aggressively, which is fatal for a warm ~22 MB session. An offscreen document gives the model a stable home without ever touching the page's main thread.
What I took from it
- The model was the easy part — the latency budget defined every other decision.
- Giving the user the thresholds rather than picking them is what makes it a tool instead of a filter.