All work
03Experiment

Lucid

On-device computer vision for calmer feeds.

A Chrome extension that filters distracting and explicit content on Instagram and TikTok using a vision transformer running entirely in the browser — no cloud calls, no images leaving the machine.

Role
Engineering + ML
Year
2026
Status
In progress
LucidDiagram pending

Overview

Lucid is a Manifest V3 extension built as an end-to-end machine-learning systems case study: data collection methodology, a labelling toolkit, the fine-tuning notebook, the trained model, and a documented architecture for in-browser inference. Two signals feed one decision — a user-curated blocklist matched against the post caption, and a fine-tuned image classifier run locally through ONNX Runtime Web. The architecture generalises to any visual classification task that benefits from staying on-device.

Problem

Feed interfaces give you almost no control over what reaches your eye. Doing something about it means classifying content faster than you can scroll past it — and doing that in the cloud would mean uploading every thumbnail a person looks at, which trades one problem for a worse one.

Approach

Run the model locally and combine it with a signal that costs nothing. A blocklist match against the caption resolves in about a millisecond; the image classifier takes around 700 milliseconds. Stage one fires immediately, stage two refines the verdict afterwards.

Architecture

A per-tab content script picks a platform strategy, observes the DOM, and extracts caption and thumbnail. Both signals meet in a decision combiner. Inference runs in an offscreen document rather than the content script, so a 700 ms forward pass never blocks scrolling. The service worker routes messages; the model is fetched once from a public GitHub Release and cached via the Cache API, backed by an L1/L2 verdict cache.

Key interactions

  • Two-stage decision — blocklist ~1 ms, classifier ~700 ms
  • Sticky-blur invariant: stage two may upgrade a blur but never removes one
  • Inference in an offscreen document, off the scrolling path
  • 4-class ViT-Tiny fine-tune — val macro F1 0.83, gore F1 0.93
  • ~22 MB model lazy-loaded and cached, keeping the extension ~11 MB
  • PlatformStrategy interface — a new platform is one file

Technical decisions

The sticky-blur invariant

Once the blocklist fires, the classifier is allowed to relabel or strengthen the blur but never to remove it. Without that rule, a caption match would blur instantly and a disagreeing model would unblur 700 ms later — showing the user exactly the thing they asked not to see.

Ship the runtime, fetch the model

Bundling a 22 MB classifier into the extension package would make every install pay for it up front. Hosting it as a versioned GitHub Release asset and caching after first use keeps the download at ~11 MB and makes model upgrades independent of extension releases.

Offscreen document over a worker

MV3 service workers are terminated aggressively, which is fatal for a warm ~22 MB session. An offscreen document gives the model a stable home without ever touching the page's main thread.

What I took from it

  • The model was the easy part — the latency budget defined every other decision.
  • Giving the user the thresholds rather than picking them is what makes it a tool instead of a filter.