← Back to the site

Client-side · OOD vs Adversarial detector

Viyog,
telling anomalies apart.

Safety-critical systems need to respond differently to two kinds of anomaly: an out-of-distribution input calls for abstention, an adversarial one demands rejection. Viyogtells them apart by reading the roughness of a model's first convolutional layer — one forward pass, no gradients, no training. Everything below runs in this tab; nothing is uploaded anywhere. See the full detector leaderboard for the complete 20-architecture comparison this demo doesn't reproduce.

Input

Upload or pick a sample

Drop an image, or click to choose one

Curated samples

Stage 1 — ID vs non-ID detector(s)

Result

Verdict

Nothing scored yet.

Idle.

So far

Loading aggregate stats…

How it works

One forward pass, two stages.

Stage 1 asks whether the input looks like the model's training distribution at all, using one or more cheap, gradient-free detectors computed straight from the classifier's own logits (MSP, MaxLogit, Energy, Entropy, GEN, KL-Matching). Pick which ones count above — a sample is flagged only once at least one selected detector crosses its calibrated threshold.

Stage 2 only runs on flagged samples. Viyog hooks the first convolutional layer and measures the roughness — total variation — of the quietest 10% of channels on in-distribution data (the "dormant band"). Gradient-based adversarial attacks inject high-frequency residue into exactly those otherwise-silent channels; natural OOD inputs don't. Higher roughness reads as adversarial, lower as out-of-distribution.

This demo ships a simplified Stage 1 (six logit-based detectors) — the thesis additionally evaluates Mahalanobis-distance, KNN, and ViM detectors across 20 architectures; see the full leaderboard linked above for that complete comparison.