← Back to the site

Real evaluation data · no live model

Catching the bit flip
no one else notices.

A single bit, flipped once inside a running ResNet-18, almost never crashes anything — the network just quietly returns a wrong answer. DrDNA profiles what "normal" activations look like at 19 points inside the network, then scores every inference against that baseline. Below are the real scores from 3,596 actual runs — drag the threshold and watch it decide, live.

Interactive

Set the detection threshold.

Each mark is one real inference — muted for the 3,000 clean runs, accent for the 596 runs with a bit flipped mid-network. Drag the threshold; every count below updates from the real 3,596 scores.

499.0

Faults caught

—

—

False alarms

—

—

Faults missed

—

—

    Interactive

    What each layer's "normal" looks like.

    The detector's baseline: a real profiled activation histogram at each of the 19 detection sites, from a 32-neuron cohort sampled over 1,500 clean inference passes. Pick a layer.

    Stem

    Stage 1

    Stage 2

    Stage 3

    Stage 4

    Head

    —

    How this works

    How the data was produced.

    Like the Feature-Space Geometry demo, this page doesn't run a model in your browser — DrDNA's fault injection (a custom single-bit-flip injector built on PyTorchFI) needs a real PyTorch runtime and a trained checkpoint neither of which fit in a client-side page. Instead, this replays the repo's own real experiment output: testwithoutSDC.csv (3,000 clean forward passes) and testwitSDCoutput.csv (596 passes with one bit flipped at a random detection site), alongside tau2.pkl, the real per-layer activation baseline computed once during profiling.

    Each score is Tau1 + Tau2 + Tau3 — three abnormality signals (per-neuron activation-frequency deviation, per-layer Earth Mover's Distance from the baseline histogram, and a shift in which neurons are most/least active) summed across all 19 layers with equal weight. Nothing on this page is simulated or interpolated: every mark in the strip above is one of the real 3,596 scores, and every bar in the layer histogram is a real bin count from the profiling run.

    Source: amanyagami/Detecting-Silent-Data-Corruptions-in-Deep-Neural-Networks.

    The two distributions overlap slightly — a few clean runs score above 499, a few faulty ones score below it — so no single threshold gets both counts to zero at once. That is the precision/recall tradeoff: drag the slider low and you'll catch every fault at the cost of more false alarms; drag it high and false alarms disappear but a few faults slip through.

    At the repo's own reported operating point (threshold 499), the real numbers are 99.5% of faulty runs caught against 5.6% of clean runs flagged — reproduced above from the same data.