A federated speech recognizer sends the server a model update and keeps the audio on the device. For models that read overlapping windows of frames, that update contains the speech. This page lets you hear it, run the attack yourself, and see where it stops.
Each clip below was attacked once: the client computed one gradient on the utterance, and the attacker saw only that gradient and the public model. The features come back exactly; a public vocoder turns them into sound. No transcript was used.
A small model of the same shape: 4 features per frame, windows of 5 frames, a 48-unit first layer. The page builds the first-layer gradient for a random utterance, takes its row space, and solves the linear system. Counting equations against unknowns gives a limit of (5 − 1) × 4 + 1 = 17 frames. Drag the length past it.
The first layer applies the same weights W to every window c of k frames. Whatever the rest of the network does, the chain rule makes its gradient a sum with one term per window:
So the rows of [∇W | ∇b] span the vectors [c, 1], and the rank counts the windows. Neighbouring windows share frames, which makes "every window lies in the span" a linear system in the features. It has at least as many equations as unknowns up to
With float32 gradients the last window of that budget is lost to rounding, so the measured limit on real models is (k − s) × F.
LibriSpeech test-clean, one float32 gradient per utterance. Blue: closed form on the first layer. Green: sequential decoding from deeper layers. Pink: recordings with digital silence, where the solution is not unique. Orange: the previous method, gradient matching with the transcript given. Hover a point.
ρ = spectral norm of the noise ÷ smallest signal singular value. The error follows ρ until the span is lost near ρ = 1.
Updates arrive as weight differences after local training. Median error of the closed form in each condition (features have a typical magnitude of 8, so anything under 0.01 is a full recovery).
Whether a layer gives away its inputs is decided by its shape. For one 5-second utterance, the rank of early-layer gradients (Whisper and wav2vec 2.0 with their released weights):