Tutorial
6 min read
How we get from a microphone to a rendered frame in under one frame, and the three places latency hides.
Ana Lindqvist, Audio Engineer

Live voice is unforgiving. If the picture moves a beat after the word, viewers notice before they can say why. Our target is simple: the frame you see belongs to the sound you just heard.
Capture without buffering
We read the microphone through an AudioWorklet in 128-sample blocks, which at 48 kHz is 2.7 ms. Anything bigger trades smoothness for lag.
Analysis on the audio thread
A 1024-point FFT is reduced to four bands and a transient flag right where the samples arrive. Only five numbers cross to the render thread, so nothing waits on a copy.
Envelopes, not raw values
Raw band energy flickers. Attack and release envelopes keep motion fast on the way up and calm on the way down, which reads as musical rather than nervous.
Rendering on the next frame
The shader reads the latest values at the start of every frame. At 60 Hz the whole path stays under 16 ms, and at 144 Hz under 7.
