Tutorial

6 min read

How we get from a microphone to a rendered frame in under one frame, and the three places latency hides.

Ana Lindqvist, Audio Engineer

Man in black crew neck t-shirt sitting on chair in front of computer

Live voice is unforgiving. If the picture moves a beat after the word, viewers notice before they can say why. Our target is simple: the frame you see belongs to the sound you just heard.

Capture without buffering

We read the microphone through an AudioWorklet in 128-sample blocks, which at 48 kHz is 2.7 ms. Anything bigger trades smoothness for lag.

Analysis on the audio thread

A 1024-point FFT is reduced to four bands and a transient flag right where the samples arrive. Only five numbers cross to the render thread, so nothing waits on a copy.

Envelopes, not raw values

Raw band energy flickers. Attack and release envelopes keep motion fast on the way up and calm on the way down, which reads as musical rather than nervous.

Rendering on the next frame

The shader reads the latest values at the start of every frame. At 60 Hz the whole path stays under 16 ms, and at 144 Hz under 7.

Create a free website with Framer, the website builder loved by startups, designers and agencies.