Track 3 — Listening
Every knob so far has moved on its own clock (sin(time*k)). Now the music takes the wheel. This track covers the audio registers, the smoothing idiom every serious preset uses, how to detect a beat without a beat detector, and how to measure whether any of it worked.
Lesson 1 · Raw audio is jittery
The simplest possible link: read a band, drive a knob.
per_frame_1=zoom = 1.01 + bass*0.15;
Watch it for a few seconds. It twitches — bass is close to instantaneous, so every sample-level flicker in the low end shows up as a visible jolt. This is the single most common beginner tell: a preset that shudders instead of breathing.
Swap in the attack-smoothed register:
per_frame_1=zoom = 1.01 + bass_att*0.15;
Same music, calmer motion. bass_att/mid_att/treb_att are envelope-followed versions of bass/mid/treb — they still rise fast on a hit, but they don't chatter. Default to the _att versions for anything continuous; reach for the raw bands only when you deliberately want the jitter (texture, noise).
Lesson 2 · Rolling your own smoothing
_att isn't the only tool. The RC low-pass filter — one line, tunable — is the idiom you'll see in nearly every hand-written preset that predates or ignores the _att registers:
per_frame_1=ra = 6/fps;
per_frame_2=bass_avg = bass_avg*(1-ra) + ra*bass;
per_frame_3=zoom = 1.01 + bass_avg*0.15;
bass_avg isn't a built-in — it's an ordinary variable. In NS-EEL, any name you assign persists from frame to frame automatically (there's no declaration step), so bass_avg*(1-ra) + ra*bass is a proper exponential moving average. At 60 fps, 6/fps is 0.1, so each frame bass_avg drifts 10% of the way from where it was toward the current bass. A per-frame ra of 0.05–0.15 (a numerator of 3–9 at 60 fps) is the useful range — smaller is smoother and laggier, larger snaps closer to raw. Dividing by fps keeps the cutoff the same regardless of frame rate: at 30 fps each frame takes a step twice as large. It's the same correction you saw as 75/fps in Track 1, written as a rate per second instead of against a 75fps reference — the more common form when a preset builds its own filter.
Why bother when _att already exists? Control. You choose the exact cutoff, and the same technique smooths anything — not just the three built-in bands, but vol, a derived signal, even another filter's output.
Lesson 3 · The volume-clock pattern
The highest-leverage idiom in the whole language: don't drive a knob from audio directly — drive an accumulating clock, and drive the knob from the clock. This is Pattern 1 from the coding guide:
per_frame_1=vol = (bass+mid+treb)*0.333;
per_frame_2=vol = vol*vol;
per_frame_3=mtime = mtime + vol*0.02*(75/fps);
per_frame_4=rot = sin(mtime)*0.03;
per_frame_5=zoom = 1 + 0.03*vol;
Two ideas stacked:
- Squaring compresses quiet and amplifies loud.
vol*volon a 0–1ish range pushes background hiss toward zero while barely touching a loud passage — the preset visibly rests between hits instead of humming continuously. mtimeaccumulates instead of following.rotdoesn't readvoldirectly — it readssin(mtime), andmtimeonly advances when there's volume to advance it. During a quiet bridge the clock nearly stops and the whole preset holds its pose; during a loud chorus it races. This is why energetic presets still feel composed rather than nervous: the music controls tempo, not just amplitude.
Lesson 4 · Detecting a beat without a beat detector
MilkDrop has no on_beat() callback. Beats are inferred from a threshold that adapts to the track, Pattern 6 in the coding guide:
per_frame_1=bass_thresh = above(bass_att,bass_thresh)*2 + (1-above(bass_att,bass_thresh))*((bass_thresh-1.3)*0.96+1.3);
per_frame_2=beat_fired = equal(bass_thresh,2);
per_frame_3=wave_r = 0.2 + beat_fired*0.7;
per_frame_4=wave_g = 0.85 - beat_fired*0.3;
Read the first line as a state machine with two states, chosen by above(bass_att, bass_thresh) (1 if a beat just cleared the bar, else 0):
- Beat fired (
aboveis 1): the whole right-hand expression collapses to1*2 + 0*(...)—bass_threshsnaps to2.0. The bar is now so high that next frame'sbass_attalmost certainly can't clear it. That is exactly the point: one hit can't retrigger itself for several frames. - No beat (
aboveis 0): it collapses to0*2 + 1*((bass_thresh-1.3)*0.96+1.3)—bass_threshdecays 4% of the way back toward a1.3floor every frame. Give it enough quiet frames and the bar is low again, ready for the next hit.
No branch, not even an if — above() returning exactly 1 or 0 is what lets both terms share one line while only one of them ever survives. This self-tuning bar is why the same threshold code works on a whisper-quiet ambient track and a wall-of-noise one: it's relative to recent loudness, not an absolute number picked for one song.
The second line reads the beat off the bar itself: equal(bass_thresh,2) is 1 only on the frame the bar snapped up. Testing above(bass_att,bass_thresh) again would compare against the new bar of 2.0, which almost nothing clears, so the flash would never fire.
beat_fired is reusable: anything that should happen once per beat can multiply by it. Track 7 dissects a Geiss preset that uses the same gate to fire a single-frame directional impulse instead of a color flash.
Lesson 5 · Measuring it
Everything above is a claim about what the preset does. Stims ships the same instrument its own agents use to check that claim — point it at a file and it reports, per variable, whether audio explains the motion:
bun run lab:reactivity -- --file docs/authoring/examples/35-reactive-vortex.milk
▶ Run the reactive vortex — RC-smoothed bass drives rotation speed, squared volume (the first half of the volume-clock pattern) drives zoom, and the adaptive threshold drives a color flash: everything in this track, combined.
The report gives each variable a verdict:
| Verdict | Meaning |
|---|---|
reactive |
measurably correlated with audio — the number you want |
autonomous |
moves, but on its own clock (sin(time)), not audio |
weak |
some correlation, but faint |
static |
never changes — not necessarily wrong (most variables in a preset should be) |
For zoom, rot, and wave_r in the vortex, you should see reactive. The waveform's own deviation (mainWave.deviation) reads reactive too — the tool is measuring the same thing your eyes were doing in Lesson 1, just with a number attached. This is the loop real authors use: write, measure, adjust the coupling strength until the verdict — not just the vibe — says reactive. It's also exactly what bun run lab:visual checks on the pixel side, and what the contributing guide asks for before a preset is submitted.
Lesson 6 · Signals only Stims has
Everything above is standard MilkDrop, portable to Winamp, projectM, and Butterchurn. Stims also feeds presets what the person watching is doing — where the pointer is, how hard it is being dragged, pinch and twist, and a set of key pulses — plus device motion. Every name is in the reference, completed and hover-documented in the editor:
| Family | Names | What it carries |
|---|---|---|
| Pointer | input_x input_y input_dx input_dy input_speed input_pressed input_count |
position (-1..1), per-frame movement, and whether anything is held |
| Hover | hover_active hover_x hover_y |
a mouse over the stage that is not pressing |
| Force | drag_intensity drag_angle wheel_delta wheel_accum |
how hard, which way, and the scroll wheel |
| Gesture | gesture_scale gesture_rotation gesture_translate_x gesture_translate_y |
pinch and twist — =/- and ,/. on a keyboard |
| Keys | action_remix action_accent action_mode_next action_mode_previous action_quick_look_1..3 |
R, Enter, X, Q/Z, 1/2/3 as pulses that decay over ~220ms |
The keys reach the preset only while the stage has focus — click the visuals first. Shift plus an arrow key steers the pointer without a mouse.
None of the 1,787 presets in the catalog read any of them. That cuts both ways: there is no prior art to copy, and a preset that answers the person watching is instantly unlike everything else in the catalog.
▶ Run the interactive drift — drag the stage to shove the field, twist to spin it, and press R to snap the color; with no hands on it, it is an ordinary slow drift.
These are not standard MilkDrop. A preset that reads them will still compile and run fine elsewhere (unknown identifiers evaluate to 0), but the interactivity silently disappears. Use them freely for Stims-native work; if portability matters, keep them behind a variable you can zero out, or skip them. The full contract, including the device-motion signals (motion_x/motion_y/motion_z), is in the signal contract.
What you can now build
Most presets take their personality from audio, and any preset that does is now legible: find the smoothing, find whether it's driving a knob directly or a clock, find the threshold logic if there's a beat effect, and measure instead of guessing.
Next: Track 4 — Warp fields, where knobs stop being one number for the whole screen and start varying per pixel.