The crowd
One voice in, up to twenty voices out: your lead, fifteen harmonies, and four doubles. This page is how the crowd is built, and, just as importantly, how its members walk on stage without anyone hearing the door.
The harmony voices
Up to fifteen harmony voices render alongside the lead: sixteen-note polyphony from one singer. Where their notes come from depends on the Source setting:
- Scale: the 3rd, 5th, and 7th toggles stack diatonic intervals above the lead's target, inside the chosen key and scale.
- MIDI: every note of the committed chord gets a voice. Play four notes, hear four voices; the chord-commit behaviour that makes this playable is on the deciding page.
- Chord: the split option. The lead keeps following the scale while the harmony takes the notes you hold. Correction stays free; the chord is worn by the crowd.
Each harmony voice is your voice, shifted to its note through the same engine as the lead. The bank is processed once, on the mono analysis sum, and then placed: each voice equal-power panned across the stereo field, scaled by Spread, with the bank's overall gain on Harm Level.
Harmony targets get their own small mercy: a ~6 ms smoother on their shift ratios. Without it, a snapper settling or a chord change steps a voice by a semitone mid-note, a stair-step you can hear. With it, the step becomes a short slew.
The doubles
Four fixed voices in the ADT tradition: slight detunes, timing lags of 11–31 ms, small formant offsets, panned wide, their width riding the same Spread control as the harmonies. Two or four of them, on the Double switch. They don't take targets of their own; they ride the lead's pitch, offset by their fixed detunes, which is why they read as thickness rather than as extra parts. Their level is their own too: Harm Level scales the harmonies only, while the doubles sit at fixed gains that rebalance with the count.
Ensemble: the scatter control
A clone stack sounds like a machine because it is one: every voice landing identically. Ensemble is one 0-to-1 control over how much the crowd's members differ: per-voice formant offsets, small pitch scatter, staggered entry lags, doubler detune. At 0 the stack is clinical; at 1 the scatter is at maximum.
Every one of those differences is a fixed constant: there is no random number generator anywhere in the plugin. The same performance at the same settings renders the same crowd, to the bit, every time.
Walking on stage
The hardest problem in this section is not the voices; it's their entrances. Start a voice the instant its note arrives and you splice stale audio against fresh: a small "tut" at every harmony entrance. The fix has three layers:
- Every voice carries a short gate envelope (~2 ms in, ~4 ms out), so no voice ever starts or stops as a hard edge.
- A newly activated voice first flushes itself and holds its gate closed until genuinely fresh audio has traversed its entire pipeline. The audible consequence: cold harmony entrances land roughly 25–60 ms behind the lead. This is deliberate, and it is the honest trade: the lag is the signature of an entrance that cannot click.
- The whole harmony/doubler bus rides a slower envelope (~15 ms in, ~50 ms out) that sustains through consonants and breaths rather than gating on them, so the crowd doesn't flicker every time you close a syllable.
What leaves this page
A stereo bed: the harmonies at the Harm Level setting, the doubles at their own fixed gains, joining the lead and the dry path for alignment and the mix.
last updated · Charlie