Skip to content

Source Separation ​

Source separation produces several audio signals from one recording. Choose HPSS when you want sustained and transient structures, NMF when you want learned spectral components, and linked NMF when several channels must share the same masks. These algorithms do not assign instrument names or guarantee isolated vocals.

Choose a method ​

There are three routes, and they answer different questions.

A/B PROCESS · HPSSIDLE
Stem decomposition — pulling the percussive part out of a mix

A sustained pad chord bed with sharp broadband hits on every beat (Full mix). The stage applied here is HPSS decomposition, and the B side is its percussive component alone: the hits survive as short vertical events while the pad's steady spectral lines are pushed out. That is one stem of a two-way split, not a full multi-stem separation (Percussive). Both averaged spectra are drawn together so you can see what the split kept. Flip Compare to audition the mix against the stem it was decomposed into.

Compare

The demo above is the harmonic/percussive split: a pad bed with broadband hits on every beat, and the B side is the percussive component alone. Sustained spectral lines are pushed out while the hits survive as short vertical events. That is one half of a two-way split, not a multi-instrument separation.

typescript
import { hpss, hpssWithResidual, decomposeStems } from '@libraz/libsonare';

// Two-way split on a fixed axis: sustained vs transient.
const { harmonic, percussive } = hpss({ samples, sampleRate });

// The same split with the leftovers exposed.
const hard = hpssWithResidual({ samples, sampleRate, hardMask: true });

// Unsupervised components, each one listenable.
const { components, w, h } = decomposeStems({
  samples,
  sampleRate,
  nComponents: 4,
  maskPower: 2, // Wiener-style; separates harder, more artefacts on shared partials
});
HPSSdecomposeStems(...)
What decides the splitA fixed axis: median-filtering along time keeps sustained content, along frequency keeps transientsNon-negative factorisation learns nComponents recurring spectral shapes from the material itself
What comes backhpss returns harmonic and percussive signals; hpssWithResidual also returns a residualOne signal per component, plus the w component and h activation matrices
LabelsKnown in advanceNone — you inspect w and h to work out which component is which
ReconstructionDefault soft-mask harmonic and percussive outputs sum back to the input; hard masks may leave a residualThe masks sum to one where the model has energy, so the components sum back to the input
PhaseOriginalOriginal — the mask is applied to the complex spectrogram, which is what makes each component listenable

The NMF demo below lets you audition each of four components. It uses 30 iterations for interactive processing; components have no instrument labels.

A/B PROCESS · NMFIDLE
NMF decomposition — four learned components

NMF learns recurring spectral shapes from this mix and returns four unnamed components. Compare each component with the full mix. A component can contain several instruments; its number does not identify an instrument. The demo uses 30 iterations and retains the original phase and output levels without separate loudness matching.

Compare
Component

Defaults worth knowing: both median kernels are 31, and under the default soft mask hpssWithResidual(...) returns a silent residual because the two masks already sum to one — pass hardMask: true for a residual that actually carries the band neither component claimed. decomposeStems(...) and decomposeStemsLinked(...) use the same NMF defaults: 4 components, nFft: 2048, hopLength: 512, 100 iterations, beta: 2 (Frobenius; pass 1 for Kullback-Leibler), init: 'random', and maskPower: 1 keeping the magnitude ratio. decomposeStemsLinked(...) defaults an omitted sampleRate to 22050.

For a multichannel recording, use decomposeStemsLinked(...). It averages the channels' magnitude spectrograms to fit one NMF model and one set of component masks. It then applies each mask unchanged to every channel's original complex spectrogram. This preserves interchannel level and phase relationships within each time-frequency bin and avoids independently chosen masks for the channels. A separated component can still have a different stereo image from the full mix because it retains different content.

Pass at least one Float32Array channel, make every channel the same length, and keep the channel count at 64 or below. The result uses components[k][c] for component k on channel c; w, h, and sampleRate have the same meaning as in decomposeStems(...). A one-channel call is bit-identical to decomposeStems(...) with the same options.

typescript
import { init, decomposeStemsLinked } from '@libraz/libsonare';

await init();

const linked = decomposeStemsLinked({
  channels: [leftChannel, rightChannel], // equal-length Float32Array planes
  sampleRate,
  nComponents: 4,
});
const firstLeft = linked.components[0][0];
const firstRight = linked.components[0][1];
console.log(linked.w.length, linked.h.length);
typescript
import { decomposeStemsLinked } from '@libraz/libsonare-native';

const linked = decomposeStemsLinked({
  channels: [leftChannel, rightChannel], // equal-length Float32Array planes
  sampleRate,
});
const firstLeft = linked.components[0][0];
const firstRight = linked.components[0][1];
python
import libsonare as sonare

linked = sonare.decompose_stems_linked(
    [left_channel, right_channel], sample_rate=sample_rate, n_components=4
)
first_left = linked["components"][0][0]
first_right = linked["components"][0][1]
print(linked["w"].shape, linked["h"].shape)
bash
# The CLI exposes mono `decompose-stems` only.
# Use a library binding when one NMF model must serve several channels.

None of these routes is a trained instrument separator, and none will hand you a clean isolated vocal from a dense mix. They are useful for making a downstream estimator's job easier: beat tracking on the percussive part, chroma and key on the harmonic part, pitch tracking on a component that isolated the lead, or multichannel processing that must keep the stereo image.

Music Analysis, Audio to Notes, Melody and Pitch