Skip to content

Node.js Native Analysis and Effects ​

This page covers the analysis, effects, and feature-extraction functions of the @libraz/libsonare-native package. For usage, errors, the Audio class, cleanup, conversions, metering, mastering analysis, and streaming classes, see Node.js Native API; for the exported TypeScript types, see Node.js Native Types.

Analysis Functions ​

FunctionReturn TypeDescription
detectBpm(samples, sampleRate?)numberTempo in BPM
detectKey(samples, sampleRate?)KeyRoot, mode, confidence
detectBeats(samples, sampleRate?)Float32ArrayBeat timestamps
detectOnsets(samples, sampleRate?)Float32ArrayOnset timestamps
detectChords(request) / detectChords(samples, sampleRate?, options?)ChordAnalysisResultChord progression with timings. Frames below threshold are returned as explicit N.C. intervals; the options form enables HMM smoothing, key context, inversions, chroma method ('stft' default), and tuning (the legacy positional form is also supported)
detectDownbeats(samples, sampleRate?)Float32ArrayDownbeat (bar-start) timestamps
detectKeyCandidates(samples, sampleRate?, options?)KeyCandidate[]Ranked key candidates with correlation scores
analyze(samples, sampleRate?, options?)AnalysisResultAll-in-one analysis in one call: BPM and ranked BPM hypotheses, key, time signature and ranked time-signature candidates, beats, chords (including romanNumeral relative to the detected key), sections, timbre, dynamics, rhythm, melody, and form. The dedicated detect*/analyze* functions below remain available for targeted or parameterized analysis
analyzeWithProgress(request) / analyzeWithProgress(samples, sampleRate, onProgress, options?)AnalysisResultSame as analyze with a (progress, stage) callback; the request form accepts options and cancel, while the positional form takes analysis options fourth
estimateMeter(request)MeterEstimateScore a meter over a caller-supplied beat series, without audio and without re-running analysis. Request-only — takes EstimateMeterRequest
analyzeBpm(samples, sampleRate?, options?)BpmAnalysisResultTempo with confidence and alternate candidates. options: bpmMin, bpmMax, startBpm, nFft, hopLength, maxCandidates
analyzeRhythm(samples, sampleRate?, options?)RhythmResultTime signature, groove, syncopation. options: bpmMin, bpmMax, startBpm, nFft, hopLength
analyzeDynamics(samples, sampleRate?, options?)DynamicsResultDynamic range, loudness range, crest factor. options: windowSec, hopLength, compressionThreshold
analyzeTimbre(samples, sampleRate?, options?)TimbreResultBrightness, warmth, density, roughness, complexity, plus per-window timbreOverTime. options: nFft, hopLength, nMels, nMfcc, windowSec
analyzeSections(samples, sampleRate?, options?)Section[]Structural sections (intro/verse/chorus…) with timings. options: nFft, hopLength, minSectionSec. Long inputs may use a pooled boundary grid; use each section's start / end for placement
detectBoundaries(request)BoundaryResultStructural transitions plus the novelty curve they were picked from, and the grid both live on. request: samples, sampleRate, nFft, hopLength, kernelSize, threshold, absoluteThreshold, nMfcc, nChroma, peakDistance, useMfcc, useChroma. Reach for it when you want to apply your own threshold rather than take analyzeSections' labelled spans
analyzeMelody(samples, sampleRate?, options?)MelodyResultLead-melody contour (F0 per frame). options: fmin, fmax, frameLength, hopLength, threshold, usePyin, center
detectAcoustic(samples, sampleRate?, options?)AcousticResultRoom acoustics from a recording (RT60 — the time reverberation takes to decay 60 dB — and related measures). options: nOctaveBands, nThirdOctaveSubbands, minDecayDb, noiseFloorMarginDb
analyzeImpulseResponse(samples, sampleRate?, nOctaveBands?, minDecayDb?)AcousticResultRoom acoustics from a measured impulse response; minDecayDb controls the decay-fit threshold (default 30)
estimateRoom(samples, sampleRate?, options?)RoomEstimateResultEquivalent-room estimate with volume, dimensions, DRR (direct-to-reverberant ratio), absorption bands, RT60 bands, and confidence
synthesizeRir(options?)RirResultMono RIR (room impulse response) from shoebox geometry
roomMorph(samples, sampleRate, options?)RoomMorphResultOffline creative morph toward a target room; read morphed samples from .audio, with .sampleRate and .diagnostics alongside them
lufs(samples, sampleRate?)LufsResultIntegrated, final momentary/short-term windows, their EBU R128 maxima (Max-M / Max-S), and loudness range
lufsInterleaved(samples, channels, sampleRate?)LufsResultChannel-weighted multichannel loudness from interleaved samples
ebur128LoudnessRange(samples, sampleRate?)numberStandards-compliant EBU R128 loudness range (LRA) in LU
momentaryLufs(samples, sampleRate?)Float32ArrayMomentary loudness (400 ms) per step
shortTermLufs(samples, sampleRate?)Float32ArrayShort-term loudness (3 s) per step
version()stringLibrary version
voiceChangerAbiVersion()numberABI version of the realtime voice-changer POD config; separate from preset JSON schemaVersion
voiceCharacterPresetId(preset)VoicePresetId | nullCanonical voice-character preset ID; an unknown numeric ordinal returns null, while an unknown string ID throws
realtimeVoiceChangerPresetConfig(preset)RealtimeVoiceChangerConfigResolved flat POD config for a built-in voice preset, without JSON parsing. Throws on an unknown preset name or out-of-range ordinal
hasFfmpegSupport()booleanWhether the loaded native addon can decode via FFmpeg

Default sample rates differ by helper family:

Helper familyDefault sampleRate
Music analysis, effects, feature, and loudness helpers22050
analyzeImpulseResponse, detectAcoustic, estimateRoom, and synthesizeRir in the native addon48000

Common helpers are also available as Audio instance methods, as noted in the Audio section.

The tables below document the Node native API. The WASM package uses the same camelCase names, but functions with a required argument after sampleRate require that sampleRate position to be supplied. See JavaScript API for the browser signatures.

analyze() options ​

analyze(...) takes an options object as its third argument, or the same fields directly on the request object. It covers the whole pipeline in one place:

OptionDefaultDescription
nFft / hopLength2048 / 512STFT resolution shared across the pipeline
bpmMin / bpmMax / startBpm60 / 200 / 120Tempo search range and prior
useTriadsOnlytrueRestrict the chord search to triads
useHpsstrueHarmonic-only chroma for chord and key detection
chromaHighpassHz80Chroma high-pass cutoff in Hz (0 disables)
useBassWeightedtrueBass-weighted chroma combination
chromaHopMultiplier4Chroma hop multiplier; larger is faster
useChordHmm, useChordKeyContext, chordHmmBeamWidth, detectChordInversions—Chord post-processing, matching the trailing options on detectChords(...)
adaptiveTempofalseTrack a locally updated tempo prior through beat tracking
tempoUpdateIntervalBeats8Local tempo context length in beats; read only when adaptiveTempo is set
computeTempoCurvefalseDecode a per-beat local tempo curve into beatLocalBpm
meterCandidateNumerators[3, 4, 6]Meter numerators the estimator scores. At most 16 entries, each in [2, 32]; an empty list is rejected rather than restoring the default, and widening the set does not force a wider meter
meterDenominator4Beat unit reported for the detected meter, a power of two in [1, 32]. The estimator still reports 8 on its own when it resolves a compound meter
tuning0Recording tuning offset in fractions of a semitone, using the unit from estimateTuning(...); must be in [-0.5, 0.5). It shifts the chroma used for key, chords, and sections

useTriadsOnly points the other way here

In the unified analyze() path useTriadsOnly is true, while the standalone detectChords(...) defaults it to false. So analyze() searches triads alone — and reports no sevenths or extensions — until the caller passes useTriadsOnly: false.

computeTempoCurve is off because the curve is an extra output rather than a better analysis: nothing else in the result changes, so a caller who never reads it would pay a decode over the beat grid for nothing. The curve also describes the beat grid it was decoded from, and beat tracking holds a fixed tempo prior unless adaptiveTempo is set as well — measuring a tempo that actually moves needs both options.

analyzeWithProgress(...) accepts the same MusicAnalyzeOptions as analyze(...). In the request form, put them under options; cancel is a separate callback that returns true to request cancellation. The positional overload puts the options object after onProgress.

estimateMeter(...) ​

estimateMeter(...) scores a meter over a caller-supplied beat series. It reads only per-beat times and accent values, never audio, so an existing analysis — or an arbitrary span of one — can be re-scored without re-running the pipeline. For the underlying concept, see Meter and grouping.

typescript
const result = analyze(samples, sampleRate);

const meter = estimateMeter({
  beatTimes: result.beats.map((beat) => beat.time),
  beatStrengths: result.beatObservations.onsetStrength,
  candidateNumerators: [3, 4, 5, 6, 7],
});

console.log(meter.searched, meter.timeSignature.numerator, meter.grouping);
FieldDefaultDescription
beatTimes—Beat positions in seconds, non-decreasing
beatStrengths—Per-beat accent, the same length as beatTimes. The series is divided by its own maximum before scoring, so it needs no pre-scaling
candidateNumerators[3, 4, 6]Numerators to score; at most 16 entries, each in [2, 32]
denominator4Beat unit reported for the detected meter
downbeatWeight / measureWeight / subdivisionWeight1 / 0.5 / 0.15Scoring weights for the first beat of each measure, measure-to-measure accent agreement, and the subdivision pattern

Two sources are supported for beatStrengths: beatObservations.onsetStrength, the windowed value the library's own downbeat pass scores, and beats[].strength, a single unwindowed frame of the same envelope. Reading onsetEnvelope(...) at timeToFrames(...) for each beat is not a third: a hop counted in samples frames a different span of time at each rate, so one waveform at 32000, 44100 and 48000 Hz — beat times identical — produced winning numerators of 6, 3 and 4, and a window around each beat does not remove the dependence.

Two properties decide whether the answer means anything:

  • The default candidate set is {3, 4, 6}. An odd meter is reported only if its numerator was asked for, so a seven needs it listed explicitly.
  • searched is false when the beat series was under eight beats. Every other field then carries a fixed fallback rather than a result — including timeSignature.confidence, which reads 0, so an unchecked read degrades toward "no idea" rather than toward a middling detection. Read searched before treating a short span's answer as a detection.

grouping reports how the bar divides into accent groups, so a seven comes back as [3, 2, 2] rather than as a bare seven, and it always sums to the numerator. A single entry means no internal division was resolved. candidateScores is standardized and signed, listed in the order the request gave the numerators: zero is the level a numerator reaches on beats carrying no meter, so only the ordering and the gaps between entries carry meaning, and a score grows with the square root of how many beats were scored. candidates is ordered by descending support instead, so match the two on numerator rather than by index. confidence also changes meaning between its two homes: on timeSignature it is derived from the margin over the runner-up, on a candidates entry it is that candidate's share of the summed support (the entries sum to one), so the two must not share a threshold.

Asynchronous variants (Node only) ​

The Node addon also exposes Promise-returning variants. They run the DSP pipeline on a libuv worker thread, so the JS event loop is not blocked.

These functions resolve with the same shape as their synchronous counterparts and are Node-native-only. Browser code can instead use OfflineWorkerClient from @libraz/libsonare/worker; it provides task-based analysis and mastering in a Web Worker rather than these identically named functions.

Progress callbacks are not available on the async path. If you need progress updates, use the synchronous call with onProgress. If you only need concurrency, run several async calls in parallel.

FunctionReturn TypeDescription
analyzeAsync(samples, sampleRate?)Promise<AnalysisResult>Async variant of analyze(...)
masterAudioAsync(samples, sampleRate?, presetName?, overrides?)Promise<MasteringChainResult>Async variant of masterAudio(...)
masterAudioStereoAsync(left, right, sampleRate?, presetName?, overrides?)Promise<MasteringChainStereoResult>Async variant of masterAudioStereo(...)

Effects Functions ​

FunctionReturn TypeDescription
hpss(samples, sr?, kernelHarmonic?, kernelPercussive?, nFft?, hopLength?, hardMask?)HpssResultHarmonic-Percussive Source Separation; nFft=2048, hopLength=512, hardMask=false by default
hpssWithResidual(samples, sr?, kernelHarmonic?, kernelPercussive?, nFft?, hopLength?, hardMask?)HpssWithResidualResultHPSS with harmonic, percussive, and residual outputs; accepts the same STFT/mask options
harmonic(samples, sr?)Float32ArrayExtract harmonic component
percussive(samples, sr?)Float32ArrayExtract percussive component
timeStretch(samples, sampleRate, rate, nFft?, hopLength?)Float32ArrayTime-stretch without pitch change; defaults to nFft=2048, hopLength=512
phaseVocoder(samples, sampleRate, rate, nFft?, hopLength?)Float32ArrayDirect phase-vocoder time scaling
pitchShift(samples, sampleRate, semitones, nFft?, hopLength?)Float32ArrayPitch-shift without tempo change; defaults to nFft=2048, hopLength=512
remix(samples, intervals, sr?, alignZeros?)Float32ArrayReorder or concatenate sample intervals
remixAlignedIntervals(samples, intervals, sr?, alignZeros?)Int32ArrayThe cut points remix would use, without cutting: one clamped (start, end) pair per input interval, flattened. alignZeros defaults to true here
normalize(samples, sr?, targetDb?, mode?)Float32ArrayNormalize to target peak or RMS dB (mode: 'peak' or 'rms', default: 'peak')
normalizeStereo(request)NormalizeStereoResultNormalize a channel pair on a level measured across both channels. Request-only — takes NormalizeStereoRequest
trim(samples, sr?, thresholdDb?, frameLength?, hopLength?)Float32ArrayTrim silence (defaults: -60.0 dB, frameLength=2048, hopLength=512)
resample(samples, srcSr, targetSr)Float32ArrayResample to target sample rate
pitchCorrectToMidi(samples, sr, currentMidi, targetMidi)Float32ArrayRetune a held note from one MIDI pitch to another
pitchCorrectToMidiTimevarying(samples, f0Hz, targetMidi, sr?, hopLength?, voiced?, voicedProb?)Float32ArrayRetune a tracked pitch contour to a fixed note, frame by frame. voiced takes the VoicedFlags union
pitchCorrectTimevarying(samples, f0Hz, sr?, hopLength?, options?)Float32ArraySnap a tracked pitch contour to a scale or a fixed note; options is PitchCorrectOptions, whose voiced field takes the same VoicedFlags union
noteStretch(samples, sr?, options?)Float32ArrayTime-stretch a single note span in place; options is { onsetSample, offsetSample, stretchRatio }
voiceChange(samples, sr?, options?)Float32ArrayPitch + formant shift for voice transformation; options is { pitchSemitones, formantFactor }

trim(...) is the simple threshold edit helper. trimSilence(...) below is the librosa-compatible frame/RMS helper that returns the original sample range.

hpss(...) and hpssWithResidual(...) default their median-filter kernels to kernelHarmonic=31 and kernelPercussive=31. The request-object forms use the same names (nFft, hopLength, and hardMask) as the positional overloads.

Cutting a stereo take with remix ​

Zero-crossing snapping is a per-signal decision, so calling remix(...) channel by channel snaps each channel to a different frame and drifts a stereo take apart. Resolve one cut set from one channel with remixAlignedIntervals(...) and apply that same set to every channel. The two entry points point opposite ways by default: remix has alignZeros false, remixAlignedIntervals has it true.

Two guards stop a slice from vanishing under snapping. A signal with no sign change at all — silence, a DC offset, any constant — is not snapped, and a slice that had content but collapses to empty after snapping keeps its unsnapped boundaries.

Normalizing a channel pair ​

normalizeStereo({ left, right, sampleRate?, targetDb?, mode? }) applies one gain to both channels, so the stereo image is preserved. A per-channel gain would lift the quieter side until the two levels matched, which is a balance change rather than a normalization. That is why the result carries a single appliedGainDb and not a pair, and why a silent pair comes back untouched at exactly 0. mode: 'peak' (the default) drives the peak of the pair to targetDb, so the louder channel lands on it and the other keeps its distance below; mode: 'rms' drives the root mean square over both channels' samples together — the quadratic mean of the per-channel figures, not their average — and hard-clips the result to [-1, 1].

Unlike the mono normalize(...), whose targetDb default is 0 in both modes, normalizeStereo defaults targetDb by mode: 0 for 'peak' and -20 for 'rms'. 0 dBFS RMS is not a usable target, since the peaks sit well above the RMS and effectively all of them would clip.

VoicedFlags is Int32Array | Uint8Array | Float32Array | readonly number[] | readonly boolean[], so the boolean[] that PitchResult.voicedFlag hands back goes straight into pitch correction with no conversion step:

typescript
const pitch = pitchPyin(samples, sampleRate);
const tuned = pitchCorrectToMidiTimevarying(
  samples,
  pitch.f0,
  69,
  sampleRate,
  512,
  pitch.voicedFlag,   // boolean[] accepted as-is
);

voicedProb derives the voicing decision only when voiced is omitted — a frame at or above 0.5 counts as voiced. When voiced is supplied, voicedProb is ignored entirely; in particular it does not scale the per-frame correction amount, so passing both is identical to passing voiced alone. The same holds for PitchCorrectOptions.voicedProb on pitchCorrectTimevarying(...).

Correction strength is not weighted by voicedProb

Code that assumes the correction is scaled by voicedProb gets stronger correction in the low register than it expects. pYIN's voiced probability is a frequency-dependent observation mass that rises with the fundamental rather than tracking confidence, so as a weight it would under-correct low registers.

voiced and voicedProb must each be the same length as f0Hz. A mismatch throws a RangeError ('voiced must have the same length as f0Hz'), not a SonareError, so isSonareError does not catch it.

Feature Extraction Functions ​

FunctionReturn TypeDescription
stft(samples, sr?, nFft?, hopLength?)StftResultShort-Time Fourier Transform
stftDb(samples, sr?, nFft?, hopLength?)StftDbResultSTFT in decibels
melSpectrogram(samples, sr?, nFft?, hopLength?, nMels?)MelSpectrogramResultMel spectrogram
mfcc(samples, sr?, nFft?, hopLength?, nMels?, nMfcc?, fmin?, fmax?, htk?, lifter?)MfccResultMel-Frequency Cepstral Coefficients (lifter default 0 = no liftering)
chroma(samples, sr?, nFft?, hopLength?)ChromaResultChroma features
spectralCentroid(samples, sr?, nFft?, hopLength?)Float32ArraySpectral centroid per frame
spectralBandwidth(samples, sr?, nFft?, hopLength?)Float32ArraySpectral bandwidth per frame
spectralRolloff(samples, sr?, nFft?, hopLength?, rollPercent?)Float32ArraySpectral rolloff per frame
spectralFlatness(samples, sr?, nFft?, hopLength?)Float32ArraySpectral flatness per frame
spectralContrast(samples, sr?, nFft?, hopLength?, nBands?, fmin?, quantile?)Matrix2dResultSpectral contrast, shape (nBands + 1) x nFrames
spectralEdit(samples, sr, ops?, options?)Float32ArrayRegion-based STFT edit with gain, attenuate, mute, or heal ops
polyFeatures(samples, sr?, nFft?, hopLength?, order?)Matrix2dResultPer-frame polynomial spectral coefficients
zeroCrossingRate(samples, sr?, frameLength?, hopLength?)Float32ArrayZero-crossing rate per frame
zeroCrossings(samples, threshold?, refMagnitude?, pad?, zeroPos?)Int32ArrayZero-crossing sample indices
rmsEnergy(samples, sr?, frameLength?, hopLength?)Float32ArrayRMS energy per frame
pitchYin(samples, sr?, frameLength?, hopLength?, fmin?, fmax?, threshold?, fillNa?)PitchResultYIN pitch estimation; unvoiced f0 stays NaN unless fillNa is true
pitchPyin(samples, sr?, frameLength?, hopLength?, fmin?, fmax?, threshold?, fillNa?)PitchResultpYIN pitch estimation; unvoiced f0 stays NaN unless fillNa is true
pitchTuning(frequencies, resolution?, binsPerOctave?)numberTuning offset from frequencies
estimateTuning(samples, sr?, nFft?, hopLength?, resolution?, binsPerOctave?)numberTuning offset from audio
cqt(samples, sr?, hopLength?, fmin?, nBins?, binsPerOctave?)CqtResultConstant-Q transform magnitude
vqt(samples, sr?, hopLength?, fmin?, nBins?, binsPerOctave?, gamma?)CqtResultVariable-Q transform magnitude (gamma controls Q)
chromaCqt(samples, sr?, hopLength?, nChroma?){ nChroma, nFrames, data }Constant-Q chromagram (librosa.feature.chroma_cqt equivalent)
nnlsChroma(samples, sr?, options?){ nChroma, nFrames, data }NNLS chromagram (note-activation chroma); options.hopLength defaults to 512
decompose(s, nFeatures, nFrames, nComponents, nIter?, beta?, init?)DecomposeResultNMF (non-negative matrix factorization) factor matrices from a row-major spectrogram, with selectable init ('random' default, 'nndsvd')
decomposeStems(request)DecomposeStemsResultNMF separation that carries the original phase, so each component is directly listenable. Request-only — takes DecomposeStemsRequest
decomposeStemsLinked(request)DecomposeStemsLinkedResultShared NMF model and soft mask for one or more same-length channels (maximum 64); applies each mask to each channel's original complex spectrum, preserving interchannel level and phase. Request-only — returns components[k][c] with the same defaults as decomposeStems; one channel is bit-identical
noteSegments(request)NoteSegment[]Segment a caller-supplied monophonic F0 track into stable note regions. Request-only — takes NoteSegmentsRequest
hybridCqt(samples, sr?, hopLength?, fmin?, nBins?, binsPerOctave?)CqtResultHybrid CQT magnitude (true CQT in low bins, pseudo-CQT in high bins)
pseudoCqt(samples, sr?, hopLength?, fmin?, nBins?, binsPerOctave?)CqtResultApproximate (pseudo) CQT magnitude (single FFT)
bassChroma(samples, sr?, hopLength?, nChroma?)ChromaResultBass-focused chroma (low-register pitch-class distribution)
chromaCens(samples, sr?, hopLength?, nChroma?)ChromaResultCENS energy-normalized/smoothed chroma
onsetStrengthMulti(samples, sr?, nFft?, hopLength?, nMels?, nBands?){ nBands, nFrames, data }Multi-band onset strength (nBands default 3; data row-major [nBands x nFrames])
nnFilter(s, nFeatures, nFrames, aggregate?, k?, width?)Matrix2dResultNearest-neighbor filtering
onsetEnvelope(samples, sr?, nFft?, hopLength?, nMels?)Float32ArrayOnset strength envelope — how sharply energy rises per frame; the input to the tempogram family

Common defaults: nFft=2048, hopLength=512, nMels=128, nMfcc=20, pitch fmin=65.0, fmax=2093.0, threshold=0.1, and rollPercent=0.85.

CQT/VQT use fmin=32.70319566 Hz (C1), nBins=84, and binsPerOctave=12. VQT's default gamma=-1 selects automatic ERB-derived bandwidth. chromaCqt defaults to nChroma=12, nBins=252, and binsPerOctave=36; bassChroma and chromaCens default to nChroma=12. onsetStrengthMulti defaults to nBands=3. decompose defaults to nIter=50, beta=2, and init='random'.

decompose versus decomposeStems ​

decompose(...) returns the W/H factors of a magnitude spectrogram. Those factors carry no phase, so reconstructing from them needs a phase estimator, and an estimated phase does not hold up as a stem. decomposeStems(...) builds a per-component soft mask from the same factorization and applies it to the original complex spectrogram, so every component keeps the source's phase. The masks sum to one wherever the model has energy and the inverse STFT is linear, so the components sum back to the input.

maskPower sets how hard the mask separates: 1 (the default) keeps the magnitude ratio, 2 is the Wiener-style power ratio, which separates harder at the cost of more artifacts on overlapping partials. decomposeStems defaults to nComponents=4, nFft=2048, hopLength=512, nIter=100, beta=2, and init='random', and returns components — one signal per component, each the length of the input — alongside the w/h matrices and sampleRate.

For multichannel input, decomposeStemsLinked({ channels, sampleRate, ... }) fits one NMF model and one mask set from the channels' averaged magnitudes, then applies each mask unchanged to each channel's original complex spectrum. This preserves interchannel level and phase. channels must contain at least one same-length Float32Array and at most 64 channels. The result uses components[k][c]; w and h describe the shared factorisation. Its option defaults match decomposeStems, and a one-channel call is bit-identical.

typescript
import { decomposeStemsLinked } from '@libraz/libsonare-native';

const linked = decomposeStemsLinked({
  channels: [leftChannel, rightChannel],
  sampleRate,
});
const firstLeft = linked.components[0][0];
const firstRight = linked.components[0][1];

NNDSVD factors will not match single-precision ones

NNDSVD seeding is computed in double precision. A magnitude spectrogram's trailing singular vectors sit at single precision's noise floor, so a float seed would depend on summation order, and wasm32 and arm64 would answer with different components; the double-precision seed makes the result reproducible across builds. Shapes, non-negativity and reconstruction quality are unaffected — this is reproducibility, not accuracy — but stored factors, or a stem render made with a single-precision seed, will not match.

noteSegments ​

noteSegments({ f0Hz, voicedProb, frameRate, ... }) segments a caller-supplied monophonic F0 track into stable note regions, one NoteSegment per region with frameStart, frameEnd, startSeconds, endSeconds, and medianCents. The tuning fields are segmentationThresholdCents (default 50), minNoteMs (default 30), referenceHz (default 440), and voicedThreshold (default 0.5), the value of voicedProb at or above which a frame counts as voiced.

Do not feed pYIN's voicedProb straight in

voicedProb is the frame's voiced observation mass, and for a fixed frame length it rises with F0 rather than tracking confidence. A fixed threshold therefore silently returns no segments at all for low-register material. Pass pitchPyin's voicedFlag converted to 0/1, or lower voicedThreshold.