Skip to content

JavaScript/TypeScript Analysis API ​

This page covers the core analysis functions for the libsonare JavaScript/TypeScript package: BPM/key/beat/onset detection, the all-in-one analyze() call and its focused helpers, structural boundary and meter estimation, plus scale quantization, unit conversion, and resampling. Two related references live alongside it: Feature Extraction covers STFT, mel/MFCC, chroma, spectral features, pitch detection, and the CQT/VQT/decomposition family, and librosa-Compatible Helpers covers the argument-for-argument librosa ports for pre-emphasis, test-signal generation, spectral reconstruction, structure, segmentation, and tempograms.

Where the sections went ​

SectionPage
Feature ExtractionFeature Extraction
librosa-Compatible Helperslibrosa-Compatible Helpers

Analysis Functions ​

detectBpm(samples, sampleRate) ​

Detect BPM (tempo) from audio samples.

Use Cases

  • DJ Software: Match tempos between tracks for seamless mixing
  • Music Players: Display tempo information, auto-generate playlists by tempo
  • Fitness Apps: Match music to workout intensity
  • Beat Sync: Synchronize visualizations or animations to music
typescript
function detectBpm(samples: Float32Array, sampleRate?: number): number
ParameterTypeDescription
samplesFloat32ArrayMono audio samples (range -1.0 to 1.0)
sampleRate?numberSample rate in Hz (default: 22050; e.g., 44100)

Always pass the real sample rate

Although sampleRate is optional here (defaulting to 22050 Hz), decoded browser audio is almost always 44100 or 48000 Hz. Pass the buffer's actual audioBuffer.sampleRate, or the reported BPM will be wrong. The same holds for detectKey, detectBeats, and analyze: their sampleRate is optional with the same 22050 Hz default, so pass the real rate to those as well.

Returns: Detected BPM as a number.

typescript
const bpm = detectBpm(samples, sampleRate);
console.log(`BPM: ${bpm}`);  // "BPM: 120"

detectKey(samples, sampleRate) ​

Detect musical key from audio samples. Returns the root note (C, D, E...) and mode (major/minor).

Use Cases

  • Harmonic Mixing: DJs match keys for smooth transitions (Camelot wheel)
  • Transposition: Suggest key changes to match vocal range
  • Music Recommendation: Find songs in compatible keys
  • Practice Tools: Display key for musicians to play along
typescript
function detectKey(samples: Float32Array, sampleRate?: number): Key  // sampleRate default: 22050

Returns: Key object

typescript
interface Key {
  root: PitchClass;      // 0-11 (C=0, B=11)
  mode: Mode;            // Major, Minor, or modal value; see Mode enum
  confidence: number;    // posterior share in [0, 1); see below
  name: string;          // "C major", "A minor"
  shortName: string;     // "C", "Am"
}

const KeyProfile = {
  KrumhanslSchmuckler: 0,
  Temperley: 1,
  Shaath: 2,
  FaraldoEDMT: 3,
  FaraldoEDMA: 4,
  FaraldoEDMM: 5,
  BellmanBudge: 6,
} as const;
typescript
const key = detectKey(samples, sampleRate);
console.log(`Key: ${key.name}`);        // "C major"
console.log(`Confidence: ${(key.confidence * 100).toFixed(1)}%`);

What confidence measures

confidence is a softmax over every scored candidate's profile correlation. Its range is [0, 1), all candidates' confidences sum to 1, and with 24 candidates in the set no single share can reach 1 — a decisive result looks like 0.3, not 0.95.

It says how sharply the chroma picked one candidate out of the set, not how often that pick is right: nothing here is calibrated against annotated recordings, so a high share on a modally ambiguous piece is a confident wrong answer, not a reliable one. Compare shares against each other rather than against a fixed threshold you carried over from a calibrated classifier.

KeyCandidate.key.confidence carries the same posterior share, while KeyCandidate.correlation stays the raw profile correlation the softmax was computed from.

detectBeats(samples, sampleRate) ​

Detect beat times from audio samples. Returns exact timestamps of each beat.

Use Cases

  • Music Visualization: Trigger effects on each beat
  • Rhythm Games: Generate note charts from audio
  • Video Editing: Auto-cut to the beat
  • Loop Creation: Find perfect loop points
typescript
function detectBeats(samples: Float32Array, sampleRate?: number): Float32Array  // sampleRate default: 22050

Returns: Float32Array of beat times in seconds

typescript
const beats = detectBeats(samples, sampleRate);
console.log(`Found ${beats.length} beats`);
for (let i = 0; i < beats.length; i++) {
  console.log(`Beat ${i + 1}: ${beats[i].toFixed(3)}s`);
}

detectOnsets(samples, sampleRate) ​

Detect onset times (note attacks) from audio samples. More granular than beats - captures every note/hit.

Use Cases

  • Drum Transcription: Detect individual drum hits
  • Audio-to-MIDI: Convert audio to note events
  • Sample Slicing: Automatically segment audio at transients
typescript
function detectOnsets(samples: Float32Array, sampleRate?: number): Float32Array  // sampleRate default: 22050

Returns: Float32Array of onset times in seconds

analyze(samples, sampleRate) Heavy ​

Perform the all-in-one music analysis. Returns BPM, key, beats, chords, sections, timbre, and more.

Use Cases

  • Music Library Management: Auto-tag songs with metadata
  • Music Production: Analyze reference tracks
  • DJ Preparation: Get all track info at once
  • Music Education: Study song structure

Performance

This is the heaviest API. For long audio files (>3 minutes), consider using analyzeWithProgress to show progress, or analyze only relevant segments.

typescript
function analyze(request: MusicAnalyzeRequest): AnalysisResult
function analyze(
  samples: Float32Array,
  sampleRate?: number,             // default 22050
  options?: MusicAnalyzeOptions,
): AnalysisResult

interface MusicAnalyzeOptions {
  // Framing
  nFft?: number;                   // default 2048
  hopLength?: number;              // default 512

  // Tempo and beats
  bpmMin?: number;                 // default 60
  bpmMax?: number;                 // default 200
  startBpm?: number;               // default 120
  adaptiveTempo?: boolean;         // track a local tempo prior; default false
  tempoUpdateIntervalBeats?: number; // local tempo context, in beats; default 8
  computeTempoCurve?: boolean;     // fill beatLocalBpm; default false

  // Meter
  meterCandidateNumerators?: number[]; // default [3, 4, 6]; 1-16 entries, each 2-32
  meterDenominator?: number;       // default 4; a power of two in [1, 32]

  // Chroma and chords
  useTriadsOnly?: boolean;         // default true here; see below
  useHpss?: boolean;               // default true
  chromaHighpassHz?: number;       // default 80; 0 disables
  useBassWeighted?: boolean;       // default true
  chromaHopMultiplier?: number;    // default 4
  useChordHmm?: boolean;           // default false
  useChordKeyContext?: boolean;    // default false
  chordHmmBeamWidth?: number;      // default 24
  detectChordInversions?: boolean; // default false
}

Returns: Complete AnalysisResult. A single analyze() call returns the full result — chords, sections, timbre, dynamics, rhythm, melody, form, and per-beat strength — on every binding, so you rarely need the focused helpers unless you only want one field. Chords in the result carry romanNumeral relative to result.key; it is an empty string for N.C..

useTriadsOnly defaults to true in analyze()

The unified path defaults useTriadsOnly to true even though the standalone chord API defaults it to false. Until you pass useTriadsOnly: false, analyze() searches triads alone — a seventh chord comes back as the triad inside it, and a result that looks harmonically flat next to detectChords() on the same audio is usually this and not a detection failure.

meterCandidateNumerators widens the search, it does not force the result: the estimator still has to find support for a numerator before reporting it. But the reverse holds absolutely — a numerator you did not ask for cannot be reported, so a 5/4 or 7/8 piece analysed with the default [3, 4, 6] comes back as a four. meterDenominator is reported as requested; the estimator will still report 8 on its own when it resolves a compound meter.

computeTempoCurve is the switch that fills beatLocalBpm, and it is off because the curve is an extra output rather than a refinement. The curve describes the beat grid it was decoded from, so a track whose tempo genuinely drifts wants adaptiveTempo as well — otherwise you get a local reading of a grid that was laid down at a fixed tempo.

typescript
const result = analyze(samples, sampleRate);
console.log(`BPM: ${result.bpm}`);
console.log(`Key: ${result.key.name}`);
console.log(`Chords: ${result.chords.length}`);
console.log(`Form: ${result.form}`);  // e.g., "IABABCO"

analyzeWithProgress(samples, sampleRate, onProgress) Heavy ​

Perform the same all-in-one analysis with progress reporting.

typescript
function analyzeWithProgress(
  samples: Float32Array,
  sampleRate: number | undefined,  // undefined applies the 22050 default
  onProgress: (progress: number, stage: string) => void
): AnalysisResult

sampleRate is positional (before the callback) but accepts undefined, which uses the same 22050 Hz default as analyze. Pass the buffer's real rate.

Progress Stages:

StageDescriptionProgress
"features"Feature precomputation0.0
"bpm"BPM detection0.15
"key"Key detection0.15
"beats"Beat tracking0.25
"chords"Chord recognition0.40
"sections"Section detection0.55
"timbre"Timbre analysis0.70
"dynamics"Dynamics analysis0.80
"rhythm"Rhythm analysis0.90
"melody"Melody contour extraction0.95
"complete"Finished1.0
typescript
const result = analyzeWithProgress(samples, sampleRate, (progress, stage) => {
  console.log(`${stage}: ${Math.round(progress * 100)}%`);
});

transcribe(request) ​

transcribe turns mono audio into note-on/note-off ProjectMidiEvent[] on a constant-tempo PPQ grid. Pass tempoBpm to choose that grid, or omit it to detect the tempo. The result reports noteCount separately (always half events.length) and does not install a tempo map on a project.

typescript
interface TranscribeRequest {
  samples: Float32Array;
  sampleRate: number;
  tempoBpm?: number;  // omit to detect
  polyphonic?: boolean;
  // Other TranscribeOptions detector settings are optional.
}

interface TranscribeResult {
  events: ProjectMidiEvent[];
  noteCount: number;
  tempoBpm: number;
}

function transcribe(request: TranscribeRequest): TranscribeResult
typescript
const { events, noteCount, tempoBpm } = transcribe({
  samples,
  sampleRate,
  tempoBpm: 120,
  polyphonic: true,
});
project.setMidiEvents(clipId, events);
console.log(noteCount, tempoBpm);

When the PPQ grid should follow a project tempo map, use Project.transcribeToClip(...) instead; its request has no tempoBpm and it returns the number of notes written.

Focused analysis helpers ​

One call is usually enough

analyze() already returns chords, sections, timbre, dynamics, rhythm, melody, form, and per-beat strength. Reach for a focused helper only when you want a single field or need options the high-level call hides.

Use the focused helpers when the default analyze(...) result is either too broad or not detailed enough. They share the same mono Float32Array input model but expose options that are hidden by the high-level call.

TaskFunctionNotes
Downbeat/bar startsdetectDownbeats(samples, sampleRate)Returns seconds for likely bar starts. Pair with detectBeats for grid displays.
Ranked key candidatesdetectKeyCandidates(samples, sampleRate, options?)Useful when the top key is ambiguous or when you want profile/mode filtering.
Detailed tempo candidatesanalyzeBpm(samples, sampleRate, ...)Returns the best BPM plus alternate candidates and tempo evidence.
Rhythm characteranalyzeRhythm(samples, sampleRate, ...)Reports groove, syncopation, and regularity style features.
DynamicsanalyzeDynamics(samples, sampleRate, ...)Dynamic range, loudness range, crest factor, and compression flag.
TimbreanalyzeTimbre(samples, sampleRate, ...)Brightness, warmth, density, roughness, and complexity.
ChordsdetectChords(samples, sampleRate, options?)Returns { chords } of chord segments; options include HMM smoothing, key context, inversions, and chromaMethod: 'stft' | 'nnls'.
SectionsanalyzeSections(samples, sampleRate, ...)Song-structure sections such as intro, verse, chorus, bridge, and outro. Long inputs keep accurate start / end times even when the internal boundary grid is pooled.
Structural boundariesdetectBoundaries(request)The transitions alone, plus the novelty curve they were picked from. Use it when you want to apply your own threshold.
MelodyanalyzeMelody(samples, sampleRate, ...)Monophonic melody contour based on pitch tracking.
typescript
const keys = detectKeyCandidates(samples, sampleRate, {
  modes: [Mode.Major, Mode.Minor],
  profile: 'krumhansl',
  genreHint: 'pop',
});

const { chords } = detectChords(samples, sampleRate, {
  useHmm: true,
  useKeyContext: true,
  keyRoot: keys[0].key.root,
  keyMode: keys[0].key.mode,
  chromaMethod: 'nnls',
});

const sections = analyzeSections(samples, sampleRate);

detectBoundaries(request) ​

Where analyzeSections hands back labelled spans, detectBoundaries hands back the transitions themselves and the novelty curve they were picked from. That curve is the reason to reach for this one: a caller that wants its own threshold cannot recover it from a list of finished sections.

typescript
const { boundaries, noveltyCurve, noveltyPeak, sampleRate: gridRate } =
  detectBoundaries({ samples, sampleRate: 44100 });

for (const { time, strength } of boundaries) {
  console.log(`${time.toFixed(2)}s  ${strength.toFixed(3)}`);
}
OptionDefaultDescription
sampleRate22050Sample rate of samples
nFft2048FFT size for the structural features
hopLength512Hop length in samples
kernelSize64Checkerboard kernel size, in frames
threshold0.3Relative novelty threshold, applied after the curve is scaled by its own maximum
absoluteThreshold0.005Novelty floor applied to the raw response before that scaling
nMfcc13Number of MFCC coefficients
nChroma12Number of chroma bins
peakDistance2.0Minimum spacing between peaks, in seconds
useMfcctrueUse MFCC features
useChromatrueUse chroma features

The two thresholds do different jobs

threshold asks how prominent a peak is within this track; because the curve is scaled by its own maximum first, it says nothing about how much the features actually changed. absoluteThreshold is the floor that asks whether anything changed at all. Setting it to 0 does not make the detector more sensitive in a useful way — self-scaling turns residual fluctuation into peaks of 1.0, so a stationary input then segments anyway.

Lowering the floor also will not recover level-only structure. A level change turns the feature vector about five times less than a comparable pitch change, which lands it below what steady noise produces, so the noise is admitted first.

The result carries the grid it was measured on, not the source's: input above 22050 Hz is resampled before any feature is computed, so boundary.frame is uninterpretable without the returned sampleRate, hopLength, nFrames, and frameStride. Use boundary.time for sample or second mapping either way. noveltyCurve is scaled by its own maximum; recover the raw response with noveltyCurve[i] * noveltyPeak.

Passing useMfcc: false together with useChroma: false throws SonareError with InvalidParameter: the two feature streams are combined frame for frame, so with neither enabled there is nothing to combine.

estimateMeter(request) ​

Score a beat series against a set of candidate time signatures and report the winner, the phase of the first measure, and how the measure divides internally. It takes beats and accents rather than audio, so it runs on a grid you already have — from analyze(), from a DAW, or from a tapped tempo — without a second pass over the samples. For what a meter is and how grouping differs from a time signature, see Meter and grouping.

typescript
function estimateMeter(request: EstimateMeterRequest): MeterEstimate

interface EstimateMeterRequest {
  beatTimes: ArrayLike<number>;      // Beat positions in seconds, non-decreasing
  beatStrengths: ArrayLike<number>;  // One accent value per beat, same length
  candidateNumerators?: number[];    // default [3, 4, 6]; 1-16 entries, each 2-32
  denominator?: number;              // default 4; a power of two in [1, 32]
  downbeatWeight?: number;           // default 1
  measureWeight?: number;            // default 0.5
  subdivisionWeight?: number;        // default 0.15
  compoundSubdivisionThreshold?: number;  // default 0.85
}

interface MeterEstimate {
  timeSignature: TimeSignature;   // The selected signature; confidence is margin-derived
  downbeatPhase: number;          // Beat index the first measure starts on
  searched: boolean;              // false when the series was too short to score
  grouping: number[];             // Beats per accent group; sums to the numerator
  candidateScores: number[];      // One score per requested numerator, in order
  candidates: TimeSignature[];    // Ranked by descending support; confidence is a share
}

There is no positional form — estimateMeter takes a request object only. Two sources are supported for beatStrengths: AnalysisResult.beatObservations.onsetStrength, the intended one, and beats[].strength, the raw single-frame value described under Beat. Neither needs pre-scaling — the series is divided by its own maximum before scoring, so only the accent contrast within it is read.

A hand-assembled accent series is sample-rate dependent

Reading onsetEnvelope at timeToFrames(beatTime, sr, hopLength) for each beat is not a third source. A hop counted in samples frames a different span of time at each rate, so one waveform sampled at 32000, 44100 and 48000 Hz — beat times identical to the sample — produced winning numerators of 6, 3 and 4, and widening the read to a window around each beat does not remove the dependence. A browser decodes at the output device's rate, so a series built this way gives a different answer per visitor for the same clip. Use one of the two sources above.

An odd meter is only reported if you asked for its numerator

The default candidate set is {3, 4, 6}. A numerator outside the set cannot win, so a 5/4 or 7/8 piece analysed with the defaults comes back as a four with no indication that the right answer was never on the ballot. Widening the set does not force an odd reading — the estimator still needs support for it — so pass the numerators your material might actually use.

searched: false means every other field is a fallback

When the beat series holds fewer than eight beats, there is not enough to score any candidate and the estimator returns a fixed fallback instead of a result: 4/<requested denominator>, downbeatPhase 0, grouping [4], all-zero candidateScores, and a single-entry candidates.

timeSignature.confidence is part of the fallback — it reads 0, so an unchecked read degrades toward "no idea" rather than toward a middling detection. searched is still what separates a fallback from a result; check it before you show a confidence or branch on one. An empty beatTimes throws rather than falling back; one to seven beats returns the fallback.

timeSignature.confidence and candidates[k].confidence share a field name but carry different quantities. On timeSignature it is derived from the margin over the runner-up; on a candidates entry it is that candidate's share of the summed support, so the entries sum to one. The two are not comparable and must not share a threshold — read the field you mean rather than whichever is to hand.

grouping is where an aksak meter shows itself: [3, 2, 2] is a 7/8 grouped three-two-two, while [2, 2] is an ordinary four. A single entry means no internal division was resolved — either the numerator has none, or the measure was too wide to search, or the series was too short — so [4] and [2, 2] are different answers, not two spellings of one.

candidateScores and candidates are not indexed alike, and mixing them up is the easy mistake here. candidateScores[k] is the score of the k-th requested numerator, in the order you passed them; candidates[k] is the k-th best hypothesis, ordered by descending support. The scores are standardized and signed — zero is the no-meter level and a negative score sits below the noise — and they grow with the square root of the number of beats scored, so compare them within one result, never across two.

chordFunctionalAnalysis(samples, keyRoot, keyMode, sampleRate?, options?) ​

Functional (Roman-numeral) harmonic analysis of the detected chord progression, relative to the given key. It runs chord detection internally and labels each detected chord, so pass the same keyRoot/keyMode you get from detectKey(...) and the same options you would give detectChords(...).

The labels carry no times; they pair with detectChords(...) one for one only when both calls get the same options, and different options can give a list of a different length. analyze(...) performs a separate chord detection with its own settings, so its chord list does not pair with these labels. Each chord in the analyze(...) result carries its own romanNumeral, relative to result.key; it is an empty string for N.C..

typescript
function chordFunctionalAnalysis(
  samples: Float32Array,
  keyRoot: PitchClass,
  keyMode?: Mode,
  sampleRate?: number,
  options?: ChordDetectionOptions,
): string[]   // one Roman-numeral label per detected chord, e.g. ["I", "IV", "V", "vi"]
typescript
const key = detectKey(samples, sampleRate);
const roman = chordFunctionalAnalysis(samples, key.root, key.mode, sampleRate);
console.log(roman);  // e.g. ["I", "IV", "V", "vi"]

detectKey(...) and detectKeyCandidates(...) accept the same KeyDetectionOptions includes:

Option groupValues
Controlsmodes, profile, genreHint, useHpss, loudnessWeighted, highPassHz
Profile namesks, krumhansl, temperley, shaath, keyfinder, faraldo-edmt / edmt, faraldo-edma / edma, faraldo-edmm / edmm, bellman-budge / bellman
Genre hintsauto, edm, electronic, dance, pop, classical, jazz

Scale Quantization ​

12-TET (twelve-tone equal temperament) scale helpers for building pitch-correction targets. modeMask is a 12-bit mask where bit i enables the i-th pitch class relative to root (a PitchClass, C = 0); natural major is 0b101010110101. referenceMidi is the tuning anchor (pass 0 for A4 = 69).

typescript
// Snap a (possibly fractional) MIDI number to the nearest enabled pitch class
function scaleQuantizeMidi(root: number, modeMask: number, midi: number, referenceMidi?: number): number
// Correction (quantized − input), in semitones
function scaleCorrectionSemitones(root: number, modeMask: number, midi: number, referenceMidi?: number): number
// Is pitchClass (0..11) enabled by modeMask relative to root?
function scalePitchClassEnabled(root: number, modeMask: number, pitchClass: number): boolean

Pair scaleQuantizeMidi(...) with pitchCorrectToMidi(...) to retune a detected note to the nearest scale degree.

Unit Conversion ​

These functions are lightweight and fast.

typescript
// Hz <-> Mel (Slaney formula)
function hzToMel(hz: number): number
function melToHz(mel: number): number

// Hz <-> MIDI note number (A4 = 440 Hz = 69)
function hzToMidi(hz: number): number
function midiToHz(midi: number): number

// Hz <-> Note name
function hzToNote(hz: number): string      // "A4", "C#5"
function noteToHz(note: string): number

// Time <-> Frames
function framesToTime(frames: number, sr: number, hopLength: number): number
function timeToFrames(time: number, sr: number, hopLength: number): number

// Frames <-> Samples (librosa.frames_to_samples / samples_to_frames)
function framesToSamples(frames: number, hopLength?: number, nFft?: number): number
function samplesToFrames(samples: number, hopLength?: number, nFft?: number): number

// dB conversions (vectorised)
function powerToDb(values: Float32Array, ref?: number, amin?: number, topDb?: number): Float32Array
function amplitudeToDb(values: Float32Array, ref?: number, amin?: number, topDb?: number): Float32Array
function dbToPower(values: Float32Array, ref?: number): Float32Array
function dbToAmplitude(values: Float32Array, ref?: number): Float32Array

Resampling ​

resample(samples, srcSr, targetSr) Medium ​

High-quality resampling using r8brain algorithm.

typescript
function resample(
  samples: Float32Array,
  srcSr: number,
  targetSr: number
): Float32Array