C++ Analysis API
Feature extraction functions and analysis result types for the libsonare C++ interface; see C++ API Reference for the rest of the C++ surface.
Feature Extraction
MelSpectrogram Medium
MelConfig config;
config.n_mels = 128;
config.n_fft = 2048;
config.hop_length = 512;
auto mel = MelSpectrogram::compute(audio, config);
// Power spectrum [n_mels x n_frames]
auto power = mel.power();
// Convert to dB
auto db = mel.to_db();
// MFCC
auto mfcc = mel.mfcc(13); // 13 coefficientsChroma Medium
ChromaConfig config;
config.n_chroma = 12;
auto chroma = Chroma::compute(audio, config);
// Features [12 x n_frames]
auto features = chroma.features();
// Mean energy per pitch class
auto energy = chroma.mean_energy();Spectral Features
// Per-frame spectral centroid (Hz)
std::vector<float> spectral_centroid(const Spectrogram& spec, int sr);
// Per-frame spectral bandwidth (Hz)
std::vector<float> spectral_bandwidth(const Spectrogram& spec, int sr);
// Per-frame spectral rolloff (Hz)
std::vector<float> spectral_rolloff(const Spectrogram& spec, int sr, float roll_percent = 0.85f);
// Per-frame spectral flatness
std::vector<float> spectral_flatness(const Spectrogram& spec);
// Zero crossing rate
std::vector<float> zero_crossing_rate(const Audio& audio, int frame_length, int hop_length);
// RMS energy
std::vector<float> rms_energy(const Audio& audio, int frame_length, int hop_length);
// Spectral contrast (difference between peaks and valleys in frequency bands)
std::vector<float> spectral_contrast(const Spectrogram& spec, int sr, int n_bands = 6,
float fmin = 200.0f, float quantile = 0.02f);Pitch Tracking Medium
PitchConfig config;
config.frame_length = 2048;
config.hop_length = 512;
config.fmin = 65.0f; // C2
config.fmax = 2093.0f; // C7
config.threshold = 0.1f;
// YIN algorithm
PitchResult yin = yin_track(audio, config);
// pYIN algorithm (probabilistic YIN with HMM smoothing)
PitchResult pyin_result = pyin(audio, config);
// Access results
float median = pyin_result.median_f0();
float mean = pyin_result.mean_f0();
const std::vector<float>& f0 = pyin_result.f0;
const std::vector<bool>& voiced = pyin_result.voiced_flag;CQT / VQT Medium
Constant-Q Transform and Variable-Q Transform for music analysis.
CqtConfig config;
config.fmin = 32.7f; // C1
config.n_bins = 84; // 7 octaves
config.bins_per_octave = 12; // Semitone resolution
auto cqt_result = cqt(audio, config);
// Access magnitude [n_bins x n_frames]
auto mag = cqt_result.magnitude();
auto power = cqt_result.power();
// Variable-Q Transform (with variable Q factor)
VqtConfig vqt_config;
vqt_config.gamma = 0.0f; // 0 = CQT behavior
auto vqt_result = vqt(audio, vqt_config);Thread Safety
CqtResult and VqtResult objects use lazy initialization for cached results. They are not thread-safe for concurrent access. Create separate copies for multi-threaded use.
Griffin-Lim
Griffin-Lim is an iterative phase-reconstruction algorithm: it recovers a plausible waveform from a magnitude-only spectrum (which stores no phase) by transforming to the time domain and back over successive passes, refining the phase estimate each time. A phase vocoder instead tracks and manipulates the STFT phase directly.
Deprecated Functions
The inverse transform functions icqt() and ivqt() are deprecated in the current C++ headers. Prefer Griffin-Lim or phase-vocoder based reconstruction paths for new code.
// Deprecated - do not use in new code
[[deprecated("Use Griffin-Lim or phase vocoder for better reconstruction quality")]]
Audio icqt(const CqtResult& cqt_result, int length = 0);
[[deprecated("Use griffinlim_vqt or phase vocoder for better reconstruction quality")]]
Audio ivqt(const VqtResult& vqt_result, int length = 0);Migration: griffinlim_cqt and griffinlim_vqt are declared in the same <feature/cqt.h> / <feature/vqt.h> headers as cqt() and vqt(), so no extra include is needed. For preview audio reconstruction, use these Griffin-Lim paths, or keep phase information in your own STFT-domain pipeline when quality matters.
const auto& cqt_magnitude = cqt_result.magnitude();
auto reconstructed = griffinlim_cqt(cqt_magnitude.data(), cqt_result.n_bins(),
cqt_result.n_frames(), config,
cqt_result.sample_rate());
auto reconstructed_vqt = griffinlim_vqt(vqt_result, vqt_result.sample_rate());NNLS Chroma
NNLS chroma uses a configurable CQT hop length. The default is 512 samples; set it on NnlsChromaConfig::cqt when matching another frame grid.
NnlsChromaConfig nnls_config;
nnls_config.cqt.hop_length = 512;
nnls_config.enable_stft_blend = true;
auto nnls_result = nnls_chroma(audio, nnls_config);Types
Key
struct Key {
PitchClass root; // C=0, Cs=1, ..., B=11
Mode mode; // Major, Minor, Dorian, Phrygian, Lydian, Mixolydian, Locrian
float confidence; // 0.0 - 1.0
std::string to_string() const; // "C major"
std::string to_short_string() const; // "C", "Am"
};Chord
struct Chord {
PitchClass root;
ChordQuality quality; // Major, Minor, Dim, Aug, 7th, etc.
float start; // seconds
float end; // seconds
float confidence;
PitchClass bass; // Bass pitch class for inversion notation
std::string to_string() const; // "C", "Am", "G7"
};Section
struct Section {
SectionType type; // Intro, Verse, Chorus, etc.
float start;
float end;
float energy_level;
float confidence;
std::string type_string() const;
float duration() const;
};AnalysisResult
struct AnalysisResult {
float bpm;
float bpm_confidence;
std::vector<BpmCandidateHypothesis> bpm_candidates;
Key key;
TimeSignature time_signature;
std::vector<TimeSignature> time_signature_candidates;
std::vector<Beat> beats;
std::vector<Chord> chords;
std::vector<Section> sections;
Timbre timbre;
Dynamics dynamics;
RhythmFeatures rhythm;
MelodyContour melody;
std::string form; // "IABABCO"
};The two perceptual sub-structs are expanded below; RhythmFeatures and MelodyContour follow the same by-value pattern.
Timbre
Perceptual sound-color descriptors, each normalized to [0, 1].
struct Timbre {
float brightness; // high = bright / harsh
float warmth; // high = warm / full
float density; // high = rich / complex
float roughness; // high = rough / harsh
float complexity; // high = harmonically complex
};Dynamics
Loudness and dynamic-range descriptors (levels in dB).
struct Dynamics {
float dynamic_range_db; // dynamic range (dB)
float peak_db; // peak level (dB)
float rms_db; // RMS level (dB)
float crest_factor; // peak-to-RMS ratio
float loudness_range_db; // loudness range / LRA (dB)
bool is_compressed; // true if the audio appears heavily compressed
};