Skip to content

Physical Modeling ​

Physical modeling is usually met as a marketing word, and the natural guess is that it means "very good samples". It means close to the opposite. A sampled instrument holds a recording of a sound that already happened. A physical model holds a description of the object that would make the sound, and computes the note while you are listening to it. There is no recording anywhere in the signal path, because nothing was ever recorded.

This page explains what that buys, what it costs, and where it stops working. It is concepts only — no code.

Two answers to the same question

Asked "what does a cello sound like", a sampler answers with a recording of one. A model answers with a string, a bow and a body, and lets the arithmetic decide. The first is accurate about one performance; the second is accurate about the mechanism, which is a different kind of accuracy and behaves differently at the edges.

Simulating the object ​

A vibrating string is not a waveform. It is a medium that carries waves which travel along it, reflect at each end, lose a little energy on every trip, and interfere with each other on the way. The tone you hear is what is left over after all of that, radiated through a bridge and a body. A model reproduces that process, sample by sample, and takes whatever comes out.

The consequence worth internalising is that a model has no notion of a "note" as a stored thing. It has a state — how much energy is where, in which direction it is moving — and a rule for advancing that state by one sample. Play a note, and the state evolves. Push on it halfway through, and the state evolves differently from there. Nothing is looked up.

The digital waveguide ​

The workhorse of the family is the digital waveguide, and it is simpler than the name suggests. Three parts:

  • A delay line stands in for the travel. A wave leaving one end takes a fixed number of samples to reach the other, and that number is the pitch: a shorter trip means a faster round trip means a higher note. These engines have no tuning oscillator, because length does the tuning.
  • A loss filter stands in for what the trip costs. Every reflection loses energy, and loses more of it at high frequencies than at low — which is exactly why a plucked note goes dull before it goes quiet. One filter in the loop reproduces the whole decay behaviour.
  • An exciter puts energy in. A hammer, a plectrum, a bow, a reed, a pair of lips, a jet of air.

These roles describe a family of architectures; they do not imply identical delay lines or feedback topology.

The waveguide loop
energy inthe wavecomes backtapExciterDelay line (travel)Loss filter(reflection)Radiation
Pitch follows each resonator's delay and feedback topology, and decay follows its losses. Bow, reed, and lip models also differ in bore or string, reflection sign and coefficients, loss filters, and radiation, so brightness and damping are model-specific controls.

These models share delay and feedback concepts, so delay, reflection, loss, and excitation are useful terms across the family. They are not one loop with only the exciter swapped. A bowed string uses two delay lines meeting at the bow. A reed uses one bore whose cylindrical and conical topologies change the feedback sign and period. Brass uses a full-period bore with a two-pole lip resonator and its own bell reflection. brightness and damping therefore map to each model's own reflection and loss stages.

Expression becomes continuous, not switched ​

This is the practical difference, and it is larger than it sounds.

A sampler is a set of recordings arranged in a grid — a few notes across the keyboard, a few velocity layers stacked at each. Playing harder means selecting a different recording. However many layers were recorded, the grid is finite, and everything between two layers is a crossfade between two performances that never happened together. Worse, once a note has started, the recording is chosen: the sampler can fade it, filter it, or start another one, but it cannot make the note that is already sounding have been played harder.

A model has no grid. Force on the bow is a number the loop reads on every sample, so pressing harder mid-note makes the string grip differently from that sample onward — and the brightness that comes with it is not an effect applied on top, it is the consequence of the string being driven harder. You push on the bow instead of picking a louder recording, and the instrument responds the way an instrument does, because the thing being pushed on is the model of the mechanism.

The cost inversion ​

Compared to sampling, physical modeling trades storage for per-voice computation:

Sampled instrumentPhysical model
DataOften large because many notes, velocity layers, and articulations are recordedA compact parameter set
CPURead memory, resample, mixSolve the physics once per sample, per voice
Budget you run out ofDisk, download, memoryPolyphony
What a new note costsAnother playback voice and its reader stateIts own delay lines, filters and resonators

That inversion is why these voices can be the floor under a MIDI file that arrives with no SoundFont at all: there is nothing to download, because there is nothing to store. It is also why a dense arrangement of bowed strings is a heavier render than the same arrangement played from samples. You are not fetching audio; you are computing it.

Struck, and sustained ​

The exciters divide into two kinds, and the division decides what a player can do mid-note.

A struck or plucked exciter is finished almost immediately. A hammer contacts the string, transfers its energy, and leaves; everything afterwards is the loop ringing down on its own. For this kind of one-shot exciter, the main expressive decisions are made at the strike — how hard, how fast, where along the string — and a mid-note excitation route has no active target.

A sustained exciter keeps feeding the loop for as long as the note lasts. A bow stays on the string, a breath keeps arriving at the reed, a bellows keeps pushing air past the tongue. The exciter is inside the loop, reacting to the wave coming back at it, which is what makes these instruments feel alive: press harder and the tone changes because the coupling changed, not because a parameter was faded.

What that buys is that the player stays in the note. A crescendo on a bowed string is one continuous gesture, not a sequence of re-articulations, and a wind player can shape a phrase after the attack has gone. On a struck model there is nothing left to shape, so an expressive control routed there has nowhere to land — and libsonare says so explicitly rather than accepting the control and quietly doing nothing with it.

The honest limits ​

Two of them, and both are worth knowing before reaching for a model.

A model sounds like what was modelled. It is not a general-purpose realism setting. A bore model gives you clarinets and saxophones very well and gives you a choir not at all. The specific voice of a specific instrument — this cello, in this hall — is not in the mechanism, and getting close to it is calibration work rather than something the method provides.

Outside its fitted range, a model does not degrade gracefully. A sampler pushed past its recorded range sounds wrong in a familiar way: too bright, too slow, obviously stretched. A model pushed past the parameter range it was fitted for can stop being the instrument altogether — a bow force nothing grips at, a blowing pressure that overblows into a mode nobody wanted, a loss filter that stops losing and lets the loop run away. The failure is not a quality gradient; it is a different object.

The acoustic models ​

libsonare's built-in synthesizer builds the following acoustic-style engines this way. Each one is organised around one physical quantity — the thing a player actually controls, which the rest of the model is arranged to respond to:

ModelWhat it modelsThe quantity it is built around
pianoFelt hammer, coupled unison strings, soundboardWhere the hammer strikes
pipe-organFlue pipe drawing on a shared wind chestWind pressure at the mouth
bowed-stringRosined bow gripping and slipping on a stringContact point and bow force
reedCane reed valving a cylindrical or conical boreBlowing pressure and reed stiffness
brassLip valve on a flared boreBlowing pressure and embouchure
fluteAir jet across an edge, driving an open pipeJet transit time against the bore period
plucked-stringString grazing a curved bridgePluck point and bridge buzz
vocalGlottal source through five vowel formantsWhich vowel
free-reedMetal tongue swinging through a slotBellows pressure and tongue stiffness
harpsichordQuill-plucked string choirs with a short rear segmentPluck position and rear coupling

The parameters each one exposes, their defaults and ranges, and how settled each model is today all live on Physical Models. The piano model is tuned; every other physical model still awaits adjustment and calibration, with more work planned for future patch releases. Treat these as data-free preview/fallback voices rather than finished instrument simulations. Two further engines use physical ideas without being waveguides — modal strikes a bank of tuned resonators, and percussion models a circular membrane with a noise bed — and they are described alongside the rest on Built-in Synthesizer. additive is additive synthesis, not a physical model.

How libsonare implements this

The waveguide-style entries are piano, pipe-organ, bowed-string, reed, brass, flute, plucked-string and harpsichord. vocal is a source-filter model — a glottal source through a bank of vowel formant resonators, with no loop at all — and free-reed is a driven tongue swinging through a slot with no coupled air column. All of them are in the family because they are solved from a mechanism rather than drawn as a waveform.

Live expression reaches a model through four abstract excitation axes rather than through named per-voice fields: force (drive into the exciter — bow force, mouth pressure, bellows pressure), position (where the exciter meets the resonator, which only the bowed string has), brightness (timbral opening of the radiating end, held apart from loudness) and morph (registration morph, for an engine whose spectrum is drawn rather than excited). A modulation route names an axis, not an engine, so one dispatch line serves every voice and each engine reads only the axes it declares.

The accept set is declared rather than inferred: engine_axis_capability() in excitation_axes.h carries one row per engine mode, the switch behind it has no default: label so a new mode that declares nothing fails to compile, and an engine marked continuously excited with an empty axis mask fails a static_assert. The struck and plucked engines decline all four axes — karplus-strong, modal, percussion, piano, plucked-string and harpsichord, alongside the non-physical subtractive, fm and sample — because their exciters are one-shot and do not expose a mid-note excitation target. additive is the in-between case: drawbar tonewheels sustain but are not excited, so it takes the morph axis alone. vocal takes brightness only, force having no target that is not already the brightness tilt or the amplifier.

Related: Physical Models, Built-in Synthesizer (NativeSynth), Sound Sources, Synthesis Basics, SoundFont and Sampled Instruments