A clarinet made of physics: reimagining physical modelling synthesis for the AI era

- Could AI help recreate the sound of a clarinet from nothing but its physical measurements?
- My first attempt sounded like a nervous beginner, so I built a virtual clarinettist and had it learn through taking music exams.
- Hear it play Mozart, Gershwin and Poulenc straight from the sheet music, then play the clarinet yourself, live in your browser.
Rebirth of the cool
I was a child of the 1980s and 90s, which was the golden age for digital music and synthesisers. I used to pore over magazines like Sound on Sound and brochures from manufacturers such as Yamaha, Roland and Korg to read about the latest developments, witnessing the birth of a new era in music technology. Julian Colbeck’s Keyfax 3 was my bible, a compendium of all the synthesisers available at the time, with detailed descriptions, specifications and reviews that fuelled my obsession.
The Yamaha DX7 was one of the best-known synths of the time, renowned for its distinctive FM synthesis and iconic electric piano sounds, which made the sound of the 1980s instantly recognisable. FM synthesis itself was relatively basic by modern standards, involving the modulation of one waveform by another to create complex harmonic content. As the 1990s progressed, synthesisers became increasingly sophisticated and evolved beyond simple FM and similar techniques to using samples and richer waveforms to achieve more realistic sounds. Many synthesisers today still use samples (they were called ROMplers back then), which are essentially hundreds of recordings of a real instrument at many different pitches and velocities, replayed on demand by the instrument.

There was one technique that stood apart from these, though: Yamaha’s Virtual Acoustic synthesis, introduced with the VL1 in 1993, which used physical modelling to recreate the sound of real acoustic instruments from scratch. Unlike FM or sample-based synthesis, Virtual Acoustic synthesis attempted to simulate the actual physics of how an instrument produces sound, resulting in a more realistic and expressive performance. It was way ahead of its time, and the idea of actually simulating instruments using physics and maths in real time was truly awe-inspiring. The VL1 and its keyboardless companion, the VL1-m, were very expensive (the VL1 was £3,999 at launch, which in today’s money would be almost £9,000) and out of reach for many; as such, not many were made and they did not sell particularly well – despite the walnut control surface, reminiscent of the dashboard of a Jaguar saloon! They could also only play two notes at a time. Nevertheless, they stand out as a remarkable achievement in the history of synthesis. One can only imagine what it must have been like to be part of the highly talented team from Yamaha back in the day, breaking new ground and pushing the boundaries of what was technically possible in music technology.

Now, thirty years on from the launch of the VL1, I wondered if it was time to take a fresh look at this kind of technology. Though advanced for its time, the VL1 had to squeeze its physics into the DSP chips of the early 1990s, using compact waveguide models of the instrument. Research has advanced since then, and I wondered what could be achieved with modern computational power, more sophisticated physical models, and the integration of AI to fine-tune the parameters for a realistic and expressive performance. Could an authentic physical model of an instrument be created that responds to a player’s nuances as naturally as the real instrument, using maths and physics alone? Could the sound of an instrument be created entirely from actual measurements of the instrument itself?
Physical modelling never went away, of course: Audio Modeling’s SWAM instruments and Modartt’s Pianoteq have made it a practical tool for musicians, and university labs have pushed the physics much further. What I wanted to try was different: to build an instrument from its measurements and teach a virtual player to play it from the score.
Attempting Anatomic Synthesis
In my youth, I played clarinet, saxophone and piano in many ensembles and orchestras, so I decided to start with an instrument I knew, the clarinet. I found some careful measurements of a professional Buffet Crampon B♭ clarinet, published by Vincent Debut, Jean Kergomard and Franck Laloë in 2005. Those measurements trace the bore from the tip of the reed to the rim of the bell, and record the position, size and chimney height of every one of the 24 tone holes. I tweaked them a bit to resemble my own Buffet Crampon “Elite” clarinet, then drew on the long-established physics of sound in narrow tubes (the Zwikker–Kosten theory of how the walls absorb energy), formulas for how sound leaves the open end of the instrument and each tone hole (Silva and colleagues, Dalmont and Nederveen) and a model of the tone holes derived from computer simulations by Lefebvre and Scavone.
The clarinet is a single-reed instrument, the reed being a thin piece of cane which vibrates to make the sound. Similar research models the effect of the reed on the sound and how it interacts with the mouthpiece and tongue, and I used this to create a digital version of the mouthpiece.
Once a basic model was up and running, I was able to compare the model’s calculations with an independent open-source acoustics solver (Inria’s OpenWInD) and with laboratory measurements of real clarinets from the University of New South Wales.
The first version of the model, which I named “Anatomic Synthesis” (after the detailed anatomy of the instrument on which it was based), was ready to try. You can hear it below.
For a first attempt, it was actually quite convincing. As I continued to improve the model and it progressed to playing longer phrases, I realised something very surprising. The sounds I heard reminded me of a beginner playing the clarinet, struggling to control the reed and produce a steady tone. It became clear at this point that it was not enough simply to model the air column and the tone holes accurately; the interaction between the player’s breath, the reed and the mouthpiece also needed to be captured in detail. We needed a virtual musician!
Building a Virtual Musician
When you learn to play the clarinet, you learn to use the lips, tongue, mouth, throat, breath column and diaphragm to control the reed and shape the airflow, producing the desired pitch, tone and dynamics. The control of your lips and mouth (called the embouchure) takes many years of practice to master if you are to avoid the squeaks, toe-curling rasps and unintended noises that can so easily occur, some of which are like fingernails scraping down a chalkboard!
To train a virtual player, where better to start than music practice and exams? When you are learning an instrument like the clarinet, you learn to play scales, arpeggios, long sustained notes and pieces of varying difficulty. I created synthetic practice routines for the virtual musician to follow, allowing it to gradually improve its control over the reed, embouchure and breath, just like a human student would, and I scored each session. Over time, I taught the virtual player to learn more complex pieces, developing a nuanced understanding of how to produce a steady tone, articulate notes clearly, and express musicality through dynamics and phrasing.
When the virtual musician was connected up to the virtual instrument, the results were very good indeed. I was able to train the player through a mixture of my own recordings and synthetic practice routines, and watch it improve over time: to my ear, first to about Grade 2–3 standard, then progressing to Grade 6–7 standard.
Getting to an approximate Grade 8 standard (the highest of the graded exams) needed quite a bit more work. Whilst the physical model of the sound was largely complete by this stage, I had to adapt the virtual musician’s learning algorithms and practice routines to play with a level of musicality appropriate for that grade.
The virtual performance
To conclude the experiment, I made the full end-to-end system: a MusicXML file as input (MusicXML is a way of encoding regular sheet music as machine-readable data), a virtual player that could read and interpret it and convert it into mouth, breath and lip signals, and finally a virtual clarinet which could convert these into sound.
I picked several classic pieces of repertoire, familiar to all budding clarinet students, to showcase what anatomic synthesis can do: Mozart’s Clarinet Concerto, the opening of Gershwin’s “Rhapsody in Blue”, and the three movements of Poulenc’s Clarinet Sonata. Many advanced students would find the opening of “Rhapsody in Blue” challenging, given its “glissando”, a musical “smear” that requires precise control of the lips and throat and which is really hard to explain or teach! The virtual musician has learned the lip, throat and breath control to handle it with aplomb.
You can have a listen below and see the simulation in action. You can change the piece being played using the dropdown at the top.
Trying it out yourself
What purpose does this serve ultimately? I think this experiment demonstrates that it is now possible to create realistic and expressive virtual instruments, reconstructed from their physical form. Whilst I used a clarinet for this example, how about CT scanning an extremely valuable Stradivarius violin and enabling musicians worldwide to practise and perform on it virtually, without risking damage to the priceless instrument?
Today, this clarinet model can be played as a virtual instrument over MIDI, using a standard keyboard or breath controller. In a way, though, the model is now more advanced than the hardware available to drive it. MIDI wind controllers such as the Yamaha WX7/11/5 and their modern-day equivalents, like the Akai EWI and Roland Aerophone, still measure only breath, bite and, on some, lip pressure or motion, while the model can also respond to the tongue, the throat and how firmly the lips hold the reed. A device with a greater array of sensors would be needed to fully exploit the capabilities of the virtual instrument.
In the meantime, though, you can try it out in your browser using the controller below. Again, this isn’t playing recordings or samples; it is actually simulating the physical behaviour of the clarinet in real time to create the sound waves you hear.
I think this kind of approach not only opens up new possibilities for replicating existing instruments, but also for creating entirely new ones that were previously unimaginable: a Frankenstein instrument made up of half oboe and half French horn, for example.
The VL1 was right about what a synthesiser could be, and others have carried that idea forward since. What’s new is that measurements, simulation and AI can now do much of the heavy lifting: we can build instruments from the inside out and bring the nuance, expression and subtlety that real instruments have and sample-based synthesis often lack. This, in turn, enables greater expression, new musical palettes, and greater performance freedom for musicians everywhere.
If you’re interested in this experiment, discussing this technology further, or would just like to geek out on the golden age of digital synthesis, you can reach me at jim@serendipity.ai.
Text written by hand, reviewed by Claude
References
The instrument
- V. Debut, J. Kergomard, F. Laloë, “Analysis and optimisation of the tuning of the twelfths for a clarinet resonator”, Applied Acoustics 66(4), 365–409 (2005). doi:10.1016/j.apacoust.2004.08.003, open preprint arXiv:physics/0309051. The clarinet’s measurements.
- E. Moers, J. Kergomard, “On the cutoff frequency of clarinet-like instruments: geometrical versus acoustical regularity”, Acta Acustica united with Acustica 97(6), 984–996 (2011). doi:10.3813/aaa.918480. Fingerings.
- C. Zwikker, C. W. Kosten, Sound Absorbing Materials, Elsevier (1949). Losses at the walls of the tube.
- F. Silva, P. Guillemain, J. Kergomard, B. Mallaroni, A. Norris, “Approximation formulae for the acoustic radiation impedance of a cylindrical pipe”, Journal of Sound and Vibration 322(1–2), 255–263 (2009). doi:10.1016/j.jsv.2008.11.008.
- J.-P. Dalmont, C. J. Nederveen, N. Joly, “Radiation impedance of tubes with different flanges: numerical and experimental investigations”, Journal of Sound and Vibration 244(3), 505–534 (2001). doi:10.1006/jsvi.2000.3487.
- A. Lefebvre, G. P. Scavone, “Characterization of woodwind instrument toneholes with the finite element method”, Journal of the Acoustical Society of America 131(4), 3153–3163 (2012). doi:10.1121/1.3685481.
The reed and the player
- P. Guillemain, J. Kergomard, T. Voinier, “Real-time synthesis of clarinet-like instruments using digital impedance models”, Journal of the Acoustical Society of America 118(1), 483–494 (2005). doi:10.1121/1.1937507.
- V. Chatziioannou, M. van Walstijn, “Estimation of clarinet reed parameters by inverse modelling”, Acta Acustica united with Acustica 98(4), 629–639 (2012). doi:10.3813/aaa.918543.
- J.-M. Chen, J. Smith, J. Wolfe, “Pitch bending and glissandi on the clarinet: roles of the vocal tract and partial tone hole closure”, Journal of the Acoustical Society of America 126(3), 1511–1520 (2009). doi:10.1121/1.3177269. How players bend notes and play the Gershwin glissando.
Checking against the real thing
- OpenWInD, open-source software for wind instrument acoustics, Inria.
- Clarinet acoustics, Music Acoustics, UNSW Sydney: measured impedance for every fingering (P. Dickens, with J. Cavanagh, R. France, J. Tann and J. Wolfe).
- University of Iowa Musical Instrument Samples, recordings by Lawrence Fritts. Reference notes for fitting the instrument.
Background
- A. Chaigne, J. Kergomard, Acoustics of Musical Instruments, Springer (2016). doi:10.1007/978-1-4939-3679-3.
- J. O. Smith III, Physical Audio Signal Processing, CCRMA, Stanford University (online book). The waveguide approach behind the VL1.
- Yamaha VL1 review, Sound on Sound.
- Yamaha Corporation, Evolution of Tone Generator Systems and Approaches to Music Production, Synth 50th Anniversary history, chapter 3. The Virtual Acoustic tone generator and the VL1 (1993).
- Audio Modeling, SWAM Clarinets and the SWAM (Synchronous Waves Acoustic Modelling) instruments: real-time woodwinds, brass and strings combining physical and behavioural modelling.
- Modartt, Pianoteq: a physically modelled piano.
Acknowledgements
- Head scan: “Infinite, 3D Head Scan” by Lee Perry-Smith (Infinite-Realities), licensed under Creative Commons Attribution 3.0 Unported; based on a work at www.triplegangers.com. Used for the virtual clarinettist in the player widgets.
- Score engraving: OpenSheetMusicDisplay. 3D rendering: three.js.