← writing

A clarinet made of physics: reimagining physical modelling synthesis for the AI era

Jim Marshall·synthesisphysical modellingmusicAI
The anatomic synthesis player mid-glissando in Gershwin's Rhapsody in Blue: a porcelain bust playing a clarinet, with live graphs of breath, reed and pitch, and the score scrolling beneath
A virtual clarinettist playing a clarinet using anatomic synthesis, an approach to physical modelling that uses AI. Hear it below. Head scan: Lee Perry-Smith / Infinite-Realities, CC-BY 3.0.
  • Could AI help recreate the sound of a clarinet from nothing but its physical measurements?
  • My first attempt sounded like a nervous beginner, so I built a virtual clarinettist and had it learn through taking music exams.
  • Hear it play Mozart, Gershwin and Poulenc straight from the sheet music, then play the clarinet yourself, live in your browser.

Rebirth of the cool

I was a child of the 1980s and 90s, which was the golden age for digital music and synthesisers. I used to pore over magazines like Sound on Sound and brochures from manufacturers such as Yamaha, Roland and Korg to read about the latest developments, witnessing the birth of a new era in music technology. Julian Colbeck’s Keyfax 3 was my bible, a compendium of all the synthesisers available at the time, with detailed descriptions, specifications and reviews that fuelled my obsession.

The Yamaha DX7 was one of the best-known synths of the time, renowned for its distinctive FM synthesis and iconic electric piano sounds, which made the sound of the 1980s instantly recognisable. FM synthesis itself was relatively basic by modern standards, involving the modulation of one waveform by another to create complex harmonic content. As the 1990s progressed, synthesisers became increasingly sophisticated and evolved beyond simple FM and similar techniques to using samples and richer waveforms to achieve more realistic sounds. Many synthesisers today still use samples (they were called ROMplers back then), which are essentially hundreds of recordings of a real instrument at many different pitches and velocities, replayed on demand by the instrument.

A 1985 Yamaha advertisement showing the DX7 and the KX5 remote keyboard, headlined 'The performance has begun. But it's just the beginning.'
”The performance has begun. But it’s just the beginning.” A 1985 Yamaha advertisement for the DX7 and the KX5 remote keyboard. Scan: retrosynthads.blogspot.com.

There was one technique that stood apart from these, though: Yamaha’s Virtual Acoustic synthesis, introduced with the VL1 in 1993, which used physical modelling to recreate the sound of real acoustic instruments from scratch. Unlike FM or sample-based synthesis, Virtual Acoustic synthesis attempted to simulate the actual physics of how an instrument produces sound, resulting in a more realistic and expressive performance. It was way ahead of its time, and the idea of actually simulating instruments using physics and maths in real time was truly awe-inspiring. The VL1 and its keyboardless companion, the VL1-m, were very expensive (the VL1 was £3,999 at launch, which in today’s money would be almost £9,000) and out of reach for many; as such, not many were made and they did not sell particularly well – despite the walnut control surface, reminiscent of the dashboard of a Jaguar saloon! They could also only play two notes at a time. Nevertheless, they stand out as a remarkable achievement in the history of synthesis. One can only imagine what it must have been like to be part of the highly talented team from Yamaha back in the day, breaking new ground and pushing the boundaries of what was technically possible in music technology.

Yamaha brochures: the VL1 Virtual Acoustic Synthesizer on a burr-walnut cover, and the VL Series Version 2 with the VL1-m tone generator and two VL1 keyboards
Yamaha’s original brochures for the VL1 Virtual Acoustic Synthesizer (left) and the VL Series Version 2, with the VL1-m tone generator (right).
If you are interested in learning more about the VL1, Wolf Design Review has some beautiful photography and documentation of the instrument on its website. Or alternatively, learn more about the big sister of the VL1, the Yamaha VP1 (which cost as much as a small car!).

Now, thirty years on from the launch of the VL1, I wondered if it was time to take a fresh look at this kind of technology. Though advanced for its time, the VL1 had to squeeze its physics into the DSP chips of the early 1990s, using compact waveguide models of the instrument. Research has advanced since then, and I wondered what could be achieved with modern computational power, more sophisticated physical models, and the integration of AI to fine-tune the parameters for a realistic and expressive performance. Could an authentic physical model of an instrument be created that responds to a player’s nuances as naturally as the real instrument, using maths and physics alone? Could the sound of an instrument be created entirely from actual measurements of the instrument itself?

Physical modelling never went away, of course: Audio Modeling’s SWAM instruments and Modartt’s Pianoteq have made it a practical tool for musicians, and university labs have pushed the physics much further. What I wanted to try was different: to build an instrument from its measurements and teach a virtual player to play it from the score.

Attempting Anatomic Synthesis

In my youth, I played clarinet, saxophone and piano in many ensembles and orchestras, so I decided to start with an instrument I knew, the clarinet. I found some careful measurements of a professional Buffet Crampon B♭ clarinet, published by Vincent Debut, Jean Kergomard and Franck Laloë in 2005. Those measurements trace the bore from the tip of the reed to the rim of the bell, and record the position, size and chimney height of every one of the 24 tone holes. I tweaked them a bit to resemble my own Buffet Crampon “Elite” clarinet, then drew on the long-established physics of sound in narrow tubes (the Zwikker–Kosten theory of how the walls absorb energy), formulas for how sound leaves the open end of the instrument and each tone hole (Silva and colleagues, Dalmont and Nederveen) and a model of the tone holes derived from computer simulations by Lefebvre and Scavone.

The clarinet is a single-reed instrument, the reed being a thin piece of cane which vibrates to make the sound. Similar research models the effect of the reed on the sound and how it interacts with the mouthpiece and tongue, and I used this to create a digital version of the mouthpiece.

Once a basic model was up and running, I was able to compare the model’s calculations with an independent open-source acoustics solver (Inria’s OpenWInD) and with laboratory measurements of real clarinets from the University of New South Wales.

The first version of the model, which I named “Anatomic Synthesis” (after the detailed anatomy of the instrument on which it was based), was ready to try. You can hear it below.

The first version of the model: a single note, middle C, played mf.

For a first attempt, it was actually quite convincing. As I continued to improve the model and it progressed to playing longer phrases, I realised something very surprising. The sounds I heard reminded me of a beginner playing the clarinet, struggling to control the reed and produce a steady tone. It became clear at this point that it was not enough simply to model the air column and the tone holes accurately; the interaction between the player’s breath, the reed and the mouthpiece also needed to be captured in detail. We needed a virtual musician!

An early attempt at the opening of Mozart’s Clarinet Concerto, II. Adagio, before the virtual musician: the instrument is kind of there, but being played by a beginner. If you have children who are learning the clarinet, you will recognise this!

Building a Virtual Musician

When you learn to play the clarinet, you learn to use the lips, tongue, mouth, throat, breath column and diaphragm to control the reed and shape the airflow, producing the desired pitch, tone and dynamics. The control of your lips and mouth (called the embouchure) takes many years of practice to master if you are to avoid the squeaks, toe-curling rasps and unintended noises that can so easily occur, some of which are like fingernails scraping down a chalkboard!

To train a virtual player, where better to start than music practice and exams? When you are learning an instrument like the clarinet, you learn to play scales, arpeggios, long sustained notes and pieces of varying difficulty. I created synthetic practice routines for the virtual musician to follow, allowing it to gradually improve its control over the reed, embouchure and breath, just like a human student would, and I scored each session. Over time, I taught the virtual player to learn more complex pieces, developing a nuanced understanding of how to produce a steady tone, articulate notes clearly, and express musicality through dynamics and phrasing.

When the virtual musician was connected up to the virtual instrument, the results were very good indeed. I was able to train the player through a mixture of my own recordings and synthetic practice routines, and watch it improve over time: to my ear, first to about Grade 2–3 standard, then progressing to Grade 6–7 standard.

The same opening of Mozart’s Adagio, with the virtual musician at around Grade 6–7.

Getting to an approximate Grade 8 standard (the highest of the graded exams) needed quite a bit more work. Whilst the physical model of the sound was largely complete by this stage, I had to adapt the virtual musician’s learning algorithms and practice routines to play with a level of musicality appropriate for that grade.

The virtual performance

To conclude the experiment, I made the full end-to-end system: a MusicXML file as input (MusicXML is a way of encoding regular sheet music as machine-readable data), a virtual player that could read and interpret it and convert it into mouth, breath and lip signals, and finally a virtual clarinet which could convert these into sound.

I picked several classic pieces of repertoire, familiar to all budding clarinet students, to showcase what anatomic synthesis can do: Mozart’s Clarinet Concerto, the opening of Gershwin’s “Rhapsody in Blue”, and the three movements of Poulenc’s Clarinet Sonata. Many advanced students would find the opening of “Rhapsody in Blue” challenging, given its “glissando”, a musical “smear” that requires precise control of the lips and throat and which is really hard to explain or teach! The virtual musician has learned the lip, throat and breath control to handle it with aplomb.

You can have a listen below and see the simulation in action. You can change the piece being played using the dropdown at the top.

Mozart’s Clarinet Concerto (timing and dynamics analysed from tens of professional recordings), the opening of Gershwin’s Rhapsody in Blue with its lipped glissando, and the three movements of Poulenc’s Clarinet Sonata. The score is shown as the clarinettist reads it.

Trying it out yourself

What purpose does this serve ultimately? I think this experiment demonstrates that it is now possible to create realistic and expressive virtual instruments, reconstructed from their physical form. Whilst I used a clarinet for this example, how about CT scanning an extremely valuable Stradivarius violin and enabling musicians worldwide to practise and perform on it virtually, without risking damage to the priceless instrument?

Today, this clarinet model can be played as a virtual instrument over MIDI, using a standard keyboard or breath controller. In a way, though, the model is now more advanced than the hardware available to drive it. MIDI wind controllers such as the Yamaha WX7/11/5 and their modern-day equivalents, like the Akai EWI and Roland Aerophone, still measure only breath, bite and, on some, lip pressure or motion, while the model can also respond to the tongue, the throat and how firmly the lips hold the reed. A device with a greater array of sensors would be needed to fully exploit the capabilities of the virtual instrument.

In the meantime, though, you can try it out in your browser using the controller below. Again, this isn’t playing recordings or samples; it is actually simulating the physical behaviour of the clarinet in real time to create the sound waves you hear.

The model runs on your device in real time. Play with the on-screen or computer keyboard, change how the notes join (slurred, tongued, staccato), smear between them, or try the scale and arpeggio buttons.

I think this kind of approach not only opens up new possibilities for replicating existing instruments, but also for creating entirely new ones that were previously unimaginable: a Frankenstein instrument made up of half oboe and half French horn, for example.

The VL1 was right about what a synthesiser could be, and others have carried that idea forward since. What’s new is that measurements, simulation and AI can now do much of the heavy lifting: we can build instruments from the inside out and bring the nuance, expression and subtlety that real instruments have and sample-based synthesis often lack. This, in turn, enables greater expression, new musical palettes, and greater performance freedom for musicians everywhere.

If you’re interested in this experiment, discussing this technology further, or would just like to geek out on the golden age of digital synthesis, you can reach me at jim@serendipity.ai.

Text written by hand, reviewed by Claude

References

The instrument

  • V. Debut, J. Kergomard, F. Laloë, “Analysis and optimisation of the tuning of the twelfths for a clarinet resonator”, Applied Acoustics 66(4), 365–409 (2005). doi:10.1016/j.apacoust.2004.08.003, open preprint arXiv:physics/0309051. The clarinet’s measurements.
  • E. Moers, J. Kergomard, “On the cutoff frequency of clarinet-like instruments: geometrical versus acoustical regularity”, Acta Acustica united with Acustica 97(6), 984–996 (2011). doi:10.3813/aaa.918480. Fingerings.
  • C. Zwikker, C. W. Kosten, Sound Absorbing Materials, Elsevier (1949). Losses at the walls of the tube.
  • F. Silva, P. Guillemain, J. Kergomard, B. Mallaroni, A. Norris, “Approximation formulae for the acoustic radiation impedance of a cylindrical pipe”, Journal of Sound and Vibration 322(1–2), 255–263 (2009). doi:10.1016/j.jsv.2008.11.008.
  • J.-P. Dalmont, C. J. Nederveen, N. Joly, “Radiation impedance of tubes with different flanges: numerical and experimental investigations”, Journal of Sound and Vibration 244(3), 505–534 (2001). doi:10.1006/jsvi.2000.3487.
  • A. Lefebvre, G. P. Scavone, “Characterization of woodwind instrument toneholes with the finite element method”, Journal of the Acoustical Society of America 131(4), 3153–3163 (2012). doi:10.1121/1.3685481.

The reed and the player

  • P. Guillemain, J. Kergomard, T. Voinier, “Real-time synthesis of clarinet-like instruments using digital impedance models”, Journal of the Acoustical Society of America 118(1), 483–494 (2005). doi:10.1121/1.1937507.
  • V. Chatziioannou, M. van Walstijn, “Estimation of clarinet reed parameters by inverse modelling”, Acta Acustica united with Acustica 98(4), 629–639 (2012). doi:10.3813/aaa.918543.
  • J.-M. Chen, J. Smith, J. Wolfe, “Pitch bending and glissandi on the clarinet: roles of the vocal tract and partial tone hole closure”, Journal of the Acoustical Society of America 126(3), 1511–1520 (2009). doi:10.1121/1.3177269. How players bend notes and play the Gershwin glissando.

Checking against the real thing

Background

Acknowledgements