← Back

Samuel: A silly speech model

I trained a model that parrots speech using Pink Trombone, a silly vocal tract simulator.

Vocal tract diagram

Try it here, or read the code here.

Update 12 September 2026: What follows is the text of a paper I wrote about Samuel for ICLC 2027. I kept the editorial "we" I (we?) used in the paper.

Introduction

Live coding software is typically loop-based: the state of the code at a given time defines music that can keep playing forever. Movement in the music is then achieved by the performer modifying the code during the performance. This approach is well-suited for creating instrumental electronic music, but it is unclear how to integrate vocals into the performance in a tighter manner than simply singing over the live-coded instrumental.

One solution is speech synthesis: using a text-to-speech model, vocals can be described by code. An example of such a system is the samda command in nudel.cc (currently defunct) by pastagang. Calling samda('hello_world') synthesizes the text "hello world" using a JavaScript reimplementation of SAM (Software Automatic Mouth), a 1982 text-to-speech by Mark Barton. The synthesized audio can then be used in Strudel as a sample.

Figure
Figure 1:

Samuel's reproduction of selected phonemes. Notice that /ɑ/ and /i/ match how a human would produce the sound. Other sounds do not have realistic mouth positions: Samuel manages to pronounce /ɔ/ without rounding the lips, and to pronounce /n/ without letting air through the nasal cavity. Samuel is not explicitly trained to be physically realistic, only to reproduce sounds. Therefore, even partial realism is not an obvious property.

To achieve a deeper dialogue between vocals and live coding, we introduce Samuel, a speech synthesis model conceived as a modern version of SAM. Samuel applies retro speech synthesis methods for their distinct timbre, but anachronistically improves upon them using modern machine learning for superior naturalness. In particular, Samuel is a model that receives a speech audio sample as input and resynthesizes it using Pink Trombone, a simplified vocal tract simulator. It is a learning-based approach, trained on public-domain data from LibriVox.

For live coding, Samuel's output is easy to modify: since it predicts the parameter movements for Pink Trombone, we can modify any of these for creative purposes. The most straightforward modification of this sort is changing the pitch contour to make the voice sing a given melody. Other, more experimental modifications are also possible, such as making everything unvoiced, or overriding the position of the lips. In addition, Samuel is a causal model, meaning that it can be used for real-time vocal processing. We provide details of potential live coding applications in the Applications section.

Note that despite being named after SAM, Samuel is a speech-to-speech rather than a text-to-speech model. However, a text-to-speech model could be easily created by first running a small TTS (e.g. Pocket TTS), and applying Samuel to the synthesized speech.

We invite the reader to try Samuel interactively using the webapp at samuel.vvolhejn.com, watch the demo video showcasing an example of a live coding use at https://youtu.be/q6DIBpIn6RY, or read the code at github.com/vvolhejn/samuel.

Pink Trombone

Pink Trombone is a simplified physical model of a mouth, and serves as a playful educational tool. It is controllable through a webapp (see Figure 2) where users can move the tongue and lips, as well as control the parameters of the throat such as pitch and whether the sound should be voiced or unvoiced.

Figure
Figure 2:

The original Pink Trombone interface by Neil Thapen.

Pink Trombone can reproduce a wide variety of speech sounds despite being a heavily simplified model. For instance, moving the tongue changes the vowel, closing and opening the lips makes a /b/ sound, and touching the soft palate with the tongue while letting the air through the nasal cavity produces a /ŋ/ sound (sing).

Previous work on speech synthesis using Pink Trombone exists, but it either relies on hand-crafted parameter paths for phonemes, or, when using learning-based approaches, does not reach intelligibility beyond individual words.

Speech synthesis model

Figure
Figure 3:

Samuel predicts the trajectories of seven parameters to reproduce speech using Pink Trombone. The depicted trajectories were predicted by Samuel for the words "bread and some apples". The eighth parameter, frequency, is estimated using the pYIN pitch detector.

Pink Trombone is a source–filter model: it represents speech as a sound produced by a source (the vocal cords), which is then filtered by the vocal tract.

Three parameters control the source, i.e. the vocal cords: the F0 (pitch), voiceness (continuous value where 0 is unvoiced, 1 is voiced) and intensity (gain). Voiced sounds are generated using a Liljencrants-Fant model, a mathematical model describing the airflow through the vibrating vocal cords. The unvoiced portion is generated using filtered white noise, and is interpolated with the voiced sound based on the voiceness parameter.

The filter, i.e. the vocal tract, is a "digital waveguide" model: it simulates a physical tube through which the sound travels, filtering it. Samuel controls the shape of the vocal tract using five parameters: the tongue base position and height, the tongue tip position and height, and how open the lips are.

The vocal tract is visualized as continuous (see Figure 2), but internally, it is modeled as a tube composed of 44 discrete segments. The position of the tongue, tongue tip, and lips determines the radii of these segments.

In total, there are therefore eight parameters to control Pink Trombone. Seven of these are predicted by Samuel, and the last one, frequency, is provided externally through a pitch detector. See Figure 3 for an example of predicted parameter trajectories.

Sound passes through the vocal tract sample-by-sample, with one wave going rightward (out of the mouth) and another going leftward (back into the throat). At every step, we update these two waves; in each segment, part of the wave continues in its direction and part is reflected back; see Figure 4 for an illustration. This is known as a Kelly–Lochbaum vocal tract model.

Figure
Figure 4:

Illustration of a single step of the Kelly–Lochbaum vocal tract model used by Pink Trombone. A waveform comes from the throat on the left, travels through the vocal tract, and exits through the mouth on the right. Two waves are traveling in the tract; one rightwards (red), one leftwards (blue). At each step, the waves move by one segment, and are partly reflected. For instance, the right-going value in segment 38 travels to segment 39 in the next step, but is also partially reflected into the left-going wave in segment 38.

Figure 5:

An impulse traveling through the vocal tract over many steps of the model from Figure 4.

Pink Trombone also has a junction leading to the nasal cavity for nasal sounds like /n/ in cone, but it remains unused in our experiments (see the Phonetic plausibility section).

The Kelly–Lochbaum vocal tract model played an important role in the history of speech synthesis applied to singing: in 1962, Kelly and Lochbaum used the model to create what is likely the first singing computer voice. The recording is available at https://ccrma.stanford.edu/~jos/wav/daisy-klm.wav. Arthur C. Clarke later heard the demo when visiting Bell Labs and used it in 2001: A Space Odyssey. HAL9000 sings the same song, Daisy Bell, in the scene where it gets disconnected (albeit in the film, the HAL9000 voice is not done using speech synthesis).

Method

The model takes a speech recording as an input and outputs parameters to emulate the original speech through Pink Trombone. Two aspects of this are challenging: defining a metric that formalizes what "as similar as possible" means, and actually finding the Pink Trombone parameters that optimize that metric.

Choosing a distance metric

To provide training signal to the model, we need a notion of distance between sounds. In our initial experiments, we compare distance metrics based on three different representations of the audio: spectral, mel-spectral, and MFCC. The spectral loss computes the L1 distance between log-magnitude spectrograms. The mel-spectral loss additionally applies an 80-bin mel filterbank to the spectrogram to reweight it so that distances better correspond to human perception. Refer to Radkoff (2021) for an in-depth explanation of these losses. MFCCs (mel-frequency cepstrum coefficients) apply a discrete cosine transform to the log-mel spectrogram, obtaining a "spectrum of a spectrum". Qualitatively, the model trained using MFCC loss yields the best results, so this loss is used for the rest of the experiments.

Training using MFCC loss results in speech-like sounds that roughly follow the original. However, in most cases it is not possible to understand what the model is saying. To improve intelligibility, we introduce a perceptual loss based on wav2vec 2.0, a self-supervised speech model that computes features whose distance should match human perception. We again refer the reader to Radkoff (2021) for an introduction to perceptual losses.

In our experiments, we find that adding the perceptual loss makes the model significantly more intelligible, with a quantitatively measurable gain. Using a speech-to-text model (Whisper, via faster-whisper), the MFCC-based Samuel models are essentially never transcribed correctly, reaching a word error rate of 1. Early experiments with the perceptual loss reduce the word error rate to around 0.3, with the final model reaching a WER of 0.20.

The fact that WER reaches non-trivial values is important, since it allows for quantitative comparison of models and therefore rapid experimentation. Before non-trivial WER is reached, comparing two models must be done qualitatively, by listening to multiple samples and deciding manually.

Efficient training

Pink Trombone's Kelly–Lochbaum vocal tract model requires sample-by-sample computation. A direct implementation in Python would make training inefficient: since the parameter choice at step tt affects the output for all steps t′>tt' > t, backpropagation would have to traverse the whole sequence.

To accelerate training, we instead use an approximation using finite impulse response (FIR) filters. When the Pink Trombone mouth does not move, it acts as a linear and time-invariant filter with an infinite impulse response. We can approximate this filter by measuring its impulse response over a finite horizon; we find that 256 samples suffice.

During training, we proceed as follows: we divide time into 512-sample chunks and have the model predict parameters for each chunk. At our sample rate of 44.1 kHz, this means the model runs at 86 Hz. In parallel, we compute the impulse response for each chunk: this is a sample-by-sample computation, but can be batched across the chunks and only requires 256 sequential steps (the impulse response length), no matter the length of the original audio.

The model is then used to control the original JavaScript Pink Trombone, whose behavior slightly different from the training-time FIR approximation. In practice, we find that the approximation is faithful enough to transfer well.

To make the model's task easier, the fundamental frequency (F0) contour of the audio is estimated separately using pYIN rather than having the model estimate it. pYIN estimates the F0 contour as well as a binary signal predicting whether the speech is voiced or unvoiced at a given time.

Regularization

Using the two losses described so far, Samuel does reach intelligibility, but no constraints are put on physical plausibility. For instance, nothing prevents the model from having lips fully closed in one frame and fully open in the next, 11.6ms later. This has no effect on the audio (in fact, the model can reproduce audio better when unconstrained) but manifests itself as jitteriness in the visualization. To reduce this jitter, we introduce three regularization losses that encourage Samuel to respect the physical limits on the speed of movement of a mouth.

Smoothness loss: An L1 loss on the first difference of the parameter trajectories. This penalizes the model for changing the parameters, encouraging it to move as little as necessary.

Acceleration loss: An L1 loss on the second difference, penalizing changes in velocities. We introduce this loss to combat the fact that the smoothness loss does penalize large jumps, but not small jitter around a fixed value.

Rest loss: In addition, we define a set of parameters as a rest pose: tongue in the back, tongue tip lowered, and lips nearly closed, and add a small rest loss that penalizes L1 distance from the rest pose. The weight of the loss is small enough to only matter if there is nothing else to do, i.e. when the original audio is nearly silent.

The weights of these regularization losses are tuned to the highest values possible that still have a minimal effect (<1%) on the word error rate. We also apply per-parameter weights to all three losses: higher weight for visually significant parameters such as the tongue index, and lower for voiceness and intensity.

Training details

Architecture. The Samuel encoder is a 1D convolutional neural network based on Pocket TTS, which in turn is based on SEANet. We refer the reader to Figure 2 of the SEANet paper for an overview of the architecture. We use downsampling ratios of [4, 4, 4, 8], with channels going from 32 in multiples of two to 512 at the end, with a final conv layer projecting to 128 channels.

Direct prediction head. We evaluate two different architectures for the prediction head. The simpler approach is direct prediction: on top of the convolutional neural network, add a linear layer followed by a tanh activation to predict the normalized values of the seven parameters. While simple, this approach is not very expressive: the model cannot express predictions such as "with 80% probability, the mouth is nearly closed, and with 20% probability, it is fully open". With this approach, we achieve a word error rate of 0.24.

Gumbel Softmax prediction head. The other prediction approach increases expressivity by letting the model predict a categorical distribution (similar to e.g. Figure 6 of PixelRNN). For each frame and each of the seven parameters, the model predicts a distribution over 32 buckets spanning the parameter's range. Then, we sample from the distribution to obtain a single set of parameters that is passed to Pink Trombone for synthesis. Since sampling from a distribution is a non-differentiable operation, we use Gumbel Softmax; see this post for an informal introduction. Unlike the original Gumbel Softmax, we keep the sampling temperature at 2 and replace temperature annealing by an entropy floor loss that prevents the model's prediction collapsing to predicting a single value for a given parameter. At inference time, we deterministically select the parameters with the highest confidence. This is the approach used in the published model, which reaches a word error rate of 0.20.

Model size. The model consists of 3.4 million parameters, making it tiny by modern machine learning standards, and inexpensive to run on CPU. In comparison, Pocket TTS, considered a very small text-to-speech model, uses 100 million parameters.

Data. Samuel is trained on a 1000-hour subset of Libri-Light. Libri-Light, in turn, is a curated collection of recordings from LibriVox, a public-domain database of audiobooks.

Compute. We train the final model in 18 hours on a single H100 GPU. Across all experiments, we use around 900 GPU-hours, all on spare capacity of the Kyutai cluster that would otherwise remain unused.

Applications

Using the methods described above, we successfully train a model that efficiently reproduces human speech using Pink Trombone. The code and model are available at github.com/vvolhejn/samuel under a GPL-3.0 license. A webapp demonstrating Samuel can be found at samuel.vvolhejn.com. The user can either record themselves speaking and let Samuel imitate their voice, or use one of the pre-recorded examples.

Samuel can be thought of as a vocal effect, a synthesizer, or a vocoder, although it does not neatly fit into any of the categories. This unique position offers several interesting live coding applications revolving around dialogue between live coding and the voice. The live coding community has already shown some interest in Pink Trombone: a SuperCollider port exists, as well as an efficient JavaScript version optimized for live performance. Building blocks that would allow integrating Samuel into live coding have thus already been developed.

Here are ideas for how to use Samuel in a live coding context. We will use Strudel throughout as an example, although the proposed ideas are not Strudel-specific.

Looper: allow the user to record a sample, pass it through Samuel, and repeat it, similar to a looper pedal. Synchronize the playback with Strudel. Since Samuel's output is a list of trajectories for the various parameters, we can then override some of them using Strudel patterns. The obvious choice is overriding the pitch/melody of the sample in an autotune-like fashion, letting the voice sample play a different melody. Overriding other parameters such as the tongue position and the voiceness is also possible and creates novel kinds of distortion. A demo video of a proof-of-concept of this idea is available at youtu.be/q6DIBpIn6RY.

The looper setup could be particularly interesting for online jam platforms such as Flok, or nudel.cc (currently defunct) by pastagang. Jamming online using traditional instruments and the voice is difficult because of latency. Online live coding jams are possible as long as participants share the same code, since small delays between participants are not an issue. Recording a vocal sample into a Samuel looper is not latency-sensitive and therefore compatible with online jamming.

Processing samples: Samuel can also be used to process pre-recorded vocal samples, an application technically similar to the looper but different in usage. Currently, Strudel's sample-manipulation capabilities are geared more towards drum loops than pitched samples. Re-tuning to a specific pitch requires using a method like .speed() to multiply the playback speed and therefore frequency. Playing a melody requires knowledge of the original pitch and complex conversions between semitones and frequency ratios. Retuning a Samuel-processed sample is trivial in contrast.

Real-time vocal processing: Samuel is a causal model, meaning it does not need to see the audio in advance to process it. It processes audio with a delay of 11.6ms, low enough for real-time use cases. A singer's voice could be routed through Samuel and partially processed via Strudel (e.g. repitching), using the methods mentioned above.

Usage as a vocoder: Samuel can be thought of as a vocoder, with the throat acting as the carrier and the mouth as the modulator. One could replace the throat with an arbitrary audio source (chords, bassline, synth lead...) and use the mouth as a filter.

Using Samuel as the input: One could also invert the control direction and use Samuel as a signal source. For example, one could take the position of the tongue and map it to the cutoff frequency of a low-pass filter.

Phonetic plausibility

Since Samuel learns from data, it is never taught how humans pronounce different phonemes. Any similarity between human mouth positions and those predicted by Samuel for Pink Trombone is not explicitly trained for. In Figure 1 we see that we achieve partial phonetic plausibility: /ɑ/ and /i/ match how a human would produce the sound, whereas /ɔ/ is pronounced unrealistically, with unrounded lips.

Humans pronounce plosives such as /p/ or /t/ by momentarily closing the air flow out of their mouth before releasing it again. Samuel instead emulates the sound without closing the vocal tract. This is not by design, since the model could choose to close the lips or press the tongue against the roof of the mouth; it simply manages to find another way to pronounce the same phoneme.

The same is true of nasal sounds like /n/ in cone. Nasal sounds are pronounced by letting air flow through the nasal cavity rather than the mouth. As seen in Figure 1, Samuel instead pronounces /n/ without using the nasal cavity, instead narrowing the mouth significantly. Samuel does have the option to open the velum and use the nasal cavity, but the model converges to not doing so.

These effects are likely explainable by the non-convexity of the optimization problem, and the loss used to train the model. Increasing phonetic accuracy would be an interesting direction for future work.

Conclusion

We introduced Samuel, a model trained to reproduce speech using Pink Trombone, a simplified speech synthesis model. The visual and playful nature of the model has attracted the attention of the live coding community in the past. Samuel enables new speech synthesis applications using Pink Trombone by making intelligible speech from an arbitrary vocal sample, with a tiny 3.4M-parameter model that can run in real time on a CPU.

Samuel's parameter-path output lends itself naturally to modification via code, establishing a dialogue between live coding software and the voice. Samuel can be used with pre-recorded clips, or in real time. We hope that Samuel inspires future creative speech synthesis work.

Usage of generative AI

Generative AI was not used for the creative aspects of this work. The paper itself is hand-written, the UI is hand-designed, and the concept of the work was not discussed with an LLM.

For the implementation and training, we made heavy use of Claude Code (mostly with Opus 5). The majority of the code is machine-written, with light human review. The ability to quickly write code is incredibly useful for experimentation: trying out a new idea takes almost no time, and if proven by experiment, the code can be cleaned up to be integrated into the main codebase. Claude Code was also used to "babysit" experiments: occasionally, training runs crash with out-of-memory issues or other errors, or the experiment simply turns out not to be promising. Such issues can easily be automatically fixed with no downtime and without human intervention, tightening the feedback loop of the experiments.

Acknowledgments

I would like to thank Kyutai for letting me utilize unused compute capacity of the cluster to run experiments while training Samuel. Kyutai is a non-profit open-source AI lab based in Paris; Samuel is not an official Kyutai product. To not take compute from official projects, I made sure to use only 1-2 GPUs from the partition dedicated to interactive jobs, which always has a few unused GPUs. A big thank you to my colleagues for all the speech AI wisdom without which I'd have no idea how to approach this project!

I would also like to thank Neil Thapen for creating Pink Trombone, and Zack Qattan for the programmable version that served as the starting point of our implementation.

← Back