
Guitar Audio Transcription
A pipeline that turns a guitar recording into the symbolic information needed to build a tab: notes with timing, string and fret positions, rhythm, and playing techniques. It has no interface of its own and ships as the Import Audio feature inside Tab Renderer.
Stack: Python, Demucs, Spotify Basic Pitch, a reproduced FretNet, librosa, scikit-learn, mir_eval / mirdata, deployed on Modal.
Why this needed a separate project
I wanted to build audio transcription functionality into my Tab Renderer project, and this is a hard problem I was never going to solve quickly or get right on the first attempt. That is exactly why I split the project into four stages — ingest, separation, transcription and techniques — from the start, so I could build with intent: each stage would get a clean environment of its own, where I could iterate and push accuracy without worrying about the rest of the project.
Every stage does one job and hands on the same fixed set of information, so I can swap out how any one of them works and put in a different method without touching the others, and I choose which method to use in a settings file rather than by editing code. Each stage also has its own test that scores a new method against the one it would replace, so a change survives when the accuracy increases. There is always a working version of every stage, which means there is always a baseline to beat.
The pipeline
Stage 1: ingest
Read the audio file, mix it down to a single audio channel, and resample it to 22.050 kHz. That is half of CD quality, which is still comfortably more than this needs: it captures everything up to a pitch of about 11 kHz, and a guitar's notes and the harmonics that identify them sit well below that.
Stage 2: separation
Seperate the guitar track from the rest of the audio. There are two options:
passthrough: if the recording is already just a guitar, separation can only damage it, so the right move is to pass it to the next stage untouched.
Demucs is a neural network that takes a mixed recording and splits it back into separate instrument tracks. I use the htdemucs_6s version specifically because it is the only widely available model with a dedicated guitar output. The more common four-track models lump guitar together with everything that is not drums, bass, or vocals, which is no use here.
Testing it needed a recording where I already knew the correct answer, so I generated one: a synthetic guitar phrase saved on its own as the reference, then mixed with synthetic drums and bass. I ran that mix through Demucs and compared its guitar output against the original phrase. The measure used is SDR, which compares the energy of the true signal against the energy of the error; this run generated 5 dB, which is significant.
Stage 3: transcription
Detect notes in the audio that can be mapped to frets. I tested two note detectors: Basic Pitch and FretNet. Basic Pitch is the better note detector whilst FretNet predicts string and fret directly, which is the information Tab Renderer actually needs. The default implementation is neither of them alone: fusion pairs Basic Pitch's notes with FretNet's string and fret head, then withholds the assignments FretNet is not confident about.
The confidence threshold is a config value rather than a constant buried in code, because it trades accuracy against coverage. At gate: 0.05, string accuracy on correctly-pitched notes rises from 0.79 to 0.83 while 94.7% of notes still keep a string. The remaining 5% get no string rather than a wrong one.
Getting a FretNet model. FretNet's authors published their code but never a trained model. This meant I had to train it myself. Training needs recordings where the right answer is already known, and for string and fret that is a demanding requirement. GuitarSet is the standard dataset for exactly this. It is acoustic guitar recorded so that each string is captured separately, which makes its per-string labels a measurement rather than somebody's transcription.
That training is normally run as six-fold cross-validation. GuitarSet contains six players, and a fold holds one player back, trains on the other five, then tests on the player it never heard. Running all six shows whether the model generalises to an unfamiliar guitarist instead of learning the habits of whoever it trained on.
Two practical problems got in the way. One of the original author's dependencies does not build at all, so it has to be replaced with a stub for the import to succeed. In the end I ran two folds rather than six, trained on Kaggle: they agreed closely enough that more were unlikely to change the answer, and each one took considerably longer on the free GPU quota than I had planned for. I needed one usable model, not a publication-grade evaluation, so I stopped there. fold-0/model-2500.pt is what the pipeline loads.
FretNet's research code also cannot be installed alongside gtab's own dependencies, so it runs in a separate conda environment as a subprocess, with fretnet_client.py marshalling audio in and predictions back out.
Stage 4: techniques
Detect playing techniques to give style to the transcription. Again I utilised two detectors, because timbre and pitch gestures need different evidence. A RandomForest over a 30-dimensional per-note feature vector handles palm-mutes and harmonics. Bends and slides use FretNet's per-string pitch instead, which is monophonic per string and therefore separable in polyphonic audio.
This is the stage that is not finished. Bends and slides cannot be told apart because FretNet's continuous pitch output is capped at plus or minus one semitone, which truncates the large glide that distinguishes them.
Running in production
In production, the models run on a Modal GPU service rather than on the web host, which has neither a GPU nor the runtime for a long job. Tab Renderer Vercel build posts an audio file to POST /api/transcription-jobs and polls GET /api/transcription-jobs/{job_id} until the job completes. The service scales to zero between jobs, so the first request after an idle period takes a few seconds to wake.
What comes back is not raw notes but a ready-to-edit score: measures, beats, and string and fret positions, plus a warning when a note cannot be assigned to a unique string.
Final thoughts
The first thing I would change is testing Stage 2 properly, on a real multi-instrument dataset where the isolated guitar track is published alongside the full mix, telling me whether separation is worth its runtime cost on real music or whether it is quietly damaging the signal it is meant to clean up. I would also cache FretNet's per-string pitch on the fusion result, since running fusion and glide detection together currently invokes FretNet twice on the same clip for data the first pass already computed. Additionally, I would have started on electric guitar much sooner. FretNet was trained on acoustic recordings, which is where the drop in string accuracy on anything else comes from, and technique detection then hits a ceiling caused by that same model, so the two biggest limitations in the project both trace back to one gap in the training data.
What's in the repo
src/gtab/types.py: the data contracts every stage agrees onsrc/gtab/stages/base.py: the three stage interfacessrc/gtab/stages/transcription.py: stub, Basic Pitch, FretNet and fusionsrc/gtab/stages/fretnet_client.py: subprocess bridge to the isolated FretNet environmentsrc/gtab/stages/techniques.py: learned, glide and contour detectorssrc/gtab/eval/: SDR, note F1, string accuracy, per-technique precision and recallsrc/gtab/config.py: implementation registries and YAML pipeline buildingscripts/: pipeline CLI, evaluation harnesses, classifier trainingdocs/: per-stage specs and handoffs, including the ordered frontier list