How the App Listens

The microphone is the most important piece of gear in this app. Everything Mankunku does — the scoring, the per-note timing, the level adjustments — flows from a single question, asked sixty times a second: what note are you playing right now? This page is about how the app answers that question, why it sometimes gets confused, and what you can do to give it the cleanest signal.

What the microphone hears

Your computer's audio system hands the app a stream of sound: the air pressure your mic detected, sampled tens of thousands of times a second. The app pulls a small sliding window from that stream — about 90 milliseconds of recent audio — and asks: is there a periodic waveform in this window, and if so, at what frequency?

If you're playing a steady note on a horn, the answer is usually clear: the air column in your instrument vibrates at a fundamental frequency, your mic captures that vibration, and the algorithm picks it up. The frequency converts to a MIDI note (440 Hz is concert A4; an octave up is 880 Hz; one semitone up from 440 is about 466). The app rounds to the nearest note and also reports how many cents flat or sharp you are.

The algorithm Mankunku uses is called the McLeod Pitch Method. It's an autocorrelation technique — it asks how well each segment of audio matches a delayed copy of itself, and the delay that matches best corresponds to the period of the note. It's particularly good at single-instrument signals like a sax or a trumpet, which is why it's the right tool here.

Each frame, the app gets a frequency and a clarity score between 0 and 1. Clarity tells the app how confident it is — a clean, sustained note has clarity above 0.92; a noise burst, an attack transient, or two notes overlapping might score 0.5. Mankunku ignores any frame with clarity below 0.80, so room noise and embouchure adjustments don't trigger phantom notes. Frames between 0.50 and 0.80 are set aside rather than thrown away: they're how the app hears a ghosted note (see below).

When the app starts listening

In ear training the microphone opens the moment your turn starts — before you play a note. That matters more than it sounds: a recording that waits for a confident pitch reading before it starts can't contain the attack of the note that triggered it, because a confident reading needs most of one analysis window of the note first. So the app arms the mic early and then trims the take back to a fixed third of a second before your first played note. Your reaction time is discarded, not scored, and you can take a breath before coming in.

"Played" is the operative word. The pitch detector's confidence has nothing to do with loudness, so the faint ring the metronome leaves in a quiet room reads as a perfectly confident note — and once anchored the trim there, three phantom notes out of a click's tail were scored against the line. Now any stretch of readings that never gets within 30 dB of the loudest thing in the take is thrown away before the trim looks for your entrance, wherever it sits — before you come in, in a rest, after the last note. A run that reaches playing level keeps everything, decay tail included.

Lick Practice and Tune Practice solve the same problem differently — the detector runs for the whole session and each playing window is sliced out of it on the bar line — and Record a lick schedules your entrance (the downbeat of bar 3, after a two-bar woodblock count-in), so it needs no reaction-time trim: a note that comes in a hair early is kept and pulled onto the beat.

Detecting where each note begins

Knowing the pitch of every frame isn't enough — the app also needs to know when one note ends and the next one begins. A scale played fast on a horn might have notes lasting 100 ms each, with crisp attacks; a ballad has long notes connected by gradual transitions. Both need to be sliced into discrete events.

The app does this with an onset detector that runs alongside the pitch detector, on a separate audio thread for low latency. It listens for sudden energy spikes in the high-frequency content of the signal — note attacks have more high-frequency content than the sustain of a held note, so a ratio jump in that energy is a reliable cue that a new note has just started. The detector enforces a small dead time (about 60 ms) after each onset so that fast trills don't trigger one onset per cycle.

If the onset detector is unavailable (older browsers, some mobile devices), the app falls back to inferring onsets from gaps in the pitch stream — a stretch of silence followed by a clarity spike means a new note. It's slightly less accurate but works.

Once the app has both the pitch readings and the onset times, segmenting them into notes is straightforward: each onset starts a new note; each note's pitch comes from a clarity-weighted vote over the pitch readings that fell inside its window — the pitch class first, then the octave, with a tie broken toward the previous note — so a single octave glitch in one frame can't change the answer (the cents figure is the median of the winning readings); each note's duration runs to the next onset or until the player stops playing. An onset only counts if the pitch detector actually heard a note in it — and heard it there: a reading whose analysis window already reaches the next onset is that attack's evidence, not this one's. That matters because the capture opens before you play, so the click you come in on is inside it; without the rule, your first note's own attack used to vouch for the click's onset and split the note in two.

After segmentation, the app runs cleanup passes that handle real-world failure modes the raw detector can't avoid:

  • Same-pitch consolidation. If two adjacent segments share the same MIDI pitch and the boundary between them has no AudioWorklet onset within ±75 ms — meaning there was no actual attack — they're merged back into one note. This catches octave glitches and clarity dropouts that briefly split a single sustained note into two.
  • Octave-boundary collapse. When the McLeod method temporarily locks to the wrong octave for a few frames, the app spots the artifact (three or more raw frames inside the segment match the lower fundamental) and merges the stray segment back into its neighbour.
  • Octave respell of a re-attack. A saxophone re-attack often speaks on its second harmonic for the first 50–100 ms before the fundamental fills in. When that transient gets cut into a note of its own — very short, exactly an octave above the note before it, with the lower fundamental still showing in the raw frequencies around it and the next note not continuing the upper octave — it keeps its attack and takes the neighbour's octave, rather than scoring as a wrong pitch.
  • Cracked attacks. A saxophone attack sometimes speaks in the wrong octave for its first tenth of a second before the octave vent takes over — no new attack, just a jump. The detector hears that as a short note an octave away, but the wrong octave never lasts long enough to be confirmed: once the detector's settling window after an attack ends, it wants three steady frames of an octave before it believes it. So the crack folds into the note it settles on, which starts where you tongued it. A real slur up or down an octave holds its first note long enough to be confirmed, so an octave leap you actually play stays two notes.

These passes are deliberately conservative: they only fire when the absence-of-attack evidence is unambiguous, so genuine re-articulations of the same pitch still register as separate notes.

Ghost notes

A ghosted note — half-fingered, breathed rather than blown, the swallowed off-beat in a bebop line — is exactly what the 0.80 clarity cutoff was built to ignore: breathy, quiet, and pitched somewhere between two keys. Left to the confident frames alone, a ghost leaves nothing but a short gap, and the notes on either side close over it. So the app looks inside those gaps at the frames it set aside. When a short gap (under about 0.4 s) holds a steady run of them — at least three frames agreeing on one pitch, at least three-quarters of a semitone from the notes on both sides, and no more than 20 dB quieter than they are — that's a ghost note, and it goes into the line where you played it.

Once found, a ghost is scored like any other note: by the nearest semitone to the pitch it actually sounded. Being quiet doesn't earn it any leeway — a ghosted C that sounds more than half a semitone sharp is heard as a C♯ and counts as a wrong note. Finding the ghost still matters even then: the rest of the line stays lined up with what you played, instead of the app treating the note as missing.

Telling a glitch from a real re-articulation

This is the hard part, and it's where most of the app's listening intelligence lives. When you play two notes of the same pitch back to back, something changes between them — but what changes depends entirely on how you tongue, and a hard tongue and a legato "doodle" tongue leave almost nothing in common.

So the segmenter looks for the re-attack in several different ways, from the most obvious evidence to the least, and takes the first one that fires:

  1. The onset detector saw it. A hard tongue puts a broadband spike in the signal and the worklet catches it. Done.
  2. The signal dropped out. The air column destabilised enough to break the pitch track for a frame or two, and the energy clearly steps back up afterwards — or, for a tongue that restarts the note a shade softer, the instrument-band floor collapses across the hole (something a click can never do, since a click only adds energy) and the note then holds its level rather than fading.
  3. The envelope dipped and recovered. No dropout, but a real dip in the note's loudness — measured on a short sliding window, because the ~93 ms analysis window smooths a 20 ms tongue stop completely out of view.
  4. The brightness spiked. No envelope dip at all, but a burst of high-frequency energy — the signature of a light tongue that reshapes the tone without interrupting it.
  5. The waveform shape broke. Nothing above fires. This is the legato tongue: the airflow never stops, the note gets louder across the re-attack rather than quieter, brightness climbs smoothly instead of spiking. What the ear hears is the reed being damped and restarting — the cycle-to-cycle waveform shape breaks for a few milliseconds and then settles into a new, brighter shape. The app measures that similarity directly and can split the note on it alone.

When waveform shape is the only evidence, that last tier requires a shallow break. A metronome click, a key click or a thump can also drive similarity toward zero, so a deep break alone is insufficient. A deeper reset can count when the instrument-band energy also dips and recovers. A quiet partial reset across a tracking gap needs a clean tone beforehand, a band-energy dip, no broadband noise burst, and sustained energy afterwards.

The envelope scan keeps looking for a tongue inside a gradual decrescendo. An existing onset near a metronome click is also preserved when the instrument-band floor stops and recovers on the same pitch class, even if octave tracking briefly glitches. For short notes, an unconfirmed octave held by the pitch stabilizer yields to a more strongly supported raw fundamental; a confirmed octave keeps its vote. The October 10 UTC Blue Note Step-Up, Climb Through the Blue and Flat Five Chromatic Down recordings cover these cases in the audio regression suite.

The net result: legato held notes stay one note, repeated notes stay separate, and you don't have to tongue hard to make the app hear what you're doing.

Why the room matters

The pitch detector works best when it's clearly hearing one note at a time. Three things commonly degrade that:

  • Background noise. A loud HVAC hum, a fan, traffic outside — these add broadband energy that drops the clarity score and confuses the autocorrelation. The app raises the clarity threshold to 0.80 specifically to filter out signals that aren't periodic enough to trust, which means in a noisy room the detector will simply miss notes rather than report wrong ones.
  • Speaker bleed. If your speakers play the original phrase loudly enough that the mic re-hears it, the detector treats those notes as if you played them. The bleed filter (see below) helps, but headphones eliminate the problem entirely. Earbuds work fine for this purpose. Worth knowing: the rhythm section is deliberately mixed to sound like a band in a small room — the snare and piano carry an ambience send, and the instruments are spread across the stereo field. That's the right choice for your ears and the wrong one for your microphone, so if you're on speakers, keep the backing level modest.
  • Multiple sources. Two horns playing at once will trip up the pitch detector — it's designed for monophonic signals. So is most other pitch detection software; this is a fundamental limit of the autocorrelation technique, not a Mankunku-specific issue.

The bleed filter

For people who can't or don't want to use headphones, the app runs a bleed filter between detection and scoring. When you're on Side B with a backing track playing through speakers, the filter knows what notes the backing track is currently playing and which onsets the app is generating. For each note your microphone detected, it asks:

  • Is the backing track playing this same pitch class right now?
  • Is the clarity below the threshold for "definitely you playing"?
  • Did a backing-track event start within 50 ms of when this note was detected?

If yes to all three, the filter drops the note as bleed. If your clarity is high (≥ 0.92), the filter keeps the note even if the pitch matches — that's you playing the same note as the backing, which is musically correct.

The filter is conservative on purpose. False positives (dropping notes you actually played) are worse than false negatives (keeping a few bleed notes), so the threshold is biased toward keeping ambiguous notes.

The filter always runs — it's how the /diagnostics A/B comparison gets its two scores — but whether its result becomes your actual score is a separate internal flag, currently off by default. Tune Practice turns it on at every strictness level. If you're using headphones none of this matters: there's no bleed to filter, so the filter finds nothing either way.

Metronome and backing-track bleed

The metronome click is its own kind of bleed, and it's a nastier problem than speaker bleed because of where it lands: on the beat, which is exactly where notes start.

Even on headphones, the playback engine schedules the click on the same internal timeline the app records from. So rather than logging what it played, the app computes when each audible event must have fired, and treats those moments as suspect:

  • With the band playing, the band is the time source. The click only sounds for the count-in bar and then stops, so what the microphone can hear from that point on is the rhythm section — and the app knows the exact instant every bass note, comp chord and drum stroke was triggered, including swung ride eighths and off-beat piano pushes. (A plain quarter-note click grid would be worse than useless here: it claims a click on beats where none sounded, and never covers the off-beat content that actually bleeds.)
  • With the metronome alone, the clicks land on exact multiples of the beat, which is trivial to compute from the tempo.

Any onset landing inside a tight 50–200 ms speaker-to-microphone window after one of those moments is treated as bleed and won't split a note. Demo and melody playback don't feed the filter. Recordings keep their backing onsets alongside the audio, so replaying one on the diagnostics page segments it with exactly the evidence the live take had.

That handles a click inventing a note. The subtler failure is a click sitting on top of the evidence for a note you really did play, and it goes in both directions:

  • A click can mask a real articulation. The downbeat kick is the worst offender — a synth membrane sweeping ~2 kHz down to 33 Hz, roughly twice a ride cymbal's level, blanking pitch tracking for 100–150 ms once per bar. Tongue a note under one of those and the evidence for your attack is buried.
  • A click can fake one. The kick wipes clarity mid-note; the segmenter's clarity tier reads the wipe as a tongue stop, pairs it with a vibrato trough a moment later, and splits a held note in two — which slides the whole alignment by one and reports the next note as a wrong pitch.

The fix is to measure in a band the metronome cannot reach. The ride is high-passed at 8 kHz, the hi-hat at 6 kHz, and the kick's body sits below 250 Hz — so a 250–5000 Hz "instrument band" hears your horn at full strength and a bare cymbal about 25 dB down. Since a click can only ever add energy, a dip in that band's floor is evidence no click can manufacture. The segmenter uses it to keep trusting real articulations that happen to land on the beat, and to stop believing dips that are just the kick.

The result: your sustained notes stay sustained, your on-the-beat tonguing still registers, and the scorer doesn't penalise rhythm for phantom subdivisions you didn't play.

Two honest caveats. First, in ear training — where the mic opens before your turn — the app's reckoning of when each click sounded has been measured running about a third of a second off the real clicks, so the click-suppression window described above isn't landing on them there yet. It's a known, open issue, and one more reason to wear headphones on Side A. Second, the instrument-band reasoning above is calibrated against the metronome's voices. The backing track's piano, bass and snare do carry energy in that 250–5000 Hz band, so through loud speakers they can fill or fake the dips this tier reads. Suppressing onsets near known backing events covers the common cases, but if you practise on speakers with the band up and see notes splitting where you didn't tongue, that's the reason. Headphones sidestep all of it.

Latency and reaction time

There's a small delay between you blowing a note and the app registering it: the microphone's sampling, the buffer that pitch detection works on, the screen refresh, and the speed of sound across the room all add up to typically 50–150 ms. This is constant — it's the same on every note — so the scorer subtracts the median delay across your matched notes and only judges your relative timing. You don't get docked for reaction time. (The math is in How Scoring Works.)

What you do get docked for is timing variation between notes: rushing one note and dragging another. That's because the latency correction subtracts the median; what's left is your actual jitter relative to your own internal clock.

Tuning

There's no live tuner in the practice rooms — deliberately; the score is the feedback. But tuning is still measured: every note the detector hears carries a cents offset — how flat or sharp you were relative to the nearest note (±5 cents is "in tune"; ±20 starts to sound off; ±50 is the edge between two notes) — and the scorer pays a small bonus for landing near zero. Each note's reading is the median over the note, so vibrato and a bend that resolves are read as the note they resolve to. The note-by-note comparison under an ear-training session on the Progress page, and the diagnostics page, show what was heard.

What this means for getting clean scores

A few practical things:

  • Use headphones for Side A if you can. It removes the speaker-bleed problem entirely.
  • Sit close to the mic but not so close that you saturate the input. A laptop's built-in mic works from a few feet away; a USB condenser is better.
  • Quiet the room as much as is reasonable. Turn off the fan, close the window. The app handles a moderate room tone, but it's listening for the clean periodic signal of your horn — anything else is competition.
  • If clean takes score low, check the detector. The diagnostics page (/diagnostics) replays your saved recordings and lists the notes it heard with their clarity. Sustained notes that come back chopped up or low-clarity mean the detector is struggling: move closer to the mic, or check whether something else is making sound in the room.
  • Don't overdrive. Most laptop mics will distort if you blow too loud into them. The pitch detector handles distorted signals badly because the harmonics get mangled. If the replay shows notes breaking up on loud passages, back off.