The framework-agnostic engine for practising from a score: take capture, score following and grading, timing against a set tempo, and sight-reading exercises. It needs no framework and no room; @hiyve/react-music-performance builds the React hooks and components on it.
Framework-agnostic take capture for score-informed performance analysis. Records what a performer plays — MIDI events from a connected instrument and unprocessed microphone audio — on one shared clock, and packages the result as a take that can be saved, downloaded, and analysed later.
import { AudioCapture, MidiCapture, TakeClock, finalizeTake } from '@hiyve/music-performance';
// Create the context inside a user gesture (a click) so every browser lets it run.
const ctx = new AudioContext({ sampleRate: 48000, latencyHint: 'interactive' });
const clock = new TakeClock(ctx);
const midi = new MidiCapture({ inputId: selectedMidiInputId });
const audio = new AudioCapture({ profile: 'instrument', onLevel: (rms) => meter.set(rms) });
await Promise.all([midi.start(clock), audio.start(ctx, clock)]);
// … the performer plays …
const bundle = finalizeTake({
profile: 'instrument',
score: { name: 'Minuet in G' },
midi: { inputId: selectedMidiInputId, events: midi.stop() },
audio: await audio.stop(),
clock: clock.snapshot(),
});
bundle.take; // JSON-serialisable record (see `serializeTake`)
bundle.wav; // Blob of the captured audio, or null
| Field | Description |
|---|---|
midi.events |
Every MIDI message received, stamped in milliseconds since the take started. Note-on with velocity 0 is recorded as noteOff. System messages are not recorded. |
audio |
Sample rate, channel count, length, the take-clock time of the first sample, and the settings the browser actually applied to the microphone |
clock |
The performance.now() and AudioContext.currentTime reference points that tie the two streams together |
alignment |
The measured MIDI-to-audio offset, once one has been estimated |
score |
The piece the take was performed against (name, storage id, or embedded MusicXML) |
| Class | Description |
|---|---|
MidiCapture |
Records from one or all MIDI inputs. Devices that connect mid-take are picked up; disconnected ones are released. MidiCapture.listInputs() lists what is connected. |
AudioCapture |
Opens the microphone with a capture profile (instrument by default — no echo cancellation, noise suppression or automatic gain), records raw samples through an AudioWorklet, and reports a level for a meter. |
TakeClock |
One timeline for the take; converts MIDI timestamps and audio frames to take milliseconds. |
LiveSession |
Follows a performance through a score as it is played, from streamed audio or MIDI, and reports a verdict for every note — see Following a performance live. |
| Function | Description |
|---|---|
finalizeTake(input, wavOptions?) |
Build the take record and encode the audio as a WAV blob |
createTake(input) |
Build the take record only |
serializeTake(take) / parseTake(json) |
Write and read the JSON form |
takeFileName(take, extension) |
A filesystem-safe name for a take artefact |
summarizeTake(take) |
Note, pedal and duration counts |
encodeWav(channels, sampleRate, options?) |
Planar float samples → WAV (16-bit PCM or 32-bit float) |
compareCaptureSettings(profile, settings) |
Which processing flags the browser honoured for a profile |
readTrackSettings(track) |
What the browser applied to a microphone track |
estimateTakeAlignment(audio, midiEvents, options?) |
Measure a take's MIDI-to-audio offset by matching audio attacks to MIDI note-ons; one note gives a calibration-note alignment, several an onset-match |
estimateOffsetMs(reference, observed, options?) |
Constant offset between two onset lists (median of nearest-neighbour differences) |
detectOnsetsMs(samples, sampleRate, options?) |
Every clear attack in a signal, for clock alignment |
detectFirstOnsetMs(samples, sampleRate, options?) |
First rise above the noise floor, for a calibration note |
mixToMono(channels) |
Average channels into one |
parseMidiMessage(bytes, timeMs) |
One Web MIDI message → MidiEvent |
rms(samples) |
Level of a block of samples |
AudioCaptureOptions| Option | Type | Default | Description |
|---|---|---|---|
profile |
'voice' | 'music' | 'instrument' |
'instrument' |
How the microphone is opened |
deviceId |
string |
browser default | Microphone to use |
chunkFrames |
number |
4096 |
Frames per chunk handed to the main thread; rounded up to a multiple of 128 |
onLevel |
(rms: number) => void |
— | Level of each chunk, for a meter |
onChunk |
(channels: Float32Array[], frame: number) => void |
— | Every chunk as it arrives, for analysis while recording |
onError |
(error: Error) => void |
— | Capture failures |
MidiCaptureOptions| Option | Type | Default | Description |
|---|---|---|---|
inputId |
string |
all inputs | Capture from one input only |
onEvent |
(event: MidiEvent) => void |
— | Every recorded event |
onInputsChanged |
(inputs: MidiInputInfo[]) => void |
— | A device appeared or disappeared |
onError |
(error: Error) => void |
— | Access failures and missing inputs |
parseMusicXml(xml, options?) turns a MusicXML document (partwise, what notation apps export) into the analysis model: every part's notes with sounding pitch (transposition applied), onset and duration in quarter notes and in milliseconds at the written tempo, measure and beat, chord membership, grace flags, and tie-merged durations. Repeats are reported (hasRepeats), not unfolded.
import { parseMusicXml, scoreEvents, expectedWindow } from '@hiyve/music-performance';
const score = parseMusicXml(musicXmlText);
score.notes; // ScoreNote[] — sorted by onset
scoreEvents(score); // notes that start together, in order
expectedWindow(score, 7.5); // { previous, current, next } events around a position
| Function | Description |
|---|---|
parseMusicXml(xml, options?) |
MusicXML → Score. Options: defaultBpm (120), includeGrace (true), applyTransposition (true) |
scoreEvents(score, { fromQuarters?, toQuarters? }) |
Notes grouped by onset, optionally within a region |
notesAtQuarters(score, q) / notesSoundingAtQuarters(score, q) |
Notes starting at, or still sounding at, a position |
expectedWindow(score, q, before?, after?) |
The event at or before a position and its neighbours |
quartersToMs(score, q) / msToQuarters(score, ms) |
Convert through the score's tempo map |
tempoAt(tempos, q) |
Tempo in effect at a position |
pitchToMidi(step, alter, octave) |
Written pitch → MIDI note number |
Parsing needs a DOM (DOMParser); in the browser that is always present.
matchPerformance(score, played, options?) compares played notes with the score and gives every expected note a verdict — correct, missing, or wrong (a different note played in its place) — and marks unexplained notes extra. A tempo is fitted through the matched notes, so timing is reported relative to how the performer actually played, not to the printed tempo.
import { matchPerformance, playedNotesFromMidi } from '@hiyve/music-performance';
const played = playedNotesFromMidi(take.midi.events, take.alignment?.midiToAudioOffsetMs ?? 0);
const report = matchPerformance(score, played);
report.summary; // counts, pitchAccuracy, meanTimingMs, timingStdMs, tempoBpm, targetBpm
report.events[3]; // { onsetQuarters, predictedMs, playedTimeMs, timingMs, verdicts: [...] }
| Option | Default | Description |
|---|---|---|
chordWindowMs |
50 |
Played notes closer than this count as one chord |
gapCost |
0.9 |
Cost of leaving an event unmatched; lower favours missing/extra over wrong |
timingScaleMs |
300 |
Timing deviation treated as a full mismatch when refining |
refine |
true |
Re-align with timing once a tempo is known |
fromQuarters / toQuarters |
— | Grade only a region of the score (a practice loop) |
Verdicts carry timingMs (played minus expected after the tempo fit; positive is late) for correct and wrong notes.
| Function | Description |
|---|---|
gradeTake(score, take, options?) |
matchPerformance fed directly from a take's MIDI events |
gradeTakeFromAudio(score, audio, options?) |
The same grade from the audio alone — for instruments without MIDI; also returns what was heard |
retimeReport(report, { bpm, startMs, atQuarters? }, offsetMs?) |
The same grade with its timing judged against a set tempo — a metronome's or a backing track's — instead of the tempo fitted from the playing; offsetMs moves the played times onto the tempo's clock first |
performanceScore(report, { timingToleranceMs?, offTimeCredit? }) |
The score a graded take earns, 0–1: a right note on time earns its point, a right note outside the tolerance half (by default), a wrong or missing note nothing; with the pitch accuracy and, when timing counts, the share of right notes on time |
gradeTakeFinal(score, bundle, { live? }) |
The score a take ends up with: from its MIDI when any was recorded, else from its audio when something was heard in it, else the real-time session's report — and every grade it could have had, with which one stood |
LiveSession |
The same verdicts while the performer plays — see Following a performance live |
playedNotesFromMidi(events, offsetMs?) |
A take's note-ons as PlayedNotes, optionally shifted |
midiToNoteName(midi, { flats? }) |
60 → C4, 66 → F#4 |
scoreNoteName(note) |
A score note spelled as written, e.g. Eb5 |
The audio analysis finds attacks in a take's audio and decides which notes sound at each — either with no help (analyzeTakeAudio) or by checking the notes a score expects at known moments (verifyTakeAudio). It works from the harmonic series of each note: how strongly the spectrum just after an attack supports a note's fundamental and partials, measured two ways — relative to the loudest bin (which note dominates) and relative to each partial's local noise floor (whether a quiet note is there at all). Accepted notes are subtracted so shared partials are not counted twice; a lone dominant fundamental still counts as a note; a candidate that is really the note an octave below yields to it; a quiet chord tone counts when its partials stand clear of their floors and are not just the harmonics of a note already accepted; attacks that are not tones (a knock, a click) are skipped; and — when a score is given — a semitone neighbour or the octave below that is clearly louder overrules the expected note. The score lowers the bar for the notes it predicts; it never invents one — and an expected note needs a series of its own (its fundamental, or two of its first three partials), so a note played an octave or a twelfth above an expected one is not mistaken for it.
gradeTakeFromAudio applies this in two passes: the audio alone is aligned to the score, then each matched moment is re-checked with the notes the score expects there — the offline form of follow-then-verify.
import { analyzeTakeAudio, audioNotesToPlayed, matchPerformance, verifyTakeAudio } from '@hiyve/music-performance';
// What the audio alone hears, then graded like MIDI would be
const events = analyzeTakeAudio(bundle.audio); // [{ timeMs, notes: [{ midi, confidence, evidence }] }]
const report = matchPerformance(score, audioNotesToPlayed(events, 0.3));
// Given expected notes at known moments (from a score position, or from a take's MIDI when evaluating)
const checks = verifyTakeAudio(bundle.audio, [{ timeMs: 1381, midis: [60] }]);
checks[0]; // { timeMs, present: [...], missing: [...], extra: [...] }
| Function | Description |
|---|---|
analyzeTakeAudio(audio, options?) |
Attacks and the notes heard at each |
verifyTakeAudio(audio, expected, options?) |
For each expected moment: which expected notes are present, which are missing, what else was heard |
audioNotesToPlayed(events, minConfidence?) |
Heard notes as PlayedNotes for matchPerformance |
estimatePitches(spectrum, binHz, options?) / verifyPitches(spectrum, binHz, expectedMidis, options?) |
The same decisions on one spectrum |
SpectrumAnalyzer |
Magnitude spectra at any position of a signal (reusable FFT and window) |
harmonicEvidence / subtractHarmonics / midiToHz / spectrumMax |
The building blocks |
synthesizeTone / synthesizeNotes |
Harmonic test signals |
Key options (EstimatePitchesOptions): partials (6), toleranceCents (50), tuningCents (0), presentThreshold (0.1, for expected notes), extraThreshold (0.2, for unexpected notes), weakPresentThreshold (0.01) with consistentPartials (3) and partialSnrFloor (10×) for quiet expected notes, consistentExtraThreshold (0.1) for quiet unexpected notes, minMidi / maxMidi (21–108), competitorMargin (1.25), dominantFundamental (0.5), octavePreference (0.6), attributionShare (0.75), ringCents (250). Analysis options: windowSize (8192), readOffsetsMs ([15, 35, 55, 75] after the attack, averaged), maxFlatness (0.05), onsets.
Measured so far (keyboard played through a webcam microphone): every note of a scale heard correctly with or without the score, no false notes, timing within ±2 ms of the keyboard's MIDI; on a sequence of eight triads, 21 of 24 chord tones confirmed with the score and 16 of 24 without it, no false notes — the misses are chord tones recorded 20–30 dB below their neighbours. Octave doublings of a loud note remain ambiguous without an instrument profile. Acoustic instruments are the next thing to measure.
LiveSession does the same work while the performer plays. Feed it audio as it is captured (or MIDI notes as they arrive); it finds each attack, decides where in the score it belongs, and gives every expected note a verdict — first a quick provisional one from a short look at the attack, then a confirmed one once enough of the note has been heard. Every change is reported through onUpdate with a full snapshot, so a view only ever renders the latest state.
import { LiveSession } from '@hiyve/music-performance';
const session = new LiveSession(score, {
sampleRate: ctx.sampleRate,
onUpdate: ({ type, attack, snapshot }) => {
snapshot.cursor; // index of the last matched event in session.events, −1 before the first
snapshot.next; // the event expected next
snapshot.nextStable; // the event to point at — moves one step at a time until placements are confirmed
snapshot.verdicts.get(noteId) // { status, stage: 'provisional' | 'confirmed', timingMs?, playedMidi?, ... }
snapshot.extras; // heard notes no score note explains
snapshot.tempoBpm; // the tempo the performer is holding
},
});
worklet.port.onmessage = ({ data }) => session.pushAudio(data.channels); // chunks of raw samples
// or, from a MIDI instrument:
session.pushMidi({ timeMs, midi }); // notes within chordWindowMs form one attack
session.tick(nowMs); // closes an open chord once its window has passed
const snapshot = session.finish(); // judges any attack still waiting for audio
const report = session.report(); // the same PerformanceReport the offline grader produces
How an attack is placed: it is tested against the next event, a couple of events ahead (a skip), the current and previous events (a repeat, a step back) and the starts of the region and of the current and previous measures (a jump). A position is a candidate when at least half of its notes were heard — except the next event, which is always a candidate once the performer has started: a wrong note is still the next note, and is shown as such the moment it is heard. The cheapest candidate wins, where cost is pitch mismatch plus a penalty for the move. Following the performer, that is all: a note is right or wrong by its pitch and the follower moves on — timing against their own uneven pace says nothing about what they meant, except that an attack far sooner after the previous one than the next event could be due (tooSoonFraction of the way there) is a stray unless it sounds like the next event. Under a set tempo a forward move also costs how far the attack fell from its beat, and leaving an attack unexplained (an extra) costs a gap, up to a second one when the attack fell right where the next event was due. After a pause (late by more than pauseMs) jumps and steps back are cheaper — the performer stopped, and where they resume is a matter of pitch. After lostAfter attacks that matched nothing nearby the whole region is searched. Under a set tempo a stray attack is paired with a skipped event as a wrong note when a later attack reveals the skip.
Strict sequence. With follower: { strictSequence: true } none of the above applies: every attack is the next event, played right or wrong, coloured and moved past — no skips or jumps, no searching, and nothing held for the attack after, so verdicts are final as soon as the audio has been read. It is the simplest reading of free playing: look at the note, was it played right, move on. Two things are read differently. An attack that only finishes the chord the performer is on — notes of it not yet heard, with nothing outside it, when the next event was not played — stays on that chord, so a rolled chord is one chord; a single note is never finished, it is right or wrong once. And an attack far too soon after the previous one that sounds like neither is a stray.
A set tempo. By default the session follows the performer's own tempo, fitted from the notes they play, and timing is reported against that — a performer who plays evenly at the wrong tempo is not late on every note. Pass tempo: { bpm, startMs } (the stream time at which the region's first beat is due, after any count-in) to judge against a clock instead: timingMs is then early or late against the beat, clockQuarters in every snapshot says where the beat is, an event whose beat passes unplayed shows as missing (provisional, with attack: -1) once missAfterMs (400) has gone by and turns correct if a late note claims it, lateness is never read as a pause, there are no jumps back to the start of a bar, and a performer who stops is picked up at the beat they rejoin on, with everything passed over marked missing.
A legato line. An onset detector finds attacks: a note that begins with a burst of energy. A voice, a bowed string, a wind instrument or a sustained patch often moves to the next note with no attack at all, and those notes would go unheard. analysis: { pitchChanges: {} } — on a LiveSession or passed to analyzeTakeAudio and gradeTakeFromAudio — takes a change of pitch as an attack too: the dominant pitch is read every 40 ms and a move of a semitone or more that holds for three reads, longer than a vibrato swing, is an onset at the moment it first appeared. Options: hopMs, holdHops, minSemitones, minConfidence, minIntervalMs, maxFlatness, windowSize (defaultPitchChangeOptions). For such instruments a lower onset threshold and later reads (onsets: { thresholdDb: 6, baselineWindows: 60 }, readOffsetsMs: [60, 100, 140, 180]) suit the slower start of a note.
The metronome's click. Played through speakers, a click reaches the microphone a little late and, landing inside the window a note is read from, ruins the reading — a burst of noise is not a note. Add click to tempo — the click as it was played (metronomeClick(sampleRate) makes the standard one) and how many clicks sound before the first beat — and the session finds the count-in clicks in the audio, learns how late they arrive (snapshot.clickOffsetMs), judges the beat from then on where the microphone heard it — the notes come in by the same path, so a performer playing to the clicks is on time — and silences 30 ms around every click before the onset detector and the pitch reads see it; a note that comes in on a click, its first milliseconds silenced with it, is taken to have started on the click. Through headphones no click is found and nothing is touched.
For a view that points at the performer's place, nextStable is the event to point at: the one after the last confirmed placement (committedCursor), or one further when a provisional placement has moved on by a step — a provisional skip, jump or step back does not move it until it is confirmed, so it does not leap about on a wrong note. For the same reason a provisional skip does not call the notes it passed over missing until it is confirmed; a provisional wrong note is red at once.
Two things keep a verdict provisional. The quick look only ever says heard — a note it missed may still be in the full read, so missing and wrong wait for that. And a placement the full read cannot settle on its own — a wrong note and a skipped note look alike until the note after them — stays provisional until the next attack has chosen between the alternatives (or commitAfterMs has passed); a clean next note, well ahead of any other reading, is confirmed at once.
| Option | Default | Description |
|---|---|---|
sampleRate |
— | Required for pushAudio; omit for a MIDI-only session |
tempo |
— | { bpm, startMs, missAfterMs? }: judge against a set tempo instead of the performer's own |
fromQuarters / toQuarters |
— | Follow only a region of the score |
analysis |
as analyzeTakeAudio |
Attack detection and pitch verification for the confirmed stage |
provisional |
{ windowSize: 4096, readOffsetMs: 15 } |
The quick first look; false to wait for the confirmed verdict only |
follower |
see below | How positions are weighed |
chordWindowMs |
50 |
MIDI notes closer than this are one attack |
onUpdate / onError |
— | Every change, with a snapshot; errors thrown by onUpdate |
Follower options (LiveFollowerOptions): strictSequence (false — every attack is the next event, see above), maxSkip (2), gapCost (0.9 — an event passed over, or an attack nothing explains), repeatCost (0.5), backCost (1.2), jumpCost (1.2 — dearer than a wrong note, so a wrong note that happens to match the opening stays a wrong note), jumpCostAfterPause (0.3), resyncCost (0.5), pauseMs (2000), timingScaleMs (300 — widened for slow music and uneven performers), minSupport (0.5), tooSoonFraction (0.35), unheardWeight (0.5 — an unheard expected note counts half as much against a position as an unexpected heard note, because the analysis misses quiet chord tones far more often than it invents notes), minConfidence (0.3), lostAfter (3), tempoWindow (8), attackRise (1.3 — a repeat needs the notes to have grown since just before the attack, so a note still ringing is not a note played again), decisionMargin (0.3 — how far ahead of its runner-up a placement must be to be confirmed without waiting for the next attack), commitAfterMs (1000).
Replaying the recorded takes as if live (85 ms chunks): provisional verdicts land about 150 ms after the attack and confirmed ones about 290 ms; every note of the scale and every chord of the triad sequences is followed in order, with the same verdicts the offline two-pass grader gives; processing stays under 15 ms per chunk.
The building blocks are exported for other uses: OnlineOnsetDetector (the streaming form of detectOnsetsMs), SampleRing, rankAttack / followAttack / initialFollowerState (the pure placement decision — every way of placing an attack, likeliest first, or just the best — with exactVerifier for notes known exactly), fitTempo / predictMs / fitBpm, and pairVerdicts.
AudioCapture.start() returns what was actually applied; compareCaptureSettings tells you which flags differ from the request.TakeClock records both reference points; estimateTakeAlignment measures the residual offset from the take itself (a single calibration note, or every note-on) and the recorder stores it as take.alignment. Add midiToAudioOffsetMs to a MIDI time to find that note in the audio — on a USB keyboard it is typically negative, because the sound reaches the microphone before the MIDI message reaches the browser.getUserMedia, AudioWorklet, and (for MIDI) Web MIDI. Web MIDI is not available in Safari; audio capture works everywhere.@hiyve/rtc-client (capture profile constraints).generateSightReading(options) writes a short tune to order and returns it as MusicXML, ready for the same viewer, follower and grading as any score, together with its notes and rests (events, each with every note sounding) and the key's name. The tune reads like something written rather than a run of random notes: a chord for every bar from a four-bar phrase pattern, chord tones on the strong beats and steps between them, leaps that turn back by step, a line that does not run one way for long, the tonic at the start and the end with a leading note or supertonic before it, and a rhythm for each bar from a vocabulary for the time signature — the bar that closes a phrase settling on a longer note, bars one and three of a phrase sharing a motif. Asked for chords, it puts a three-note chord under the tune where it sits on the bar's chord — on the first beat of each bar, on the strong beats, or on every note long enough, as the difficulty rises — in close position with the tune on top, so what is played can be heard note by note. The same seed and options make the same exercise.
| Option | Default | Meaning |
|---|---|---|
fifths, mode |
0, 'major' |
The key, as sharps (positive) or flats (negative), −6 to 6, major or minor |
timeSignature |
'4/4' |
'4/4', '3/4', '2/4' or '6/8' |
measures |
8 |
Bars, up to 1023 |
clef, lowestMidi, highestMidi |
treble, two octaves from C4 | The clef — treble, bass, alto or tenor — and the range; by default two octaves from C4, C2, F3 or D3 |
difficulty |
2 |
1: steps and small leaps, quarters and halves. 2: eighths, dotted notes and rests. 3: busier rhythms, leaps to a fifth. 4: sixteenths, syncopation, leaps to an octave. 5: chromatic neighbour notes too. With chords on, the level also says where they fall |
rhythms |
what the difficulty allows | Which rhythms may appear: 'whole', 'half', 'quarter', 'eighth', 'sixteenth', 'dotted', 'rests', 'syncopation' |
chords |
false |
Three-note chords under the tune, on notes of a quarter or longer that sit on the bar's chord: at level 1 on the first beat of each bar, at 2 on the strong beats, from 3 on every such note. In minor the dominant's third is raised unless the tune has the natural seventh in that bar |
tempoBpm |
80 |
Written on the score, in quarter notes per minute |
seed, title |
1, made from the settings |
Reproducibility, and the title on the score — by default the key and the settings |
Every exercise has a code: sixteen letters and digits carrying every option that decides its notes, the seed included. sightReadingCode(options) is the code those options make and parseSightReadingCode(code) the options a code stands for (null for anything that is not one; case, spaces and dashes do not matter), so a code can be kept with a result, searched for, or handed to someone else to get the very same exercise. The tempo written on the score is not part of the code.
sightReadingKeyName(fifths, mode) gives the key's name, as "F♯ major"; createRandom(seed) is the seeded random source the generator uses.
@hiyve/music-performance — framework-agnostic take capture for score-informed performance analysis.
Records what a performer plays — MIDI events from a connected instrument and unprocessed microphone audio — on one shared clock, and packages the result as a take that can be saved, downloaded, and analysed later.
Example