For language teachers
Give tonight’s feedback without playing thirty files.
You do not need the waveform. You need three filters in this order: confidence under 0.5 (re-record, do not grade), decodedTranscript versus the assigned phrase (they said something else), then words below 50 and the phonemes inside them. Phone PCC 0.682 versus human consensus 0.555 is why those phonemes are a better note than “listen again.”
Three filters, then a note — never a second listen first
A folder of homework is not a playlist. It is a table. Confidence, decodedTranscript, word scores, phonemes. The ear comes in only for the rows that survive the table — and even then, only if you still disagree with a red cell. Most nights you will not need the ear.
- Confidence under 0.5 — throw away. The scoring guide says the audio may not match the reference. Grading it trains the student to game a broken take.
- decodedTranscript far from the phrase — they did not say the homework. That is a task error, not a phoneme error. Ask for the phrase again.
- Words below 50 — open `phonemes[]`. Write the two worst phones in IPA. That is the note.
The note is a phoneme, not a pep talk
The engine’s own coaching prompt is the procedure we use internally: strengths at 80–100, work items below 60–79, top three sounds, minimal pairs, IPA. It also says: if confidence is under 0.5, state that the assessment may be unreliable before you say anything pedagogical.
That order is the whole job. “Practice more” is what you write when you skipped the tape. DH /ð/ in weather, then a re-record of the same sentence, is what you write when you read it.
Twenty students, one phrase, one pass
Same reference text for everyone or you cannot sort. Phone PCC 0.682 is phone-level agreement with trained humans at 0.555. You are not claiming the engine is a person. You are claiming it is more consistent with the human consensus than two humans are with each other — enough to decide who gets a note tonight without a thirty-file listen.