For language teachers
A high score with confidence under 0.5 is not a high score.
Phone PCC 0.682 versus human consensus 0.555 is an agreement figure on clean, matched takes. The same engine publishes the conditions under which that figure does not apply. If you ignore those conditions, you are grading noise and calling it phoneme work.
Confidence is a gate, not a decoration
Confidence runs from 0 to 1. It is built from transcript agreement, posterior quality, whether the score distribution looks like a real take, and the audio-quality block. The scoring guide is explicit: below 0.5, the audio may not match the reference text. That is a reject, not a “maybe.”
overallScore is itself confidence-weighted. A mismatched take is penalised there too. Read confidence first anyway. Teachers who sort by overallScore and skip confidence will promote lucky noise.
The recording has to be a recording
The engine states the floor: SNR 15 dB required, 20 dB recommended. Below 15 dB the take is rejected. Between 15 dB and 20 dB you get a `low_snr` warning — grade at your own risk, which means do not grade. Duration 0.5 s to 60 s. Sample rate 16 kHz (other rates are resampled). Mono. Microphone 15–30 cm from the mouth. Quiet room. No clipping.
`high_wer` means the decoded transcript diverges from the phrase you sent. That is not a student who “has an accent.” That is a student who did not say the sentence, or a take too dirty to align. Either way, not a phoneme lesson.
- `low_snr` — noise, not vowels. Re-record.
- `high_wer` — the heard text is not the assigned text. Re-assign or re-record.
- `low_confidence_reject` — the engine already refused. Do not override it with your ear to be kind.
PCC does not bless a dirty take
Phone PCC 0.682 is measured against trained human consensus at 0.555 on a public test set of matched utterances. It is not a promise that a kitchen recording of the wrong sentence is a 80–100 phoneme. The proof number and the reject rules are the same product. Use both.