Headline result (§5b): the model was only saturated on clean studio audio. On real,
in-the-wild user audio it had huge headroom (24% CER) — and a small batch of consented
user recordings closed a third of it (24% → 16% CER, 86% → 34% WER) with no loss on the
studio benchmarks. The data flywheel works.
Off-the-shelf Sanskrit ASR is trained on conversational IndicVoices-style data and degrades badly
on recitation and śāstra — dense compounds, sandhi, retroflex/aspirate contrasts, pitch, and
long metrical utterances. The goal here was a model and tooling good enough for scholars:
accurate on chant and prose, and useful as a practice aid rather than a transcription toy.