
A 54.9% relative reduction in word error and a 67.3% reduction in character error. The project reports identical clips and decoding settings for both models. These are repository-reported results, rather than a new benchmark run. Evaluation write-up ↗
A small update.
A multilingual objective.
- Trainable parameters
- 18.35M
- Optimizer steps
- 12,761
- Training epochs
- 1
| Base | Gemma 4 E4B / Unsloth / 4-bit |
|---|---|
| LoRA | Rank 8 · alpha 16 · dropout 0 |
| Batch size | 16 / gradient accumulation 1 |
| Learning rate | 0.0002 / warmup ratio 0.03 |
| Compute | Colab A100 and RunPod B200, as reported by the project |
The published Hugging Face artifact is a PEFT adapter and requires its compatible base model. Training-loss histories are not published in this repository. Training script ↗

Source audio comes from Chatterbox multilingual data ↗. Each clip produces three tasks: 66,814 transcription examples plus 133,628 translation examples. The validation split contains 3,712 source clips and 11,136 task examples. Translation labels partly depend on IndicTrans2; they are generated targets, not independently verified human translations.

Combined WER is 12.35% and CER is 3.44% across 15 clips. This run has no matched baseline and must not be compared directly with the 200-clip result above. Maithili has the highest word error in this small sample. A larger held-out evaluation, including noise, dialects, and mixed-language speech, is the next useful test. Model card and limitations ↗