Fine-tuning Qwen3-ASR for Sinhala – Part 8: Conversational Speech, and a 9M-Parameter Model

· #asr #sinhala #machine-learning #mlx · 14 min read

In part 7 the model was finished and on Hugging Face, and the next step was getting it into the app. Before that, I wanted to know how it handles the speech it will actually get. Every test so far was read speech: volunteers reading news and blog sentences aloud. Dictation is closer to conversation, and I had nothing to test that with.

This part covers that test, what a few hours of conversation did to the model, and the experiment it led to: a much smaller model, meant for phones, which is running on my Mac as I write this.

A Sinhala dataset on Hugging Face

SPEAK-ASR publishes 11 Sinhala speech datasets on Hugging Face. Most of them turned out to be data I already had:

The YouTube set is the interesting one. It’s people talking, not reading: news, technology, cooking, law, business. 54% of its clips mix English words into Sinhala, and its transcripts write those words in English letters. That’s the dictation style I’m after, and OpenSLR has almost none of it (2% of its sentences, part 3).

NB: SPEAK-ASR’s own model reports 10.8% WER, but on their random split, so its test speakers were most likely in its training data too, and its text is normalised differently. It can’t be compared with the 27–30% in part 7.

YouTube, after all

In part 1 I ruled out YouTube. The recordings belong to their channels, and a model trained on them is a legal risk for an open source app. That’s still true, and SPEAK-ASR publish this set without a licence.

So I trained on it as a local experiment first, only to see what it would change, with nothing published. Before publishing anything trained on it, I asked SPEAK-ASR’s team, and they were happy for me to go ahead. The model card lists the dataset, says it has no licence and was used with its authors’ permission, and that the recordings belong to their channels. The model’s own licence stays CC BY-SA 4.0, because of OpenSLR 52.

Splitting by channel

Part 7’s lesson was to check what a test set has already seen. SPEAK-ASR’s own split puts clips of the same video on both sides: 58 of the 111 videos are in both their training and their test set. So I split it again, by channel, so that no test channel’s speakers or videos are heard in training:

The transcripts got the same clean-up as OpenSLR’s, including English words in Sinhala letters rewritten in English letters (part 3).

The published model on conversation

I ran the published model, the first revision, over the five test channels first. On read speech it gets 7.08% of characters wrong. On conversation it got 32.99%, and 56.39% of words. Two things stood out.

English words came out in Sinhala letters. On read speech it wrote 82% of English words in English letters. In conversation it managed 16%. “fundamental ideas” came out as “ෆන්ඩමෙල අයිඩාස්”, and “focus” as “ෆෝගස”. OpenSLR, with English words in only 2% of its sentences, hadn’t taught it that.

7 of the 566 transcripts ran away. The model repeated a word (“ඕනේ ඕනේ …”) until it hit the length limit. That’s 1 recording in 80, and those 7 held about 70% of all its character errors: counted as they came out, the CER was 104.87%. The 32.99% above cuts each loop where it starts, as a loop guard would. It’s the loop from part 5, and it gets past the app’s guard for the same reason: the speech library stops a transcript only when its last 24 tokens hold 3 or fewer different tokens, and one Sinhala word is several tokens.

Training on conversation

The run continued from the published model for three passes of 71 steps of 128 recordings, at a peak learning rate of 1×10⁻⁵, half the first run’s. Each pass mixed the 3,251 conversation clips with 3,251 random OpenSLR recordings and 2,500 replay recordings, so the model kept hearing read Sinhala and its other languages (part 3). It took just under three hours on the Mac.

On the dev channels, the checkpoints at steps 106, 177 and 213 all scored about 18% CER, so it had stopped improving part way through the second pass.

Setting the dial again

As in parts 5 and 7, I blended the trained weights back towards the model they started from, and tested a few blends of step 177:

Bar chart of the character error rate on the 566 conversation recordings: 18.16% keeping all of the new training, 20.46% at 75%, the published blend, 25.00% at 50%, and 32.99% for the first revision, which keeps none of it.
Conversation CER by the share of the new training kept. 0% is the first revision, and 75% is the one I published.

I published the 75% blend as the model’s second revision, 047217c, on 28 September.

Note: this time I chose the checkpoint and the blend by their test results, which part 7 warned against. The model card says so, and its figures may flatter this revision a little.

The results

Test First revision Second revision
Conversation, 5 unseen channels (566 recordings): CER 32.99% 20.46%
Conversation: WER 56.39% 41.47%
The 323 that mix in English: CER 43.21% 24.35%
The 243 without English: CER 17.46% 14.55%
English words written in English letters (of 1,148) 16% 59%
Transcripts that ran away 7 1
Read Sinhala, new sentences (2,547 recordings): CER 7.08% 6.86%
Read Sinhala, all 8,748 test recordings: CER 6.36% 6.11%
English, FLEURS test: WER 5.24% 5.29%

NB: the conversation figures cut each runaway transcript where its loop starts. Counted as they came out, the CERs are 104.87% and 31.38%. The original English-only model gets 4.96% on the same English test.

Most of the gain is where English is mixed in, which is also where most of the errors were. The recordings without English improved too, from 17.46% to 14.55%, so it isn’t only about how English words are spelt. All five test channels improved.

Two more limits on these numbers, besides the choice made on the test sets:

Spanish, French and German slipped a little further (5.61% to 5.91%, 10.07% to 10.25% and 8.70% to 8.85% WER). Chinese didn’t move.

A habit it picked up: digits

OpenSLR spells numbers out. SPEAK-ASR’s transcripts often use digits: 5.8% of its clips have one, against 0.4% of OpenSLR’s. The model learnt both styles. 9 of the 8,748 read-Sinhala transcripts now have digits where the reference has none, against 2 before. Most are numbers that were said, written as digits. Three times, though, it heard a number that wasn’t there: “1970” for “නිදහස ඔබ සතුය”, “23” for “විසිතුරු” and “150” for “පහසුව”.

One fix is to write the YouTube transcripts’ numbers out in words before training, so the two data sets agree.

A guard for the loops

The loops needed their own fix. I tried a stronger guard on the saved transcripts: stop when a phrase of 1 to 8 words repeats 5 times in a row, or when the output passes 25 letters per second of audio (plus 20). The fastest clip in the YouTube training set has 22.6 letters a second.

It stopped every loop in the tests, and cut short one correct transcript in about 220,000 (“left, right, left, right, …”). Stopping at 4 repeats instead of 5 cut “එක එක, එක එක”, which is a real phrase. The French loop, a 13-word phrase, is longer than the phrase check looks for, but the length cap stops it. In fact the cap alone catches every loop so far. It isn’t in the apps yet.

Publishing it

NB: uploading the 1 GB model hung for 1 hour 40 minutes without committing anything. Hugging Face’s uploader starts with a single connection, and that one ran at about 285 KB/s on a link that measures 148 Mbps, so each 62 MB chunk timed out and was retried. Forcing 8 connections finished it in about four minutes:

HF_XET_FIXED_UPLOAD_CONCURRENCY=8 HF_XET_CLIENT_READ_TIMEOUT=600s \
  HF_XET_CLIENT_RETRY_MAX_DURATION=3600s HF_XET_CLIENT_RETRY_MAX_ATTEMPTS=10 \
  hf upload …

The apps still use the first revision. Moving them to the second is a separate change.

Why a 9M-parameter model

The idea is a Sinhala model that runs on phones, iPhones and Android phones alike, not just on a laptop. The Sinhala Qwen model is about 1 GB to download and 782 million parameters to run, which is a lot to ask of a phone.

The conversation test also left me with two thoughts. A few hours of the right speech changed the model more than anything since the first run. And its worst failures, the loops, come from how it writes, not from what it knows.

Qwen3-ASR writes a transcript the way a chatbot writes an answer: one token at a time, each chosen after looking at the ones before. That’s what lets the app tell it the language (part 2). It’s also why it can talk itself into a loop, and why Sinhala, at about 8 tokens per second of speech, is slower than English.

The other kind of model is the kind Omnilingual is in part 1: CTC. It looks at the whole clip and picks a piece of a word, or a blank, for every slice of audio, all at once. Its transcript can’t outgrow the audio, because it has nothing left to write with when the audio ends. Omnilingual failed in part 1 because it knew the letters of 1,600 languages and couldn’t be told which one it was hearing. A CTC model whose vocabulary holds only Sinhala and English pieces can’t write Bengali, because it has no Bengali to write.

Qwen3-ASR: the audio goes through an audio encoder into a language model, which writes the next token and feeds it back in, until it writes the end or reaches the limit. Niagara: the audio goes through an encoder that picks a piece or a blank for each 40 ms, and the transcript is those pieces with the blanks and repeats removed.
Qwen3-ASR decides when to stop, so it can repeat itself until it hits the limit. A CTC model makes one choice for each 40 ms of audio, so its transcript can’t outgrow the audio.

Earlier this month, Applied Brain Research (ABR) published the smallest of its Niagara speech models on Hugging Face, niagara-9m-batch.en. Niagara models are state space models with attention, and ABR builds them for small devices. The family goes from 84 million parameters down to this one:

That’s why I wanted to try it:

Reason 1: it’s small enough for a phone. Part 1’s goal was a model light enough for a laptop, and this goes further. At 9 million parameters, the weights are about 36 MB in 32-bit floats, against about 1 GB for the Sinhala Qwen model at 8 bits.

Reason 2: it can’t run away. The loops were most of the first revision’s errors on conversation, and a CTC model can’t write past the end of the audio.

Reason 3: everything around the model carries over. The data, the splits and the tests from the Qwen model are all reusable, so the two models can be compared on exactly the same recordings.

It gives some things up, though:

Getting Niagara to train

ABR publishes the model as a compiled PyTorch graph, for running it, not code for training it. Two things stood in the way.

The weights were baked in. The graph stores each weight twice, once as a parameter and once as a constant, and it computes with the constants. So training would only have updated parameters that nothing uses. The trainer points each of the 683 constants back at the parameter it equals, 9.07 million weights in all, so that updates reach the graph. The other constants (the position tables, the state space layers’ fixed basis and a padding value) stay as they are.

It only ran on the CPU. The graph checks 2,423 times that its data is on the CPU, and names the CPU in 2,374 more places. Dropping the checks and pointing the rest at the Mac’s GPU made it run there. Its outputs stayed within 0.025 of the CPU’s, and a training step on 16 clips took 1.1 seconds instead of 5.4. Its English still works: “unfortunately studying traffic flow is difficult because driver behavior…”

Inside, it’s small and regular. It takes 80 audio features every 10 ms, and a convolution brings that down to one step every 40 ms. Then come 18 blocks that ABR calls LMUFormers, each with a state space layer, attention and two feed-forward layers, all 128 numbers wide. The output is one choice per 40 ms, from 1,024 pieces plus the CTC blank.

The plan

As I write this, a 100-step test run is checking that the training loop works end to end on the Mac’s GPU. The full run is next.

What I learnt

  1. Test on the speech you’ll really get. Read speech put the model at 7%, and conversation at 33%.
  2. A few hours of the right data can move a model more than hundreds of hours of the wrong kind. 7.4 hours of conversation cut the errors on conversation by more than a third, and by more than 40% where English is mixed in.
  3. Check someone else’s split before trusting it, or their scores on it. SPEAK-ASR’s splits put the same speakers, and the same videos, on both sides.
  4. New data brings new habits. The YouTube transcripts’ digits taught the model to write years nobody said.
  5. A few runaway transcripts can decide the average. 7 loops held about 70% of the errors. Guard the decoder, or use one that can’t run away.
  6. Read a model’s licence before its benchmarks. Niagara’s decides what can be done with anything trained from it.

Next: how the 9M-parameter model does against the 782M one on the same tests, and whether it’s good enough to put on a phone.