NVIDIA’s workflow combines selective data cleanup, replay examples, and decoder settings for Najdi and Hijazi speech.

NVIDIA has published a technical tutorial showing how to adapt its Nemotron 3.5 automatic speech-recognition model to Saudi Arabic dialects. The initial experiment targeted Najdi and Hijazi speech while checking whether selected English and Arabic capabilities remained intact.

Automatic speech recognition, or speech-to-text software, often learns from broad datasets. Those datasets may not contain enough examples of regional accents, local recording conditions, or less common speech patterns. A model can perform well on broad Arabic or English tests and still struggle with everyday speech in a specific region.

The tutorial uses NVIDIA’s NeMo framework and focuses on a practical question: how can engineers specialize a model without training a new one from scratch or abandoning its previous work?

Keep difficult examples that matter

The first step is preparing the target speech corpus. NVIDIA’s workflow removes unusable labels, missing text, and recordings where the transcript clearly does not match the audio.

But it does not discard speech simply because the model performs poorly on it. That distinction matters. Difficult accents and noisy recordings may be exactly what a deployed system needs to understand.

For the reported experiment, NVIDIA retained 103,559 of 125,490 utterances after its initial checks. That was 82.5% of the starting set, representing 133.7 hours of speech.

The stated goal was to remove structurally bad examples, not hard dialect speech. NVIDIA also describes inspecting audio-quality score distributions before choosing filtering thresholds, rather than applying generic settings automatically. That approach can help prevent a quality filter from removing too much useful regional speech.

Replay lets the model practice old skills

Fine-tuning means continuing to train an existing model on a new task or type of data. The danger is that a model trained mostly on new examples can weaken at older tasks. This problem is often called catastrophic forgetting.

NVIDIA’s answer is replay mixing. During training, the model sees mostly Saudi speech but also revisits a smaller amount of data from earlier training. The model keeps practicing selected older behaviors while learning the new dialect.

The reported training mix contained 90% Saudi speech, 7% English data from the FLEURS dataset, and 3% Arabic FLEURS data. The replay data was mixed by explicit weights so it appeared throughout training.

This method does not protect every capability automatically. Replay protects what its examples represent. If a team needs to retain certain languages, accents, names, or speaking styles, those capabilities need suitable examples in the replay set and separate tests afterward.

That makes replay an engineering decision, not a magic safety net. The replay data is effectively a list of the older skills the team has chosen to keep checking.

Update as much of the model as the data supports

Fine-tuning can update all of a model’s speech-processing layers, or only selected layers. NVIDIA describes partial encoder unfreezing, which allows some layers to change while keeping others fixed.

Updating fewer layers can reduce training cost and memory use, but it can also limit adaptation. In this experiment, NVIDIA tested different adaptation depths on 134 hours of target speech. Updating all 24 encoder layers performed better than updating only the top six or top eight.

On the Najdi and Hijazi test split, the reported word error rate fell from 55.05% before fine-tuning to 29.96% after the full fine-tuning run. Word error rate counts incorrect, missing, and extra words, so lower is better.

The run lasted 12,000 steps and took about 4.5 hours on two GPUs. NVIDIA’s results are tied to this data, model configuration, and test setup. They do not establish that full fine-tuning will be the best choice for every corpus.

For a smaller or more constrained project, freezing more layers may be useful because it lowers the amount of computation and limits how many parameters change. The trade-off is that the model may adapt less fully. Teams would need to compare those choices on their own target and retention tests.

The reported checks found no English or Arabic drop

NVIDIA also measured selected FLEURS retention checks after fine-tuning. English word error rate moved from 11.04% to 10.42%, while Arabic moved from 12.67% to 11.41%.

In the same reported evaluation, the full SADA test set improved from 58.84% to 35.61% word error rate. These results show how the model performed on the listed checks, but they do not prove that every previous language or speech condition stayed unchanged.

That limitation follows directly from the replay design. A model can only rehearse and measure the capabilities represented in the replay and evaluation data.

Decoding can improve accuracy without retraining

Training is not the only adjustment available. After fine-tuning, engineers can change decoding: the process that turns the model’s possible outputs into a final transcript.

NVIDIA reports that switching to its highest-lookahead attention context reduced word error rate by 1.31 percentage points without retraining. The cost was approximately 800 milliseconds of additional latency because the system waited for more future audio before committing to text.

That trade-off may suit recorded meetings, media, or call archives. It may be unsuitable for live captions or voice conversations, where delays are more noticeable.

Beam search offers another option. Instead of quickly choosing one transcription, it considers several likely possibilities. In NVIDIA’s reported tests, a beam-search configuration reduced word error rate further, but it also involved a runtime trade-off. The useful setting depends on whether the application values maximum accuracy or rapid output.

What teams should take from the workflow

The tutorial is a starting point for dialect adaptation, not evidence that the same results will apply to every Arabic dialect or deployment environment. NVIDIA specifically limits the workflow’s generalization beyond the reported Najdi and Hijazi experiment.

For an engineering team, the practical next step is to build an audited test set containing the dialect, speakers, noise, and recording conditions expected in use. That set should measure both the new dialect and the older capabilities the product still needs.

The durable lesson is controlled specialization: keep representative hard speech, replay the previous skills that matter, update as much of the model as the data supports, and choose decoding settings according to the required balance between accuracy and delay.