NVIDIA Nemotron 3.5 ASR (Automatic Speech Recognition) model is out of the box capable of transcribing up to 40 language locales, including Arabic, but in real-world deployments, it is often encountered with regional dialects and local recordings that are not adequately represented in the initial training data. This problem is particularly pronounced in Saudi Arabia, where both Modern Standard Arabic and specific dialects such as Najdi and Hijazi are spoken. In this article, we review how NVIDIA used the NeMo platform to perform precise fine-tuning to improve the model’s performance in these dialects, and what conclusions can be applied to other language groups.
Why is dialect adaptation necessary?

An automatic speech recognition system must recognize not only standard speech, but also how people actually speak. Regional dialects and local recording conditions are often underrepresented in the training data of large, multilingual models. As a result, a model, while performing well on broad benchmarks, may not be able to accurately transcribe local accents or specific phonetic features. In the case of Saudi Arabia, this means that a model may be able to recognize Modern Standard Arabic or English well, but has difficulty recognizing Najdi and Hijazi.
Fundamentals of the fine-tuning process

The NVIDIA team chose a „low-resource“ corpus, a limited but carefully prepared dataset containing 103,559 out of 125,490 recordings, totaling 133.7 hours (≈ 82.5 of the % original set). This filtering was not intended to remove hard-to-understand accents, but only structural flaws, such as poorly linked audio and text fragments. With the help of automatic quality assessments (UTMOS, SIGMOS), corpus-specific thresholds were set: UTMOS ≥ 1.25, SIGMOS noise ≥ 1.5, and overall SIGMOS ≥ 1.5, which allowed us to retain about 85 %-eligible recordings.
After data preparation, the model was adapted using the NVIDIA NeMo framework and the ASR fine-tuning recipe. The key steps were:
- Weighted replay mix – including old, well-learned language data in the training cycle so that fine-tuning does not eliminate already acquired knowledge.
- Efficient batching – optimized packet size, allowing efficient use of GPU memory.
- Partial unfreezing – loosening only some of the model layers to adapt to new data, while maintaining the basic knowledge base.
Results and insights
In the initial SADA (Saudi Arabic Dialect Adaptation) experiment, before fine-tuning, the Nemotron 3.5 ASR model achieved a WER (Word Error Rate) of 49.5 % on the validation set and a WER of 59 % on the full training corpus. After the first 10 epochs, the WER decreased to 47.8 %, and remained stable in subsequent epochs. This suggests that precise fine-tuning can provide significant performance improvements, although further experiments and hyperparameter tuning may be required to achieve larger improvements.
It is important to emphasize that this workflow is not a universal tool for all Arabic dialects or other language groups. The replay mechanism only protects data that is included in the training set, and partial layer loosening may require re-alignment when the data mix changes. Also, while this method works well for Saudi dialects, other regional environments may require different filtering thresholds or other data preparation measures.
Practical tips for other projects
- Select the target dialect – start with one clearly defined language variant, as was done with Najdi and Hijazi.
- Quality filtering – use automated tools (UTMOS, SIGMOS), but always check their distribution characteristics to avoid corrupting valuable data.
- Replay mix – include data from old languages to maintain the universality of the model.
- The simplest basic experiment – start with a full fine-tuning process with a single data set and a fixed evaluation set to measure changes.
Conclusions
The adaptation of the NVIDIA Nemotron 3.5 ASR model to Saudi Arabian dialects has shown that precise fine-tuning, based on careful data preparation and appropriate training methods, can improve speech recognition accuracy in regional environments. While this process will not be directly transferable to every language variant, its principles – data quality control, weighted replay, and partial layer relaxation – provide clear guidance for other researchers and companies seeking to adapt large speech models to specific markets.
CTA: If you want to learn how artificial intelligence can help your organization automate speech recognition, contact Krikis IT.






