Thesis
We study how to build the most compute-efficient, high-quality speech recognition system for Moroccan Darija, across pretraining, post-training and decoding. The aim is not the largest model. It is the smallest system that reaches the practical accuracy frontier for Darija.
Why Darija needs its own research
- Widely spoken, badly served. 91.9 % of Morocco's residents use Darija (HCP census 2024; several languages can be reported), about 34 million of 36.8 million people. On the Moroccan part of the public Casablanca benchmark, zero-shot Whisper-large-v3 scores 87 % WER / 44 % CER (83 % / 42 % with the paper's text normalisation); SeamlessM4T and MMS do worse (Talafha et al., EMNLP 2024).
- Little open data. Openly downloadable transcribed Moroccan speech adds up to roughly 100–130 hours across the public sets we know of (our tally; licences vary).
- Spelling is not fixed, so literal CER misleads. The same spoken word is written in several valid ways. In one early measurement, orthographic harmonisation alone moved WER by about 45 % relative. That is why we built sound-aware metrics.
- The sounds that matter are hard. The emphatic consonants (ص ض ط ظ) are the hardest class for every model we tested. ص vs س is near chance in phonetic discrimination tests for every encoder, including Arabic-pretrained ones.
- Speech is fast and mixed. Moroccan is the fastest-spoken dialect in Casablanca, and the one with the most French/English code-switched segments.
What we built
- Audio collection with provenance
- Evaluation system
- Self-supervised encoder, from scratch
- Matched fine-tuning harness
- Darija-aware error analysis
- Decoding research
- Evaluation system: a Darija dev set of 94 clips from 32 channels with Gemini-derived references reviewed by the operator; a held-out test split reserved for final results; Darija-specific metrics and paired statistics.
- Collection: about 40k clip-hours of Darija audio collected (40.8k unique by audio hash on 2026-09-29), with source-level provenance. Training eligibility is tracked separately from collection: some classes are held for evaluation only, enforced in code.
- Transcription: about 6.6k hours machine-transcribed by at least one of two open teacher models (5.4k hours by both), with agreement tiers; speaker turns on part.
- Encoder: BEST-RQ Conformer, 103.5M parameters, pretrained from scratch on Darija.
- Harness: identical fine-tuning recipe for any encoder, so comparisons isolate the encoder.
- Compute used so far: free-tier GPUs (T4) and small cloud credits, about 27 T4 GPU-hours for the current encoder's final training chain (about 45 including two abandoned restarts).
How we evaluate
| Choice | Why |
|---|---|
| soundCER | CER after folding spelling-only variants. Literal CER penalises valid Darija spellings. |
| normCER | Punctuation and Arabic orthographic normalisation (hamza, ta-marbuta…). The profile is always stated. |
| Emphatic-collapse rate | How often ص ض ط ظ are written as their plain partners. It tracks the hardest sound class directly. |
| Equivalence registry | Manual, meaning-preserving spelling equivalences. Conflicts are rejected at load; raw text is never edited. |
| Paired bootstrap CIs | Item-level paired bootstrap (2,000 resamples) for the results so far; channel-clustered CIs are the registered protocol for the scaling study. A result is "supported" only if the CI excludes 0 and the effect is ≥ 2× seed noise. |
| Effective n | 11 duplicate pairs mean the 94 dev clips are about 83 independent items. The CIs reported so far resample all 94; the scaling study counts duplicates once. |
| Single-seed results | Labelled as such (CI-only), never as conclusions. |
| Versioned metrics | Any methodology change gets a new version; old numbers must reproduce bit-for-bit. |
| Held-out test | Reserved for final results. Every result of ours on this page is on the dev set. |
Results
The from-scratch Darija encoder keeps improving with steps
103.5M BEST-RQ Conformer, trained on 1.2k h of Darija (1.4k h manifest after duration filtering), about 370 s of audio per update. A frozen readout (a small probe on the frozen encoder), dev set, soundCER (lower is better):
| Steps | soundCER | Paired change vs 23k (95 % CI) |
|---|---|---|
| 11k | 0.465 | |
| 23k | 0.421 | — |
| 28k | 0.408 | −0.013 [−0.018, −0.008] |
| 31k | 0.399 | −0.022 [−0.028, −0.017] |
Still falling at the last point we could afford to measure.
Matched comparison with a 3× larger Arabic-wide encoder
Both encoders get the identical CTC fine-tune: 100 h of general Darija, top 6 of the first 12 layers, same seed and schedule.
| Encoder | Params | Pretraining | Dev CER | Dev soundCER | Emphatic collapse |
|---|---|---|---|---|---|
| Sibawayh encoder | 103.5M | 1.2k h Darija, ~27 T4-h | 0.361 | 0.359 | 0.109 |
| Ara-BEST-RQ (Elleuch et al., ICASSP 2026) | 307M (first 12 of 24 layers used in the fine-tune) | 5.6k h Arabic dialects, 16×A100 | 0.332 | 0.320 | 0.100 |
- Gap: +0.039 soundCER, 95 % CI [+0.024, +0.054]. One seed per model.
- The gap is concentrated, not uniform. 15 items carry +0.025 of the +0.039 net soundCER gap (66 %). They are noisy conversational stretches where our model drops words. Our encoder is better on about a quarter of items (25 of 94); outside the top 15 it is still behind by +0.013.
- On the fine-tune's own dev split (5.2 h), our encoder ends lower: CER 0.269 vs 0.282. That split differs from the 94-clip dev set, and there is one seed each, so this is an observation, not a claim. Our pretraining corpus and the fine-tune data also share source families, which may favour our encoder on this split.
- Deletions dominate for both models (about 20 % of reference characters).
We control the fine-tune recipe, data, seed, layers used (both encoders cut to their first 12 layers) and scoring. Model size, pretraining data and pretraining compute differ, and those are exactly what the next study varies.
Clip edges cost words; real context recovers most of it
- Effect: in both encoders, words near the cut edge of a segment are kept far less often. About 5 % of words lie within 1 s of an edge (about 10 % within 2 s), and the effect fades over about 2 s.
-
Test: re-decoding the same clips with ±2 s of the real surrounding audio
(n = 52 items with source audio available).
- soundCER 0.367 → 0.345: −0.022, 95 % CI [−0.030, −0.013].
- Edge-word keep rate within 1 s: 0.11 → 0.18 with ±2 s of context (+6.7 pp, CI [+2.0, +11.5]) and 0.19 with ±4 s, against 0.24 mid-clip. About half (±2 s) to 60 % (±4 s) of the edge loss is recovered.
- Padding with silence instead does nothing (+0.003, CI spans 0). Left-only or right-only context gives roughly a third to a half of the gain.
- Implication: on the 52 items with source audio, segment-level evaluation understates the model by about 0.02 soundCER, and long-audio decoding should use overlapping windows with real context.
Hear it
Selected examples, including failures, from our own recordings and DODa. The measured results above are the evidence.
Short clips (trimmed and loudness-normalised only), greedy decoding, no language model. Reference on top, model output below. red underline = reference text the model missed or changed; red = wrong letters in the output; blue = extra letters. Text is shown without diacritics or punctuation, as scored.
كيفاش الناس كيبينو اخلاقهم والاحترام ديالهم فهاد البلاد
Sibawayh (ours)كيفاش الناس كيبينو اخلاقهم والاحترام ديالهم في هاد البلاد
Source, licence, exposure
DODa audio dataset, AtlasIA (Hugging Face), MIT per dataset card; DODa text CC BY-NC 4.0. Speaker M2. Reference = the dataset's corrected Arabic-script text.
Licence: MIT per the Hugging Face dataset card; the upstream DODa text is CC BY-NC 4.0. Operator decision 2026-09-30: OK for a short credited clip on a non-commercial research page.
- our pretraining pool: no (0 hash matches; registry: DODa audio yielded 0 rows to any training view because the panel leak gate reserved all 7 speakers)
- our finetune data: no (no DODa origin among the 16 origins of the 100 h fine-tune; 0 hash matches)
- arabest pretraining: no known overlap (DODa is not among the public sets listed in the paper's Table 2 and is not in the crawled set by construction); the dataset is public, so unrecorded overlap cannot be excluded
- arabest finetune: no (same 100 h fine-tune data as ours)
ماشي حيت تقيل راه السنسور ديال الباب بقى لاصق وكيحبس السينيال على البواطيي
Sibawayh (ours)ماشي حيت تقيل راه السينسيور ديال الباب قلاسيق وكيحباس السينياة على ولبواط
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
هاد البياس اللي ولاو كيجيبو دابا غير الشينوا ومكتكملش حتى سيمانة وكتلقاها تفرتتات
Sibawayh (ours)هاد الكياس اللي ولا وكيجي بودا باشن واما كت كماشتاسيمعا وكتلقات فا لتا
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
سير جيب ليا الساروت ديال طرواسطاش وقطع الضو من الديزجونكتور الكبير عاد نرجعو هنا
Sibawayh (ours)سيجيب لي السارود ديال طر واستاش وقطادو ما ديجونكتيور كبير العال جعو هنا
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
هاد لافيش راه مدرحة وسلعة عيانة النحاس لي لداخل داب غير داز فيه الضو والريحة ديال الشياط عاطية
Sibawayh (ours)هاد الافيش راه من درحة وسايانة نحاس اللي داخلداب بغيد سي لدورحة ديالشاطا
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
الضو ديال الصالون والو ما بغاش يشعل وحتى السنطراليزي ديال البيبان ما بقاش كايسد كاع
Sibawayh (ours)الدوديال السارون وار ما غاش واحد شاء السونتراليزي ديال البيباه ماقاش كيسبي
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
ولكن واش عندكم شي نشاطات اللي كتكونو فيهم مجموعين العاب مسلسلات
Sibawayh (ours)ولكن مواش عندكم شيناساط ا اللي كونوا فيهم م المعيم علىاياب موصل سالات
Source, licence, exposure
DODa audio dataset, AtlasIA (Hugging Face), MIT per dataset card; DODa text CC BY-NC 4.0. Speaker M3. Reference = the dataset's corrected Arabic-script text.
Licence: MIT per the Hugging Face dataset card; the upstream DODa text is CC BY-NC 4.0. Operator decision 2026-09-30: OK for a short credited clip on a non-commercial research page.
- our pretraining pool: no (0 hash matches; registry: DODa audio yielded 0 rows to any training view because the panel leak gate reserved all 7 speakers)
- our finetune data: no (no DODa origin among the 16 origins of the 100 h fine-tune; 0 hash matches)
- arabest pretraining: no known overlap (DODa is not among the public sets listed in the paper's Table 2 and is not in the crawled set by construction); the dataset is public, so unrecorded overlap cannot be excluded
- arabest finetune: no (same 100 h fine-tune data as ours)
من قبيلة و le panneau latéral واحل عريض كاع لي boutons ديال les paramètres خازنهم وما بغاش يتطوا
Sibawayh (ours)ممقبيلة والطانلاطغال واح العريق كالي بالطون الي بام ا خا خازو وبغاش يطوى
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
ومخنق الدخان ما عندو منين يخرج وكيرجع لداخل هكا غادي يخنق بنادم يلا شعلتو الشوفو بلا ما
Decoded as cutمخق الدخامة عن دنين يخر جو كيرجا لداخل هكا غادي يخنما بنهالمام ي واشع شو لشوفو لا
مخق الدخامة عن دنين يخر جو تيرجا داخل هكا غادي يخما بنالم ي واشعت شو الشوفو بلا ما
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
والسيكيريتي ديال البيبان مخربقة ودابا حتى هاد الشاريو راه
Decoded as cutسكورتي ديال فيبان خابفة و دابا حتى ها الشايو
و سيكورتي ديال فيبان خربفة و دابا حتى ها الشايورا
Source, licence, exposure
Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).
Licence: Own recording by the speaker, published by him; no third-party licence
- our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
- our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
- arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
- arabest finetune: no (same 100 h fine-tune data as ours)
Clip-edge check on all 5 interior cuts of the operator recordings (fixed rule, no picking): mean soundCER 0.312 as cut, 0.252 with 2 s of real context; better in 4, worse in 1. Small sample; the dev-set result is in section 5.3.
Ara-BEST comparison for these clips: coming after our next matched fine-tune.
What did not work
- Continued pretraining of an Arabic encoder on Darija: null or negative in four independent attempts. Pretraining loss improved; downstream quality did not (e.g. after a supervised fine-tune, dev WER 0.42 → 0.46 with the continued encoder; preliminary). This is why we pretrain from scratch.
- The reference learning rate, used from scratch: training loss looked healthy, but the encoder was worse than a random-init control on frozen readout. Half the rate works (3 replications). Lesson: always run a random-init control.
- Bigger batches: at equal audio seen, 1,600 s per update improved the readout less than 400 s (−0.007 vs −0.016 soundCER; single run, overlapping CIs). Here, more updates beat bigger updates.
- Distillation-style continuation (keeping the model close to its starting point): worse readout (+0.016 CER).
- Streaming (chunked) attention during pretraining: hurt the emphatic sounds. Turning it off cut emphatic errors about 17 % relative.
- Reading the final layer: the wrong instrument. The final layer was the worst or near-worst layer for every model we probed.
- N-best language-model reranking (earlier Arabic decoder work): did not beat the baseline. True shallow fusion gave small, real gains.
- Silence padding for clip edges: no effect (see Results).
- Our own first negation check: false alarms on 78–82 % of harmless fused or elided spelling edits. Rebuilt; v0.3 has 0 % on those edits (0.4 % on glued verbs).
- A withdrawn claim: an early analysis said the learned codes tracked recording channel more than phonetics. A later check showed it was a counting artefact, and we withdrew it.
Research log
-
Arabic decoder research on a frozen acoustic model: normalisation dominates naive Arabic WER (~22 pts on MediaSpeech); wider beams show oracle headroom but cannot rank it; shallow LM fusion + morphological-legality rescoring give significant gains on three independent MSA sets, and 80 % of the rescoring gain over a fixed-vocabulary control traces to lexicon coverage.
-
Pivot to Moroccan Darija. Existing systems benchmarked; recording and verification tools built.
-
First Darija fine-tunes and an evaluation panel. A leakage audit found that a public benchmark overlaps a public encoder's training data, so those comparisons are reported as exposure-aware only.
-
Evaluation v2: sound-aware metrics, equivalence registry, paired bootstrap, held-out-split guard. Representation analyses; continued pretraining found negative; final-layer readout found to be the wrong instrument.
-
From-scratch BEST-RQ pretraining. Instrument errors caught by controls (learning rate, masking rate, batch size). Data engine to about 40k clip-hours. Step curve. Matched comparison. Clip-edge finding and context decoding.
-
Scaling study (see Roadmap).
Roadmap
- Now (done)
- Darija-aware evaluation · 100M encoder from scratch · matched reference comparison · edge-effect analysis
- Next: scaling
- Model size 100M / 200M / 300M × pretraining steps × unlabelled data 1.2k → ~8k → ~38k h
- Next: post-training
- Supervised data scaling · data mixtures · domain adaptation
- Next: decoding
- Context-aware decoding · phonetic and orthographic robustness · code-switch behaviour
- Later: systems
- Streaming · quantisation · efficient CPU inference · noisy/telephone speech
The compute request covers the whole program: the scaling study, long pretraining, post-training and decoding.
Protocol for the scaling study
- One pretraining run per (size × data) cell, checkpointed at 25k, 50k, 100k and 200k steps.
- Each checkpoint gets the same frozen readout and the same matched fine-tune (3 seeds), scored with paired CIs.
- One pretraining-seed replicate.
- The held-out test is used once, at the end.
The bottleneck is compute
The pipeline, evaluation and fine-tuning harness exist and already run on cloud GPUs; the matched fine-tunes run on Modal. The current encoder was trained on free-tier T4 time.
Phase I, the scaling study (about 700 H100-hours). Model size × pretraining steps × unlabelled data, with matched fine-tunes at every checkpoint. It is derived from measured throughput: 0.31 steps/s per T4 at 100M on arm a's recipe; the 300M rate is estimated at 3× slower (parameter ratio) and will be measured on the target hardware in the first hour. Each 200k-step cell needs about 20–70 H100-hours.
Next phases (about 850 H100-hours):
- long pretraining of the best configuration to 1M steps on the full collected corpus, which is growing toward 100k hours;
- supervised post-training scaling (100 h → 1k h → 7k h of labelled speech) and data-mixture and domain-adaptation runs;
- decoding research: language models, context-aware decoding and correction models.
These estimates extrapolate from measured throughput and will be re-measured on the target hardware in the first hour.
Whole program: about 1,550 H100-hours.
Data and rights
- Collected, with source-level provenance
- Eligibility classified per source class
- Training set / evaluation-only set
- Release only with a clear, compatible rights basis
- Source-level provenance is kept for all collected audio.
- The collection contains several source classes with different rights conditions. Training eligibility is tracked separately from collection, and public-service broadcast audio is held for evaluation only.
- No source audio or transcripts are redistributed.
- Any public checkpoint will be trained only on data with a clear, compatible rights basis.
- Release of model weights depends on that clearance. Results and methodology are published here.
About and references
Elwalid Aboulaakoul: Independent ML Researcher, Asayl Labs; Engineering Student, ENSA Berrechid (UH1).
- Built the full Sibawayh pipeline: data collection, labelling, evaluation, pretraining and fine-tuning.
- Earlier: Arabic ASR decoder research (see Research log), and independent computer-vision research on representation learning, with a manuscript under revision.
Contact: [email protected]
References
- Chiu et al., Self-supervised Learning with Random-projection Quantizer for Speech Recognition (BEST-RQ), ICML 2022, arXiv:2202.01855
- Elleuch et al., Ara-BEST-RQ, ICASSP 2026, arXiv:2603.21900
- Talafha et al., Casablanca, EMNLP 2024, arXiv:2410.04527
- HCP, RGPH 2024, Principaux résultats, hcp.ma