← Research Research log Ongoing

Sibawayh

Efficient speech recognition for Moroccan Darija: how far a small encoder trained from scratch gets, what failed on the way, and the scaling question that needs compute.

Illustration: a robotic hand feeds a paper tape printed with a sound waveform into a brass mechanical ear; on the other side the tape emerges as three slightly different Arabic calligraphic writings of the same sound.
Lead
Elwalid Aboulaakoul
Independent ML Researcher, Asayl Labs
Engineering Student, ENSA Berrechid (Université Hassan 1er)
Contact
[email protected]

Thesis

We study how to build the most compute-efficient, high-quality speech recognition system for Moroccan Darija, across pretraining, post-training and decoding. The aim is not the largest model. It is the smallest system that reaches the practical accuracy frontier for Darija.

Why Darija needs its own research

  • Widely spoken, badly served. 91.9 % of Morocco's residents use Darija (HCP census 2024; several languages can be reported), about 34 million of 36.8 million people. On the Moroccan part of the public Casablanca benchmark, zero-shot Whisper-large-v3 scores 87 % WER / 44 % CER (83 % / 42 % with the paper's text normalisation); SeamlessM4T and MMS do worse (Talafha et al., EMNLP 2024).
  • Little open data. Openly downloadable transcribed Moroccan speech adds up to roughly 100–130 hours across the public sets we know of (our tally; licences vary).
  • Spelling is not fixed, so literal CER misleads. The same spoken word is written in several valid ways. In one early measurement, orthographic harmonisation alone moved WER by about 45 % relative. That is why we built sound-aware metrics.
  • The sounds that matter are hard. The emphatic consonants (ص ض ط ظ) are the hardest class for every model we tested. ص vs س is near chance in phonetic discrimination tests for every encoder, including Arabic-pretrained ones.
  • Speech is fast and mixed. Moroccan is the fastest-spoken dialect in Casablanca, and the one with the most French/English code-switched segments.
Bar chart of zero-shot word error rate on the Moroccan part of Casablanca: Whisper-large-v3 87.2 percent, Whisper-large-v2 88.55, SeamlessM4T-v2-large 95.18, MMS-1b-all 96.91.
Figure 1 General-purpose ASR on the Moroccan part of Casablanca, zero-shot (Talafha et al. 2024). Different test set from ours.
Log-scale bar chart of hours of speech: open transcribed Moroccan speech, the reference encoder's pretraining data, and our data used and collected.
Figure 2 Open transcribed Moroccan speech vs the reference encoder's pretraining data vs ours (used and collected). Log scale.

What we built

  1. Audio collection with provenance
  2. Evaluation system
  3. Self-supervised encoder, from scratch
  4. Matched fine-tuning harness
  5. Darija-aware error analysis
  6. Decoding research
  • Evaluation system: a Darija dev set of 94 clips from 32 channels with Gemini-derived references reviewed by the operator; a held-out test split reserved for final results; Darija-specific metrics and paired statistics.
  • Collection: about 40k clip-hours of Darija audio collected (40.8k unique by audio hash on 2026-09-29), with source-level provenance. Training eligibility is tracked separately from collection: some classes are held for evaluation only, enforced in code.
  • Transcription: about 6.6k hours machine-transcribed by at least one of two open teacher models (5.4k hours by both), with agreement tiers; speaker turns on part.
  • Encoder: BEST-RQ Conformer, 103.5M parameters, pretrained from scratch on Darija.
  • Harness: identical fine-tuning recipe for any encoder, so comparisons isolate the encoder.
  • Compute used so far: free-tier GPUs (T4) and small cloud credits, about 27 T4 GPU-hours for the current encoder's final training chain (about 45 including two abandoned restarts).

How we evaluate

ChoiceWhy
soundCERCER after folding spelling-only variants. Literal CER penalises valid Darija spellings.
normCERPunctuation and Arabic orthographic normalisation (hamza, ta-marbuta…). The profile is always stated.
Emphatic-collapse rateHow often ص ض ط ظ are written as their plain partners. It tracks the hardest sound class directly.
Equivalence registryManual, meaning-preserving spelling equivalences. Conflicts are rejected at load; raw text is never edited.
Paired bootstrap CIsItem-level paired bootstrap (2,000 resamples) for the results so far; channel-clustered CIs are the registered protocol for the scaling study. A result is "supported" only if the CI excludes 0 and the effect is ≥ 2× seed noise.
Effective n11 duplicate pairs mean the 94 dev clips are about 83 independent items. The CIs reported so far resample all 94; the scaling study counts duplicates once.
Single-seed resultsLabelled as such (CI-only), never as conclusions.
Versioned metricsAny methodology change gets a new version; old numbers must reproduce bit-for-bit.
Held-out testReserved for final results. Every result of ours on this page is on the dev set.

Results

The from-scratch Darija encoder keeps improving with steps

103.5M BEST-RQ Conformer, trained on 1.2k h of Darija (1.4k h manifest after duration filtering), about 370 s of audio per update. A frozen readout (a small probe on the frozen encoder), dev set, soundCER (lower is better):

StepssoundCERPaired change vs 23k (95 % CI)
11k0.465
23k0.421—
28k0.408−0.013 [−0.018, −0.008]
31k0.399−0.022 [−0.028, −0.017]

Still falling at the last point we could afford to measure.

Line chart of frozen-readout soundCER against pretraining steps, falling from 0.465 at 11k steps to 0.399 at 31k, with confidence intervals versus 23k.
Figure 3 Frozen-readout soundCER of the from-scratch encoder vs pretraining steps, with paired CIs vs 23k.

Matched comparison with a 3× larger Arabic-wide encoder

Both encoders get the identical CTC fine-tune: 100 h of general Darija, top 6 of the first 12 layers, same seed and schedule.

EncoderParamsPretrainingDev CERDev soundCEREmphatic collapse
Sibawayh encoder103.5M1.2k h Darija, ~27 T4-h0.3610.3590.109
Ara-BEST-RQ (Elleuch et al., ICASSP 2026)307M (first 12 of 24 layers used in the fine-tune)5.6k h Arabic dialects, 16×A1000.3320.3200.100
Grouped bar chart of CER, normCER and soundCER for the Sibawayh encoder and Ara-BEST-RQ under an identical 100 hour fine-tune.
Figure 4 Identical 100 h fine-tune: CER, normCER and soundCER for both encoders.
  • Gap: +0.039 soundCER, 95 % CI [+0.024, +0.054]. One seed per model.
  • The gap is concentrated, not uniform. 15 items carry +0.025 of the +0.039 net soundCER gap (66 %). They are noisy conversational stretches where our model drops words. Our encoder is better on about a quarter of items (25 of 94); outside the top 15 it is still behind by +0.013.
  • On the fine-tune's own dev split (5.2 h), our encoder ends lower: CER 0.269 vs 0.282. That split differs from the 94-clip dev set, and there is one seed each, so this is an observation, not a claim. Our pretraining corpus and the fine-tune data also share source families, which may favour our encoder on this split.
  • Deletions dominate for both models (about 20 % of reference characters).
What is and is not controlled

We control the fine-tune recipe, data, seed, layers used (both encoders cut to their first 12 layers) and scoring. Model size, pretraining data and pretraining compute differ, and those are exactly what the next study varies.

Left: scatter of per-item soundCER, ours versus the reference, with the 15 largest-gap items highlighted. Right: cumulative gap curve showing the top 15 items account for 66 percent of the net gap.
Figure 5 Left: per-item soundCER, ours vs reference. Right: cumulative gap. 15 items carry 66 % of the net gap.
Fine-tune dev CER per epoch for both encoders; the Sibawayh encoder ends lower on this split.
Figure 6 Fine-tune dev CER per epoch. Same recipe; our encoder ends lower on this split (single seed).

Clip edges cost words; real context recovers most of it

  • Effect: in both encoders, words near the cut edge of a segment are kept far less often. About 5 % of words lie within 1 s of an edge (about 10 % within 2 s), and the effect fades over about 2 s.
  • Test: re-decoding the same clips with ±2 s of the real surrounding audio (n = 52 items with source audio available).
    • soundCER 0.367 → 0.345: −0.022, 95 % CI [−0.030, −0.013].
    • Edge-word keep rate within 1 s: 0.11 → 0.18 with ±2 s of context (+6.7 pp, CI [+2.0, +11.5]) and 0.19 with ±4 s, against 0.24 mid-clip. About half (±2 s) to 60 % (±4 s) of the edge loss is recovered.
    • Padding with silence instead does nothing (+0.003, CI spans 0). Left-only or right-only context gives roughly a third to a half of the gain.
  • Implication: on the 52 items with source audio, segment-level evaluation understates the model by about 0.02 soundCER, and long-audio decoding should use overlapping windows with real context.
Exact-word keep rate by distance to the clip edge for the Sibawayh encoder: as cut, with 1 s of silence padding, and with 2 s of real surrounding audio.
Figure 7 Exact-word keep rate by distance to the clip edge: as cut, silence-padded, and with ±2 s of real audio. The under-0.3 s bucket contains only 88 words.
Change in soundCER for each context variant with 95 percent confidence intervals; only real audio context helps.
Figure 8 Change in soundCER for each context variant, with 95 % CIs (n = 52). Only real audio helps.

Hear it

Selected examples, including failures, from our own recordings and DODa. The measured results above are the evidence.

Short clips (trimmed and loudness-normalised only), greedy decoding, no language model. Reference on top, model output below. red underline = reference text the model missed or changed; red = wrong letters in the output; blue = extra letters. Text is shown without diacritics or punctuation, as scored.

Best DODa clip in this pool: 7 of 8 words exact, soundCER 0.02. (see Results)
Reference

كيفاش الناس كيبينو اخلاقهم والاحترام ديالهم فهاد البلاد

Sibawayh (ours)

كيفاش الناس كيبينو اخلاقهم والاحترام ديالهم في هاد البلاد

CER 0.04 | normCER 0.04 | soundCER 0.02
Source, licence, exposure

DODa audio dataset, AtlasIA (Hugging Face), MIT per dataset card; DODa text CC BY-NC 4.0. Speaker M2. Reference = the dataset's corrected Arabic-script text.

Licence: MIT per the Hugging Face dataset card; the upstream DODa text is CC BY-NC 4.0. Operator decision 2026-09-30: OK for a short credited clip on a non-commercial research page.

  • our pretraining pool: no (0 hash matches; registry: DODa audio yielded 0 rows to any training view because the panel leak gate reserved all 7 speakers)
  • our finetune data: no (no DODa origin among the 16 origins of the 100 h fine-tune; 0 hash matches)
  • arabest pretraining: no known overlap (DODa is not among the public sets listed in the paper's Table 2 and is not in the crawled set by construction); the dataset is public, so unrecorded overlap cannot be excluded
  • arabest finetune: no (same 100 h fine-tune data as ours)
Best operator clip in this pool: 7 of 13 words exact, soundCER 0.18. (see Results)
Reference

ماشي حيت تقيل راه السنسور ديال الباب بقى لاصق وكيحبس السينيال على البواطيي

Sibawayh (ours)

ماشي حيت تقيل راه السينسيور ديال الباب قلاسيق وكيحباس السينياة على ولبواط

CER 0.18 | normCER 0.16 | soundCER 0.18
edge_loss_endnegation
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
Middle tier: 2 of 13 words exact, soundCER 0.29; many outputs are close-sounding non-words. (see Results)
Reference

هاد البياس اللي ولاو كيجيبو دابا غير الشينوا ومكتكملش حتى سيمانة وكتلقاها تفرتتات

Sibawayh (ours)

هاد الكياس اللي ولا وكيجي بودا باشن واما كت كماشتاسيمعا وكتلقات فا لتا

CER 0.41 | normCER 0.41 | soundCER 0.29
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
Emphatic collapse: the model writes ض→د, ط→ت (an emphatic letter as its plain partner). 2 of 14 words exact, soundCER 0.27. (see Why Darija and Results)
Reference

سير جيب ليا الساروت ديال طرواسطاش وقطع الضو من الديزجونكتور الكبير عاد نرجعو هنا

Sibawayh (ours)

سيجيب لي السارود ديال طر واستاش وقطادو ما ديجونكتيور كبير العال جعو هنا

CER 0.28 | normCER 0.28 | soundCER 0.27
emphatic_collapse
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
Dropped stretch: 21 of 77 reference letters have no counterpart in the output (27%); 2 of 18 words exact. Deletions are the dominant error for both models. (see Results)
Reference

هاد لافيش راه مدرحة وسلعة عيانة النحاس لي لداخل داب غير داز فيه الضو والريحة ديال الشياط عاطية

Sibawayh (ours)

هاد الافيش راه من درحة وسايانة نحاس اللي داخلداب بغيد سي لدورحة ديالشاطا

CER 0.42 | normCER 0.39 | soundCER 0.39
deletionsedge_loss_endemphatic_loss
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
Negation: the reference has 3 negation words, the output 2. Literal CER does not show that a negation was dropped. 1 of 15 words exact. (see Results)
Reference

الضو ديال الصالون والو ما بغاش يشعل وحتى السنطراليزي ديال البيبان ما بقاش كايسد كاع

Sibawayh (ours)

الدوديال السارون وار ما غاش واحد شاء السونتراليزي ديال البيباه ماقاش كيسبي

CER 0.33 | normCER 0.31 | soundCER 0.33
emphatic_collapsenegationnegation_lost
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
A studio-read clip from the second source, above the pool median error: 4 of 11 words exact, soundCER 0.34. (see Results)
Reference

ولكن واش عندكم شي نشاطات اللي كتكونو فيهم مجموعين العاب مسلسلات

Sibawayh (ours)

ولكن مواش عندكم شيناساط ا اللي كونوا فيهم م المعيم علىاياب موصل سالات

CER 0.37 | normCER 0.35 | soundCER 0.34
Source, licence, exposure

DODa audio dataset, AtlasIA (Hugging Face), MIT per dataset card; DODa text CC BY-NC 4.0. Speaker M3. Reference = the dataset's corrected Arabic-script text.

Licence: MIT per the Hugging Face dataset card; the upstream DODa text is CC BY-NC 4.0. Operator decision 2026-09-30: OK for a short credited clip on a non-commercial research page.

  • our pretraining pool: no (0 hash matches; registry: DODa audio yielded 0 rows to any training view because the panel leak gate reserved all 7 speakers)
  • our finetune data: no (no DODa origin among the 16 origins of the 100 h fine-tune; 0 hash matches)
  • arabest pretraining: no known overlap (DODa is not among the public sets listed in the paper's Table 2 and is not in the crawled set by construction); the dataset is public, so unrecorded overlap cannot be excluded
  • arabest finetune: no (same 100 h fine-tune data as ours)
Code-switching: the reference writes the French words in Latin script, the model writes them phonetically in Arabic script. The literal score is inflated by the script mismatch, not only by wrong sounds. (see Why Darija)
Reference

من قبيلة و le panneau latéral واحل عريض كاع لي boutons ديال les paramètres خازنهم وما بغاش يتطوا

Sibawayh (ours)

ممقبيلة والطانلاطغال واح العريق كالي بالطون الي بام ا خا خازو وبغاش يطوى

CER 0.61 | normCER 0.59 | soundCER 0.77
code_switch
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
Same words, decoded twice: as the cut clip alone (soundCER 0.28) and inside a window with 2 s of the real neighbouring speech on each side, keeping only the cut's frames (soundCER 0.21). (see Results)
Clip as cutSame clip inside +-2 s of the real recording Reference for the cut span

ومخنق الدخان ما عندو منين يخرج وكيرجع لداخل هكا غادي يخنق بنادم يلا شعلتو الشوفو بلا ما

Decoded as cut

مخق الدخامة عن دنين يخر جو كيرجا لداخل هكا غادي يخنما بنهالمام ي واشع شو لشوفو لا

CER 0.32 | normCER 0.32 | soundCER 0.28
Decoded with context (only the cut span is scored)

مخق الدخامة عن دنين يخر جو تيرجا داخل هكا غادي يخما بنالم ي واشعت شو الشوفو بلا ما

CER 0.26 | normCER 0.26 | soundCER 0.21
clip_edge
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)
Same words, decoded twice: as the cut clip alone (soundCER 0.35) and inside a window with 2 s of the real neighbouring speech on each side, keeping only the cut's frames (soundCER 0.24). (see Results)
Clip as cutSame clip inside +-2 s of the real recording Reference for the cut span

والسيكيريتي ديال البيبان مخربقة ودابا حتى هاد الشاريو راه

Decoded as cut

سكورتي ديال فيبان خابفة و دابا حتى ها الشايو

CER 0.33 | normCER 0.33 | soundCER 0.35
Decoded with context (only the cut span is scored)

و سيكورتي ديال فيبان خربفة و دابا حتى ها الشايورا

CER 0.25 | normCER 0.25 | soundCER 0.24
clip_edge
Source, licence, exposure

Recording: Elwalid Aboulaakoul, Asayl Labs. The reference is the prompt text he read aloud (approved by the operator as a correct reading).

Licence: Own recording by the speaker, published by him; no third-party licence

  • our pretraining pool: no (0 hash matches in the pool, train, dev and sup manifests; on the eval-quarantine list)
  • our finetune data: no (same hash check; the 100 h fine-tune has no first-party origin)
  • arabest pretraining: no known overlap (Ara-BEST-RQ 300M pretrained on 5,640 h of crawled Creative Commons YouTube, arXiv:2603.21900 section 3; this recording was never published)
  • arabest finetune: no (same 100 h fine-tune data as ours)

Clip-edge check on all 5 interior cuts of the operator recordings (fixed rule, no picking): mean soundCER 0.312 as cut, 0.252 with 2 s of real context; better in 4, worse in 1. Small sample; the dev-set result is in section 5.3.

Ara-BEST comparison for these clips: coming after our next matched fine-tune.

What did not work

  1. Continued pretraining of an Arabic encoder on Darija: null or negative in four independent attempts. Pretraining loss improved; downstream quality did not (e.g. after a supervised fine-tune, dev WER 0.42 → 0.46 with the continued encoder; preliminary). This is why we pretrain from scratch.
  2. The reference learning rate, used from scratch: training loss looked healthy, but the encoder was worse than a random-init control on frozen readout. Half the rate works (3 replications). Lesson: always run a random-init control.
  3. Bigger batches: at equal audio seen, 1,600 s per update improved the readout less than 400 s (−0.007 vs −0.016 soundCER; single run, overlapping CIs). Here, more updates beat bigger updates.
  4. Distillation-style continuation (keeping the model close to its starting point): worse readout (+0.016 CER).
  5. Streaming (chunked) attention during pretraining: hurt the emphatic sounds. Turning it off cut emphatic errors about 17 % relative.
  6. Reading the final layer: the wrong instrument. The final layer was the worst or near-worst layer for every model we probed.
  7. N-best language-model reranking (earlier Arabic decoder work): did not beat the baseline. True shallow fusion gave small, real gains.
  8. Silence padding for clip edges: no effect (see Results).
  9. Our own first negation check: false alarms on 78–82 % of harmless fused or elided spelling edits. Rebuilt; v0.3 has 0 % on those edits (0.4 % on glued verbs).
  10. A withdrawn claim: an early analysis said the learned codes tracked recording channel more than phonetics. A later check showed it was a counting artefact, and we withdrew it.

Research log

  1. Arabic decoder research on a frozen acoustic model: normalisation dominates naive Arabic WER (~22 pts on MediaSpeech); wider beams show oracle headroom but cannot rank it; shallow LM fusion + morphological-legality rescoring give significant gains on three independent MSA sets, and 80 % of the rescoring gain over a fixed-vocabulary control traces to lexicon coverage.

  2. Pivot to Moroccan Darija. Existing systems benchmarked; recording and verification tools built.

  3. First Darija fine-tunes and an evaluation panel. A leakage audit found that a public benchmark overlaps a public encoder's training data, so those comparisons are reported as exposure-aware only.

  4. Evaluation v2: sound-aware metrics, equivalence registry, paired bootstrap, held-out-split guard. Representation analyses; continued pretraining found negative; final-layer readout found to be the wrong instrument.

  5. From-scratch BEST-RQ pretraining. Instrument errors caught by controls (learning rate, masking rate, batch size). Data engine to about 40k clip-hours. Step curve. Matched comparison. Clip-edge finding and context decoding.

Roadmap

Now (done)
Darija-aware evaluation · 100M encoder from scratch · matched reference comparison · edge-effect analysis
Next: scaling
Model size 100M / 200M / 300M × pretraining steps × unlabelled data 1.2k → ~8k → ~38k h
Next: post-training
Supervised data scaling · data mixtures · domain adaptation
Next: decoding
Context-aware decoding · phonetic and orthographic robustness · code-switch behaviour
Later: systems
Streaming · quantisation · efficient CPU inference · noisy/telephone speech

The compute request covers the whole program: the scaling study, long pretraining, post-training and decoding.

Protocol for the scaling study

  • One pretraining run per (size × data) cell, checkpointed at 25k, 50k, 100k and 200k steps.
  • Each checkpoint gets the same frozen readout and the same matched fine-tune (3 seeds), scored with paired CIs.
  • One pretraining-seed replicate.
  • The held-out test is used once, at the end.

The bottleneck is compute

The pipeline, evaluation and fine-tuning harness exist and already run on cloud GPUs; the matched fine-tunes run on Modal. The current encoder was trained on free-tier T4 time.

Phase I, the scaling study (about 700 H100-hours). Model size × pretraining steps × unlabelled data, with matched fine-tunes at every checkpoint. It is derived from measured throughput: 0.31 steps/s per T4 at 100M on arm a's recipe; the 300M rate is estimated at 3× slower (parameter ratio) and will be measured on the target hardware in the first hour. Each 200k-step cell needs about 20–70 H100-hours.

Next phases (about 850 H100-hours):

  • long pretraining of the best configuration to 1M steps on the full collected corpus, which is growing toward 100k hours;
  • supervised post-training scaling (100 h → 1k h → 7k h of labelled speech) and data-mixture and domain-adaptation runs;
  • decoding research: language models, context-aware decoding and correction models.

These estimates extrapolate from measured throughput and will be re-measured on the target hardware in the first hour.

Whole program: about 1,550 H100-hours.

Bar chart of requested H100-hours by block, A to G, about 1,550 H100-hours in total; the 100M rate is measured, and the 300M rate and later blocks are estimates.
Figure 9 Requested compute by block, with 20 % headroom. Phase I uses the measured T4 rate at 100M; the 300M rate and the later blocks are estimates, to be re-measured on H100.

Data and rights

  1. Collected, with source-level provenance
  2. Eligibility classified per source class
  3. Training set / evaluation-only set
  4. Release only with a clear, compatible rights basis
  • Source-level provenance is kept for all collected audio.
  • The collection contains several source classes with different rights conditions. Training eligibility is tracked separately from collection, and public-service broadcast audio is held for evaluation only.
  • No source audio or transcripts are redistributed.
  • Any public checkpoint will be trained only on data with a clear, compatible rights basis.
  • Release of model weights depends on that clearance. Results and methodology are published here.

About and references

Elwalid Aboulaakoul: Independent ML Researcher, Asayl Labs; Engineering Student, ENSA Berrechid (UH1).

  • Built the full Sibawayh pipeline: data collection, labelling, evaluation, pretraining and fine-tuning.
  • Earlier: Arabic ASR decoder research (see Research log), and independent computer-vision research on representation learning, with a manuscript under revision.

Contact: [email protected]

References

  • Chiu et al., Self-supervised Learning with Random-projection Quantizer for Speech Recognition (BEST-RQ), ICML 2022, arXiv:2202.01855
  • Elleuch et al., Ara-BEST-RQ, ICASSP 2026, arXiv:2603.21900
  • Talafha et al., Casablanca, EMNLP 2024, arXiv:2410.04527
  • HCP, RGPH 2024, Principaux résultats, hcp.ma