Back to work
Featured

Jais Corpus Pipeline

Distributed scraping + OCR/ASR pipeline feeding the Jais multilingual LLM.

Role
Data Scientist
Impact
Delivered 1T+ Arabic and 10T+ English tokens of clean training data.
Year
2024

End-to-end data engineering pipeline for the Jais Arabic LLM: distributed crawling, dedup, PII scrubbing, Azure OCR for scanned PDFs, Azure ASR for audio, and language-aware tokenization. Produced training-ready corpora at trillion-token scale.

Stack
Data Engineering · MLOps · Azure