Featured
Jais Corpus Pipeline
Distributed scraping + OCR/ASR pipeline feeding the Jais multilingual LLM.
Role
Data Scientist
Impact
Delivered 1T+ Arabic and 10T+ English tokens of clean training data.
Year
2024
End-to-end data engineering pipeline for the Jais Arabic LLM: distributed crawling, dedup, PII scrubbing, Azure OCR for scanned PDFs, Azure ASR for audio, and language-aware tokenization. Produced training-ready corpora at trillion-token scale.
Stack
Data Engineering · MLOps · Azure