Kawn Lab
We build the models, datasets, benchmarks and training methods Arabic AI needs to exist at all.
Publications
Misraj AI at AR-MS NAKBA-NLP 2026: A State-of-the-Art VLM in Arabic Handwritten Text Recognition
Presents a specialized Arabic handwritten text recognition VLM that achieved 1st Place and established a new state-of-the-art on the Nakba OCR benchmark.
Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation
Proposes a multi-teacher distillation framework for Arabic embedding models and introduces Kawn-Embed-Light, achieving state-of-the-art performance among Arabic embedding models under one billion parameters.
Kuwain 1.5B: An Arabic SLM via Language Injection
Introduces a novel language injection method for building Arabic Small Language Models, enabling efficient Arabic specialization while preserving the original model's capabilities.
Lahjawi: Arabic Cross-Dialect Translator
Presents the first family of models for Arabic dialect-to-dialect and dialect-to-MSA translation across 15 Arabic dialects.
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model
Introduces a compact Arabic-English translation model together with Tarjama-25, a comprehensive benchmark for Arabic-English machine translation.
Sadeed: Advancing Arabic Diacritization Through Small Language Models
Presents an Arabic diacritization model and SadeedDiac-25, a new benchmark for robust evaluation across diverse Arabic text genres.
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
Introduces a specialized Arabic VLM for document OCR and Misraj-DocOCR, achieving state-of-the-art performance on Arabic document understanding.
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
Introduces an open pipeline for constructing large-scale structured Arabic multimodal datasets from Common Crawl while preserving document structure in Markdown format.
Annotating Real-World Quranic Recitations for Mispronunciation Detection and Diagnosis
Introduces a large-scale Tajweed-aware annotation methodology and corpus for Quranic Mispronunciation Detection and Diagnosis (MDD), enabling future pronunciation-aware Quranic ASR systems.
ARISQU: A Dataset for Quotation Extraction and Attribution in Islamic Jurisprudential Issues
Introduces ARISQU, the first manually annotated dataset for extracting and attributing Islamic quotations, containing 8,800 quotations across 856,000 words with a fine-grained 14-tag annotation schema.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.