logo

Kawn Lab

We build the models, datasets, benchmarks and training methods Arabic AI needs to exist at all.

Publications

Nakba-NLP 2026 @ LREC 2026
Arabic foundation models & open AI

Misraj AI at AR-MS NAKBA-NLP 2026: A State-of-the-Art VLM in Arabic Handwritten Text Recognition

Presents a specialized Arabic handwritten text recognition VLM that achieved 1st Place and established a new state-of-the-art on the Nakba OCR benchmark.

LaTell 2026
Arabic foundation models & open AI

Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation

Proposes a multi-teacher distillation framework for Arabic embedding models and introduces Kawn-Embed-Light, achieving state-of-the-art performance among Arabic embedding models under one billion parameters.

arXiv (2025), LaTell 2026
Arabic foundation models & open AI

Kuwain 1.5B: An Arabic SLM via Language Injection

Introduces a novel language injection method for building Arabic Small Language Models, enabling efficient Arabic specialization while preserving the original model's capabilities.

COLING Workshops (2025)
Arabic foundation models & open AI

Lahjawi: Arabic Cross-Dialect Translator

Presents the first family of models for Arabic dialect-to-dialect and dialect-to-MSA translation across 15 Arabic dialects.

arXiv (2025)
Arabic foundation models & open AI

Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

Introduces a compact Arabic-English translation model together with Tarjama-25, a comprehensive benchmark for Arabic-English machine translation.

arXiv (2025)
Arabic foundation models & open AI

Sadeed: Advancing Arabic Diacritization Through Small Language Models

Presents an Arabic diacritization model and SadeedDiac-25, a new benchmark for robust evaluation across diverse Arabic text genres.

arXiv (2025), LaTell 2026
Arabic foundation models & open AI

Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR

Introduces a specialized Arabic VLM for document OCR and Misraj-DocOCR, achieving state-of-the-art performance on Arabic document understanding.

arXiv (2025), LaTell 2026
Arabic foundation models & open AI

Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora

Introduces an open pipeline for constructing large-scale structured Arabic multimodal datasets from Common Crawl while preserving document structure in Markdown format.

Under review
Islamic text & acoustic intelligence

Annotating Real-World Quranic Recitations for Mispronunciation Detection and Diagnosis

Introduces a large-scale Tajweed-aware annotation methodology and corpus for Quranic Mispronunciation Detection and Diagnosis (MDD), enabling future pronunciation-aware Quranic ASR systems.

Under review
Islamic text & acoustic intelligence

ARISQU: A Dataset for Quotation Extraction and Attribution in Islamic Jurisprudential Issues

Introduces ARISQU, the first manually annotated dataset for extracting and attributing Islamic quotations, containing 8,800 quotations across 856,000 words with a fine-grained 14-tag annotation schema.

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.