Open Source
Models, tools and corpora Misraj publishes openly, and where to find them.
Where we publish
Two accounts hold everything on this page: the Hugging Face organisation for weights, datasets and benchmark sets, and the GitHub organisation for code.
Models with a public artefact
Models that have something published to download or read: weights on Hugging Face, code on GitHub, or both. The rest of the catalogue is on the models page.
Kuwain
A 1.5B-parameter Arabic SLM built by language injection, the base layer for our downstream models.
Language1.5BOpen weights
Mutarjim
Bidirectional Arabic–English translation built on Kuwain, for general and specialized domains.
Language1.5BOpen weights
Sadeed
Arabic diacritization built on Kuwain, restoring full Tashkeel for MSA and classical Arabic.
Specialized NLP1.5BOpen weights
Baseer Nakba
Handwriting-specialised variant of Baseer
Vision-Language3BCommercial
Open datasets
Corpora released so others can train and evaluate on them. Where a set has no public repository yet it is listed without a link, rather than left out.
Open Arabic Corpus
35B+ curated Arabic tokens released publicly, growing regularly, so other researchers can build on the same data foundation rather than starting from scratch.
35B+ tokens
Wasm - Structured Arabic Multimodal Corpus
An open pipeline for constructing large-scale structured Arabic multimodal corpora from Common Crawl, preserving document structure in Markdown format rather than flattening it into noise.
MSDD - Misraj Structured Data Dump
A large-scale Arabic multimodal dataset created with the Wasm pipeline, extracted and filtered from Common Crawl dumps. It uniquely preserves the structural integrity of web content by emitting Markdown, addressing the shortage of high-quality structured multimodal data for Arabic.
MUDD - Misraj Unstructured Data Dump
A large-scale Arabic plain-text dataset of 4,758,338 rows, translated from SlimPajama-627B for pre-training Arabic language models. Each entry keeps its UUID from the original source.
4,758,338 rows
Sadeed Tashkeela
A large, high-quality Arabic diacritized corpus for training and evaluating Arabic diacritization models. The training split is built from a cleaned version of the Tashkeela corpus; the test split is a refined version of the Fadel Tashkeela test set.
Arabic Image Captioning (100M)
A large-scale Arabic multimodal dataset of 100 million Arabic image captions, generated with the Mutarjim translation model. It exists to close the gap in Arabic multimodal resources so Arabic vision-language systems can be trained at all.
100M captions
KITAB PDF-to-Markdown (Reviewed)
A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. Ground-truth errors - hallucinated text, missing page numbers, omitted small-font text - were fixed and formatting standardised, while the original task and schema are kept so it remains a drop-in evaluation set.
Open benchmarks
Evaluation sets published so a result can be reproduced rather than taken on trust. Each is a dataset on Hugging Face.
Misraj-DocOCR
A curated, expert-verified benchmark for Arabic document OCR built around structure preservation: 400 expertly verified pages, real and synthetic, covering books, reports, forms and complex layouts. Each page carries ground truth for both text fidelity and layout fidelity - Markdown and HTML tables, lists, footnotes, equations, watermarks, multi-column text and marginalia.
SadeedDiac-25
A benchmark for Arabic diacritization that unifies Modern Standard Arabic and Classical Arabic in a single dataset. Existing benchmarks cover one or the other, with limited domain diversity and inconsistent quality. SadeedDiac-25 spans news, religion, politics, sports and culinary text, is annotated through a multi-stage expert review, and is built to avoid contamination from large-scale pre-training corpora.
Tarjama-25
A bidirectional Arabic-English machine translation benchmark of 5,000 expert-reviewed sentence pairs. Sentences are long (roughly 50-100 tokens) and domain-balanced across scientific, technical, healthcare, cultural and general text, and half were originally written in Arabic and half in English before being professionally translated and verified in both directions.
Code repositories
Every repository in the Misraj GitHub organisation, in catalogue order. Licences are shown only where the host reports one it can identify, and a repository that belongs to a model links to that model's page.
kawn.ai Python SDK
The official Python client for the kawn.ai Models API. It exposes an embedding service and a Baseer OCR service, each with a synchronous and an asynchronous form, and reads its key from the MISRAJ_API_KEY environment variable.
PythonMITkawn.ai
LlamaIndex Kawn integration
LlamaIndex wrappers for the Kawn SDK: KawnEmbedding for document and query embeddings, and BaseerReader, which reads PDFs and images through the Baseer OCR service into LlamaIndex documents.
PythonMITllama-index-kawn
QuranHub API
A REST API over the Holy Quran covering multiple editions and languages: translations, tafsir, audio recitations, search, tajweed rules and the structural metadata of the mushaf. Built with FastAPI and PostgreSQL.
PLpgSQL
Sadeed
The code repository published alongside the Sadeed diacritization model and its paper.
Jupyter Notebook
Mutarjim evaluation
The evaluation repository published alongside the Mutarjim translation model and its paper.
Python
Nakba pipeline
Reproduces the result of the Misraj model in the Nakba-NLP 2026 shared task on Arabic handwritten text recognition.
Python
Kuwain Arabic cleaner
The Arabic data-cleaning repository published alongside the Kuwain 1.5B model and its paper.
Jupyter Notebook
Wasm
The pipeline repository published alongside the Wasm paper on constructing structured Arabic interleaved multimodal corpora.
Python
Papers with a public preprint or repository
The papers that can be read or rerun today. Papers still under review are on the publications page instead.
Kuwain 1.5B: An Arabic SLM via Language Injection
arXiv (2025), LaTell 20262025
Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model
arXiv (2025)2025
Sadeed: Advancing Arabic Diacritization Through Small Language Models
arXiv (2025)2025
Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
arXiv (2025), LaTell 20262025
Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
arXiv (2025), LaTell 20262025
What is not open
The 2-trillion-token Arabic pre-training corpus is proprietary and is not released. Several models are licensed openly in the catalogue but have no published weights yet; where that is so, only the artefact that exists is linked above. Nothing here stands in for something that does not.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.