logo
Vision-Language
3B
Commercial

Baseer

Arabic document intelligence: scans to structured Markdown

Baseer is an Arabic document intelligence engine built on a 3-billion-parameter vision-language model fine-tuned specifically for Arabic. Unlike standard OCR tools that produce raw text strings, Baseer reads documents structurally: it understands columns, tables, headers, and layout hierarchy, then outputs clean Markdown that is usable downstream. It is trained on a large-scale dataset combining synthetic and real-world Arabic documents, using a decoder-only fine-tuning strategy that adapts a pre-trained multimodal model while preserving its visual understanding. Benchmarked on Misraj-DocOCR, an expert-verified Arabic OCR benchmark, Baseer achieves a Word Error Rate of 0.25.

Benchmarks

DatasetMetricResult
Misraj-DocOCRArabic document OCR had no expert-verified benchmark that measured structure, not just characters. Misraj-DocOCR is the evaluation Baseer was measured against, and it is public, so anyone can reproduce the comparison rather than take the number on trust.
Word Error Rate
0.25
State of the art

Access

Cloud API
On-premises

Variants

Baseer Nakba

Handwriting-specialised variant of Baseer

Baseer Nakba is an open-source vision-language model for Arabic handwritten text recognition, optimized for historical manuscripts. It accurately transcribes highly cursive handwritten Arabic from degraded archival documents. It achieved 1st Place at the AR-MS NAKBA-NLP 2026 Arabic Manuscript Understanding Shared Task, establishing a new state of the art on the Nakba OCR benchmark. Open weights are published on Hugging Face.

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.