Skip to content
MISRAJAI
What we believe

Arabic-first by design. Built for the language, not translated into it.

Baseer OCRProduct

Turn Arabic Document ImagesInto Structured Markdown

Most OCR returns a flat run of characters, so the tables, headings and columns of an Arabic page are lost before anyone can use the text.

Turn Arabic Document Images Into Structured Markdown
The numbers behind it0.25Word error rate on Misraj-DocOCR
What it reads

Tables that stay tables

Table structure comes back as HTML inside the Markdown. On Misraj-DocOCR, Baseer scores 66 on TEDS, the table-structure metric, the highest of the systems compared in the paper.

0.25

Word error rate on Misraj-DocOCR

Headings and reading order

Headings, lists and multi-column pages come back in the order a person reads them, not as one flattened block.

66

TEDS table-structure score on Misraj-DocOCR

Arabic script as printed

Cursive letterforms, diacritics, varied fonts and right-to-left layout: the cases where general OCR tools lose accuracy on Arabic.

On the evidence

Both figures come from the Baseer paper (arXiv:2509.18174) and were measured on Misraj-DocOCR, a public benchmark of 400 expert-verified Arabic document images. Both are the best of the systems the paper compares; the nearest, Gemini 2.5 Pro, scores 0.37 and 52. The model in the paper is fine-tuned from Qwen2.5-VL-3B-Instruct. On the smaller corrected KITAB-Bench PDF-to-Markdown set, Baseer also leads on TEDS and MARS, while Dots.ocr and Nanonets post lower word error rates, and we would rather say so than leave it out.

Who it is for

One product. Many teams. One shared record.

Baseer OCR gives every team the same answer from the same source, from the people who run it to the people who sign off on it.

Archives and libraries

Scanned books and records become searchable text that keeps its headings and tables.

Publishers and education teams

Books, magazines and course material become Markdown that can be edited, republished or reused.

AI and search teams

Arabic PDFs reach retrieval pipelines and training sets as structured text rather than broken strings.

Developers

Send a page image and get Markdown back, through Kawn Console or as a node in Seamless API.

OCR that returns a string of characters has done half the job. Baseer OCR reads the page the way a person does and keeps what the layout says: which line is a heading, which cells belong to which row, and where one column ends and the next begins.

Misraj AIProduct team

See our customers
Integrations

Plugs into what you already run

10 systems, deployment modes and models named on this page, so Baseer OCR lands in the environment you already have.

View integrations
Kawn Console System
Seamless API System
Seamless Enterprise System
Baseer Extract System
Cloud Deployment
On-premises Deployment
Sovereign Deployment
Cloud API Deployment
Get started

Put Baseer OCR to work

Tell us the problem. We will tell you honestly whether Baseer OCR is the right answer.

Baseer OCR | Misraj AI