Arabic paperwork piles up faster than teams can process it - manual data entry, error-prone re-keying, and growing backlogs of invoices, contracts and forms.
Baseer Extract is the enterprise product form of Baseer: OCR plus structured field extraction. Where Baseer turns a scan into clean Markdown, Baseer Extract turns it into data - invoice line items, contract clauses and dates, form fields, identity documents - delivered as structured output an ERP or workflow can consume without a human retyping it.
Document -> Baseer OCR (including the Nakba variant for handwritten text) -> field extraction -> structured output -> your system. Integrates with ERPs, Seamless API and Workforces.
How it works
Document in
A scan, a photograph or a PDF of the paperwork you already have.
Baseer OCR
The page is read structurally rather than as a flat string: columns, tables, headers and layout hierarchy survive the read. The Nakba variant covers handwritten Arabic.
Field extraction
The read becomes named fields: invoice line items, contract clauses and dates, form fields, identity documents.
Structured output
Delivered in a shape an ERP or a workflow consumes directly, rather than as text somebody has to interpret.
Your system
The receiving system takes the data without a human retyping it, which is the entire point.
What it extracts
Invoices
Line items and the dates and totals attached to them, returned as rows rather than as a picture of rows.
Contracts
Clauses and dates lifted out of the body of the document, so a contract becomes searchable at the level people actually search it.
Forms
Form fields returned as name and value pairs, including forms that were filled in by hand.
Identity documents
The fields an onboarding or verification step needs, read off the document instead of keyed in from it.
Handwritten Arabic
The Nakba variant handles handwritten Arabic, so a hand-filled form or a manuscript is in scope rather than an exception.
Layout that survives the read
Multi-column pages and embedded tables come back with their logical structure intact instead of scrambled into one block of text.
Where the output goes
ERP systems
Structured output is shaped for the system that receives it, which is the difference between extracting fields and extracting text.
Seamless API
Extraction is a node in a visual pipeline rather than a separate integration project.
Workforces
Document processing is one of the pre-built agent flows, so extraction feeds an operational process rather than a folder.
The named list, on request
This page names the classes of system the platform reaches. Which specific products, versions and authentication modes are in scope is confirmed against your estate during scoping, and that list is shared under NDA. A connector matrix published in the abstract is a claim nobody can hold us to, so there is not one here.
How it deploys
Cloud
Through Kawn Console, for teams that want the volume without running the model themselves.
On-premises
For organizations with data sovereignty requirements. The documents never leave your environment.
Sovereign
Dedicated in-Kingdom hosting, where on-premises is not practical but the jurisdiction still is.
The numbers behind it
- 0.25Word Error Rate on Misraj-DocOCR, an expert-verified Arabic OCR benchmark
- 3BParameter vision-language model, fine-tuned for Arabic
There is no published extraction accuracy per document type. The word error rate above is Baseer's, measured on Misraj-DocOCR and published with the benchmark; a field-level extraction benchmark is not something we have released, and we would rather say so than put a plausible second number beside a real one.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.