Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Blog

qalam: open-source Arabic PDF text extraction, without OCR

Misraj AI

Most PDF extractors return Arabic that is reversed or quietly corrupted. qalam, our open-source Rust and Python library, extracts Arabic text from digitally born PDFs in the right order and tells you which pages it could not read.

qalam: open-source Arabic PDF text extraction, without OCR

Take one line from a Saudi government guidance document, the disclaimer these documents carry: هذا الدليل إرشادي ولا يغني عن الرجوع إلى ("This guide is advisory and does not substitute for referring to..."). Ask pdfium, the engine inside Chrome's PDF viewer, to extract it, and you get إلى الرجوع عن يغني ولا إرشادي الدليل هذا. Every word is spelled correctly and every word is in the wrong place. Nothing raises an error.

We built qalam for that problem. It is an open-source library from Misraj for Arabic PDF text extraction. It reads digitally born PDFs, the ones created on a computer rather than scanned, and returns Arabic in logical reading order, as base letters, with ligatures and diacritics intact. When a page has no usable text layer, qalam says so. It does no OCR. The core is written in Rust, with Python bindings and a command-line tool, and it installs with pip install qalam.

If you build Arabic corpora, search indexes or RAG pipelines, this matters more than it first appears. Corrupted Arabic does not fail loudly. It tokenizes, it gets embedded, it gets indexed, and it never matches the text people actually search for.

Why Arabic PDF text extraction goes wrong

A PDF does not store text. It stores drawing instructions: use this font at this size, move to this coordinate, paint glyph #276. You get glyph numbers instead of letters and positions instead of lines. Extraction has to rebuild the text from those instructions, and three things it needs are optional in the format: the map from glyph back to Unicode (/ToUnicode), the logical structure, and the reading order, which an untagged file never stores at all.

For English, a naive extractor gets away with this most of the time. Arabic hits every weak point at once:

  • Glyphs are painted left to right, in visual order. Concatenate them in painting order and every word comes out reversed.
  • Many PDF producers store letters as Unicode presentation forms, the shaped variants of each letter, which share no codepoint with ordinary Arabic text.
  • A ligature such as lam-alef is one glyph that stands for two letters.
  • Diacritics are drawn inside the horizontal span of their base letter, so sorting by position can move them onto the wrong letter.
  • Columns sit side by side, and grouping text by baseline interleaves them line by line.
  • The /ToUnicode map can be missing or wrong, which means silent corruption or no text at all.

Before and after: real output from a government PDF

Every example below is real output from a real government PDF. The repository's scripts/compare.py reproduces them.

ProblemOther extractorsqalam
Reading orderإلى الرجوع عن يغني ولا إرشادي الدليل هذا
every word reversed
هذا الدليل إرشادي ولا يغني عن الرجوع إلى
logical order, ligature intact
Presentation formsﺍﻟﺪﻟﻴﻞ ﺍﻹﺭﺷﺎﺩﻱ
shares no codepoint with the real text
الدليل الإرشادي
base letters
Diacriticsتتضم َّن
a space in the middle of the word
تتضمَّن
marks attached to their base
Page with no text layer""
same as a blank page
page.verdict == "needs_ocr"
says what it could not read

The second row is the one people miss. To a reader the two strings look almost the same. To a computer they share no codepoint, so a tokenizer won't match them and a search index won't find them. In one 44-page document, the article behind the project counted 31,080 presentation-form characters in the output of a popular library.

The third row is easy to cause yourself. NFKC normalisation of an isolated form that carries a mark inserts a space before the mark, so a correct extraction followed by a naive normalisation step still ends up with a space inside the word.

How qalam fixes Arabic text extraction

qalam works on positioned glyphs, not on strings. Once text has been flattened into a string in the wrong order, the information needed to repair it (where each glyph sat on the page) is gone. So qalam reads the content stream, resolves each glyph code to Unicode, and only then builds lines.

Reorder before normalising

Most tools get the order of two operations wrong. NFKC normalisation folds presentation forms back to base letters, which is what you want. It also expands the lam-alef ligature ﻻ into two characters. Normalise first and reverse afterwards, and those two characters swap: ولا ("and not") becomes وال ("and the"). Both are real Arabic. The sentence still reads fluently, with one word changed.

qalam reorders while ligatures are still single glyphs, then normalises, so the ligature survives the reversal and expands in the right place. The same rule covers fonts that map one glyph to several letters; the article found fourteen such codes in a single file. qalam treats each of those mappings as one unit and never reverses inside it.

Marks, columns and tables

Diacritics stay attached to their base letter through reordering, so a shadda or a fatha stays on the letter it belongs to and no space appears inside the word.

For multi-column pages qalam uses a recursive XY-cut: find the widest empty band on the page, split there, and repeat inside each half. A full-width heading gets cut away before the search for the columns beneath it, and each column comes out whole, right to left. Gap thresholds scale with the local type size, and column direction is a property of the whole page, so a column of Latin figures can't flip the page's reading order.

Ruled tables come back as grids. Columns are numbered in reading order, so on an Arabic page rows[0][0] is the rightmost cell, where a reader starts, and table.to_rows() goes straight into csv.writer.

A verdict on every page

The failure we worry about most returns no error at all. A scanned page has no glyphs. A page whose fonts carry no usable Unicode map can't be decoded. Ask a mainstream extractor for either one and you get an empty string, the same answer it gives for a page that really is blank. In the 44-page test document from the article, four pages are unrecoverable, and pdfium and PyMuPDF both return empty strings for all four without mentioning it.

Run ten thousand documents through a pipeline like that and the scanned pages vanish. The job reports success, and nothing downstream can tell "this document didn't cover that topic" from "we lost page 12".

qalam judges every page instead:

VerdictMeaning
okEvery glyph resolved. The text can be trusted.
degradedText came out, but something is wrong. page.reasons says what.
needs_ocrNo usable text layer. Emitting text would be worse than useless.

Two choices in the scoring are deliberate. ok requires every glyph to resolve, because each unresolved glyph is a character missing from the output; a page where 99.9% of glyphs resolved is still degraded. And a page with no glyphs scores 0.0, not 1.0. "Nothing failed" is not the same as "everything worked", and a naive metric would hand a scanned page a perfect score.

We would rather a tool tell you it couldn't read page 42 than give you something that looks like page 42.

Quick start

pip install qalam
import qalam

text = qalam.extract_text("guide.pdf")

extract_text returns the whole document. Pages with no usable text layer contribute nothing, so a scanned page never turns up as text. For page-level detail, open a Document:

import qalam

doc = qalam.Document("guide.pdf")
print(doc.confidence)          # mean per-page confidence, 0.0 to 1.0
print(doc.pages_needing_ocr)   # e.g. [2, 3, 42, 43]

for page in doc:
    if page.verdict != "ok":
        print(page.number, page.verdict, page.reasons)

Each page is also an ordered list of typed blocks (paragraphs, tables and images), and doc.to_html() and doc.to_json() export the same result. The wheels are abi3, so one build covers Python 3.9 and later. The README documents the Rust API and the command-line tool.

Pairing qalam with Baseer OCR

qalam does not read scanned pages, by design. Real archives mix digital and scanned pages, sometimes inside one file, and the verdict splits the work cleanly. qalam extracts the pages that have a real text layer, and the page numbers in doc.pages_needing_ocr go to Baseer OCR, Misraj's product that reads scanned pages and photographs of Arabic documents and returns structured Markdown.

import qalam

doc = qalam.Document("archive.pdf")
text = doc.text                  # every page with a usable text layer
to_ocr = doc.pages_needing_ocr   # the page numbers to send to Baseer OCR

No page goes missing without you knowing. You take the text directly wherever the PDF has it, and you run OCR only on the pages that need it. If you are choosing an OCR model for the scanned part, our Arabic OCR benchmark results compare the options.

What qalam does not do

  • OCR. Scanned pages are detected and reported, not read.
  • PDF creation or editing.
  • Pixel-perfect layout. You get ordered, typed blocks, not a visual clone of the page.
  • Borderless tables. Only ruled tables are reconstructed for now.
  • Soft-mask compositing. A logo whose whole shape lives in its transparency mask extracts byte-correct but looks blank, and the image's dropped_transparency flag tells you so.

Help us find the PDFs that break it

qalam is young software, validated against a small corpus. Every rule in its pipeline exists because a real document proved the previous rule wrong. So the most useful contribution is an Arabic PDF that extracts badly: government publications, financial reports, theses, anything with tables. If you can't share the file, the output of qalam inspect yourfile.pdf still helps.

  • Found a bad extraction? Open an issue and attach the PDF if you can share it.
  • Not sure qalam fits your pipeline? Ask in GitHub Discussions.
  • Want to contribute code? Start with CONTRIBUTING.md, and read the findings log in PLAN.md, which records what each real document taught us.

The source is on GitHub, the package is on PyPI, and the full story of each failure is in the long-form article.

Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.