logo
Datasets

KITAB PDF-to-Markdown (Reviewed)

Open Source

A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. Ground-truth errors - hallucinated text, missing page numbers, omitted small-font text - were fixed and formatting standardised, while the original task and schema are kept so it remains a drop-in evaluation set.

huggingface.co

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.