Datasets
KITAB PDF-to-Markdown (Reviewed)
Open Source
A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. Ground-truth errors - hallucinated text, missing page numbers, omitted small-font text - were fixed and formatting standardised, while the original task and schema are kept so it remains a drop-in evaluation set.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.