logo
Datasets

Sadeed Tashkeela

Open Source

A large, high-quality Arabic diacritized corpus for training and evaluating Arabic diacritization models. The training split is built from a cleaned version of the Tashkeela corpus; the test split is a refined version of the Fadel Tashkeela test set.

huggingface.co

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.