Datasets
MUDD - Misraj Unstructured Data Dump
4,758,338 rows
Open Source
A large-scale Arabic plain-text dataset of 4,758,338 rows, translated from SlimPajama-627B for pre-training Arabic language models. Each entry keeps its UUID from the original source.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.