Datasets
Arabic Pre-training Corpus
2T tokens
A 2-trillion-token proprietary Arabic corpus - the largest clean dataset built specifically for Arabic LLM pre-training. It underlies Kawn Lab's foundation models and is not publicly released.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.