Datasets
Arabic Image Captioning (100M)
100M captions
Open Source
A large-scale Arabic multimodal dataset of 100 million Arabic image captions, generated with the Mutarjim translation model. It exists to close the gap in Arabic multimodal resources so Arabic vision-language systems can be trained at all.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.