logo
Datasets

Arabic Image Captioning (100M)

100M captions
Open Source

A large-scale Arabic multimodal dataset of 100 million Arabic image captions, generated with the Mutarjim translation model. It exists to close the gap in Arabic multimodal resources so Arabic vision-language systems can be trained at all.

huggingface.co

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.

Arabic Image Captioning (100M) | Misraj AI | Next-Gen Arabic AI Lab