Hugging Face, in collaboration with Argilla, has unveiled the Data Is Better Together (DIBT) initiative, aimed at transforming how datasets are created for the multilingual machine learning community. This initiative includes the launch of the Multilingual Prompt Evaluation Project (MPEP), a venture that addresses the critical scarcity of multilingual datasets.
Key Developments
- The DIBT initiative began with the release of a dataset featuring 10,000 ranked prompts. This launch attracted participation from over 385 contributors within just days.
- The MPEP aims to create a leaderboard to evaluate prompts across multiple languages. To kick things off, 500 high-quality prompts have already been translated into languages including Dutch, Russian, and Spanish.
- Future community efforts will focus on the development of more domain-specific datasets, enabling engineers and domain specialists to collaborate effectively on relevant data resources.
- This initiative brings to light significant inequalities in dataset representation across languages, domains, and tasks, underscoring the necessity for detailed benchmarks.
- Engagement with the community will continue via Discord channels, where contributors can share datasets, results, and provide guidance for dataset construction.
The DIBT initiative is a strategic move to close gaps in multilingual dataset availability while fostering community collaboration. As the need for diverse datasets grows, this initiative stands to play a crucial role in shaping the future of open-source machine learning.




