Sara Hooker

sarahooker

AI & ML interests

None yet

Recent Activity

liked a model 9 days ago
CohereForAI/c4ai-command-r7b-12-2024
liked a dataset 9 days ago
CohereForAI/Global-MMLU-Lite
View all activity

Organizations

Cohere For AI's profile picture Massive Text Embedding Benchmark's profile picture C4AI Community's profile picture Hugging Face Party @ PyTorch Conference's profile picture

sarahooker's activity

reacted to dvilasuero's post with ❤️🔥 18 days ago
view post
Post
2261
🌐 Announcing Global-MMLU: an improved MMLU Open dataset with evaluation coverage across 42 languages, built with Argilla and the Hugging Face community.

Global-MMLU is the result of months of work with the goal of advancing Multilingual LLM evaluation. It's been an amazing open science effort with collaborators from Cohere For AI, Mila - Quebec Artificial Intelligence Institute, EPFL, Massachusetts Institute of Technology, AI Singapore, National University of Singapore, KAIST, Instituto Superior Técnico, Carnegie Mellon University, CONICET, and University of Buenos Aires.

🏷️ +200 contributors used Argilla MMLU questions where regional, dialect, or cultural knowledge was required to answer correctly. 85% of the questions required Western-centric knowledge!

Thanks to this annotation process, the open dataset contains two subsets:

1. 🗽 Culturally Agnostic: no specific regional, cultural knowledge is required.
2. ⚖️ Culturally Sensitive: requires dialect, cultural knowledge or geographic knowledge to answer correctly.

Moreover, we provide high quality translations of 25 out of 42 languages, thanks again to the community and professional annotators leveraging Argilla on the Hub.

I hope this will ensure a better understanding of the limitations and challenges for making open AI useful for many languages.

Dataset: CohereForAI/Global-MMLU
liked a Space 2 months ago
upvoted an article 4 months ago