Cohere For AI

community

https://cohere.for.ai/

CohereForAI

for-ai

Activity Feed Request to join this org

AI & ML interests

None defined yet.

Recent Activity

shivalikasingh updated a dataset about 6 hours ago

CohereForAI/Global-MMLU

shivalikasingh updated a dataset about 6 hours ago

CohereForAI/m-ArenaHard

clefourrier authored a paper 25 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

View all activity

Articles

A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality

Oct 24, 2024

• 59

Putting RL back in RLHF

Jun 12, 2024

• 78

CohereForAI's activity

shivalikasingh

updated 2 datasets about 6 hours ago

CohereForAI/Global-MMLU

Viewer • Updated about 6 hours ago • 487k • 14.2k • 107

CohereForAI/m-ArenaHard

Viewer • Updated about 6 hours ago • 10.5k • 660 • 17

davanstrien

posted an update 3 days ago

Post

2413

📊 Introducing "Hugging Face Dataset Spotlight" 📊

I'm excited to share the first episode of our AI-generated podcast series focusing on nice datasets from the Hugging Face Hub!

This first episode explores mathematical reasoning datasets:

- SynthLabsAI/Big-Math-RL-Verified: Over 250,000 rigorously verified problems spanning multiple difficulty levels and mathematical domains
- open-r1/OpenR1-Math-220k: 220,000 math problems with multiple reasoning traces, verified for accuracy using Math Verify and Llama-3.3-70B models.
- facebook/natural_reasoning: 1.1 million general reasoning questions carefully deduplicated and decontaminated from existing benchmarks, showing superior scaling effects when training models like Llama3.1-8B-Instruct.

Plus a bonus segment on bespokelabs/bespoke-manim!

https://www.youtube.com/watch?v=-TgmRq45tW4

davanstrien

posted an update 4 days ago

Post

3462

Quick POC: Turn a Hugging Face dataset card into a short podcast introducing the dataset using all open models.

I think I'm the only weirdo who would enjoy listening to something like this though 😅

Here is an example for eth-nlped/stepverify

2 replies

davanstrien

posted an update 11 days ago

Post

2540

Hacked together a way to log trl GRPO training completions to a 🤗 dataset repo. This allows you to:

- Track rewards from multiple reward functions
- Treat the completion and rewards from training as a "proper" dataset and do EDA
- Share results for open science

The implementation is super hacky, but I'm curious if people would find this useful.

To push completions to the Hub, you just need two extra parameters:

log_completions=True
log_completions_hub_repo='your-username/repo-name'

Example dataset: davanstrien/test-logs
Colab: https://colab.research.google.com/drive/1wzBFPVthRYYTp-mEYlznLg_e_0Za1M3g

davanstrien

posted an update 15 days ago

Post

2212

Dataset descriptions for trending Hugging Face datasets? Powered by a Smol model davanstrien/Smol-Hub-tldr

davanstrien

posted an update 17 days ago

Post

1881

How do you make 1M+ Hugging Face models & datasets more discoverable?

davanstrien/Smol-Hub-tldr!

I fine-tuned HuggingFaceTB/SmolLM2-360M to generate one-line summaries from a model or dataset README.

Its own self-description?
"A model for generating concise summaries of model & dataset cards from the Hugging Face Hub"

The goal? Make it easier to find the right models and datasets for your specific needs. It's already powering a semantic search for datasets Space.

It's still a WIP but thanks to @loubnabnl , @anton-l , @eliebak et al, for cooking such a nice base model for fine-tuning small, efficient models for specific domains and tasks. 🙏

davanstrien

posted an update 18 days ago

Post

1347

Made some significant updates to my 🤗 semantic datasets search app. If you love falling into a wiki black hole, you might like this...

https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search

clefourrier

authored a paper 25 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 27 days ago • 196

davanstrien

posted an update about 1 month ago

Post

1827

Why choose between strong LLM reasoning and efficient models?

Use DeepSeek to generate high-quality training data, then distil that knowledge into ModernBERT answerdotai/ModernBERT-base for fast, efficient classification.

Blog post: https://danielvanstrien.xyz/posts/2025/deepseek/distil-deepseek-modernbert.html

davanstrien

posted an update about 1 month ago

Post

1919

Updated the ColPali Query Generator Space davanstrien/ColPali-Query-Generator to use Qwen/Qwen2.5-VL-7B-Instruct.

Given an input image, it generates several queries along with explanations to justify them. This approach can generate synthetic data for fine-tuning ColPali models.

davanstrien

posted an update about 1 month ago

Post

2035

🌍 Big step for multilingual AI data!

The Hugging Face community has rated educational content in languages spoken by 1.6 billion people! New additions:
• Japanese
• Italian
• Old High German

Learn more and contribute: https://huggingface.co/blog/davanstrien/fineweb2-community

These ratings can help enhance training data for major world languages.

1 reply

sarahooker

authored 6 papers about 1 month ago

Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress

Paper • 2408.14960 • Published Aug 27, 2024

Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning

Paper • 2410.10801 • Published Oct 14, 2024

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

Paper • 2411.19799 • Published Nov 29, 2024 • 11

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Paper • 2412.03304 • Published Dec 4, 2024 • 18

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

Paper • 2406.03368 • Published Jun 5, 2024

Bridging the Data Provenance Gap Across Text, Speech and Video

Paper • 2412.17847 • Published Dec 19, 2024 • 9

ariG23498

posted an update about 1 month ago

Post

2158

Tried my hand at simplifying the derivations of Direct Preference Optimization.

I cover how one can reformulate RLHF into DPO. The idea of implicit reward modeling is chef's kiss.

Blog: https://huggingface.co/blog/ariG23498/rlhf-to-dpo

ariG23498

posted an update about 2 months ago

Post

1936

Timm ❤️ Transformers

Wtih the latest version of transformers you can now use any timm model with the familiar transformers API.

Blog Post: https://huggingface.co/blog/timm-transformers
Repository with examples: https://github.com/ariG23498/timm-wrapper-examples
Collection: ariG23498/timmwrapper-6777b85f1e8d085d3f1374a1

AI & ML interests

Recent Activity

Articles

A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality

Putting RL back in RLHF

Team members 53

CohereForAI's activity