Anton Lozhkov

anton-l

AI & ML interests

Generative Models, Distributed Training, Photo and Video Enhancement

Recent Activity

Articles

Organizations

Hugging Face's profile picture 🧨Diffusers's profile picture Hugging Face Internal Testing Organization's profile picture superb's profile picture Anton's SUPERB Test Org's profile picture Util scripts for speech recognition's profile picture Internal Data & Models for Speech Recognition Event's profile picture Speech Recognition Community Event Version 2's profile picture OpenSLR's profile picture (De)fusing's profile picture HuggingFaceGECLM's profile picture BigCode's profile picture CompVis's profile picture CompVis Community's profile picture BigCode Data's profile picture Hugging Face TB Research's profile picture huggingPartyParis's profile picture HuggingFaceFW's profile picture Cosmopedia Stories Collab's profile picture StarCoder2 Data's profile picture Data Agents's profile picture Argilla Warehouse's profile picture smol-explorers's profile picture swissai-hf-data's profile picture Hugging Face Science's profile picture

Posts 1

view post
Post
1935
Introducing 📐𝐅𝐢𝐧𝐞𝐌𝐚𝐭𝐡: the best public math pre-training dataset with 50B+ tokens!
HuggingFaceTB/finemath

Math remains challenging for LLMs and by training on FineMath we see considerable gains over other math datasets, especially on GSM8K and MATH.

We build the dataset by:
🛠️ carefully extracting math data from Common Crawl;
🔎 iteratively filtering and recalling high quality math pages using a classifier trained on synthetic annotations to identify math reasoning and deduction.

We conducted a series of ablations comparing the performance of Llama-3.2-3B-Base after continued pre-training on FineMath and observe notable gains compared to the baseline model and other public math datasets.

We hope this helps advance the performance of LLMs on math and reasoning! 🚀
We’re also releasing all the ablation models as well as the evaluation code.

HuggingFaceTB/finemath-6763fb8f71b6439b653482c2