HuggingFaceTB
/

SmolLM2-360M-Instruct

Text Generation

Transformers.js

text-generation-inference

Inference Endpoints

Model card Files Files and versions Metrics Training metrics Community

loubnabnl HF staff commited on Oct 31, 2024

Commit

502b5cb

·

verified ·

1 Parent(s): 51b1c73

Update README.md

Files changed (1) hide show

README.md +2 -2

README.md CHANGED Viewed

@@ -22,9 +22,9 @@ language:
 SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device.
-SmolLM2 demonstrates significant advances over its predecessor SmolLM1, particularly in instruction following, knowledge, reasoning. The 135M model was trained on 2 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new filtered datasets we curated and will release soon.
-We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our own curated datasets, designed to enhance instruction following, rewriting, and summarization capabilities. We then applied Direct Preference Optimization (DPO) using a mix of UltraFeedback and DPO-ORPO.
 ### How to use

 SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device.
+SmolLM2 demonstrates significant advances over its predecessor SmolLM1, particularly in instruction following, knowledge, reasoning. The 360M model was trained on 4 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new filtered datasets we curated and will release soon.  We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our own curated datasets. We then applied Direct Preference Optimization (DPO) using [UltraFeedback](https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized).
+The instruct model additionally supports tasks such as text rewriting, summarization and function calling thanks to datasets developed by [Argilla](https://huggingface.co/argilla) such as [Synth-APIGen-v0.1](https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1).
 ### How to use