s-nlp
/

mt0-xl-detox-orpo

Text2Text Generation

Inference Endpoints

Model card Files Files and versions Community

lmeribal commited on Jul 5

Commit

e3e2445

•

1 Parent(s): 0495ca1

Update README.md

Files changed (1) hide show

README.md +3 -4

README.md CHANGED Viewed

@@ -27,6 +27,8 @@ pipeline_tag: text2text-generation
  ## Model Information
 This is a multilingual 3.7B text detoxification model for 9 languages built on [TextDetox 2024 shared task](https://pan.webis.de/clef24/pan24-web/text-detoxification.html) based on [mT0-xl](https://huggingface.co/bigscience/mt0-xl). The model was trained in a two-step setup: the first step is full fine-tuning on different parallel text detoxification datasets, and the second step is ORPO alignment on a self-annotated preference dataset collected using toxicity and similarity classifiers. See the paper for more details.
  ## Example usage
  ```python
@@ -64,7 +66,4 @@ tokenizer = AutoTokenizer.from_pretrained('s-nlp/mt0-xl-detox-orpo')
     return tokenizer.batch_decode(outputs, skip_special_tokens=True)
 ```
- ## Human evaluation
- ## Automatic evaluation

  ## Model Information
 This is a multilingual 3.7B text detoxification model for 9 languages built on [TextDetox 2024 shared task](https://pan.webis.de/clef24/pan24-web/text-detoxification.html) based on [mT0-xl](https://huggingface.co/bigscience/mt0-xl). The model was trained in a two-step setup: the first step is full fine-tuning on different parallel text detoxification datasets, and the second step is ORPO alignment on a self-annotated preference dataset collected using toxicity and similarity classifiers. See the paper for more details.
+The model shows state-of-the-art performance for the Ukrainian language on the [TextDetox 2024 shared task](https://pan.webis.de/clef24/pan24-web/text-detoxification.html), top-2 scores for Arabic, and near state-of-the-art performance for other languages. Overall, the model is the second best approach on the entire human-rated leaderboard.
  ## Example usage
  ```python
     return tokenizer.batch_decode(outputs, skip_special_tokens=True)
 ```