Upload 9 files

Browse files

Files changed (9) hide show

README.md +361 -3
config.json +9 -0
model.bin +3 -0
sentencepiece.bpe.model +3 -0
shared_vocabulary.json +0 -0
special_tokens_map.json +109 -0
tokenization_small100.py +365 -0
tokenizer_config.json +118 -0
vocab.json +0 -0

README.md CHANGED Viewed

@@ -1,3 +1,361 @@
----
-license: mit
----

+---
+language:
+- multilingual
+- af
+- am
+- ar
+- ast
+- az
+- ba
+- be
+- bg
+- bn
+- br
+- bs
+- ca
+- ceb
+- cs
+- cy
+- da
+- de
+- el
+- en
+- es
+- et
+- fa
+- ff
+- fi
+- fr
+- fy
+- ga
+- gd
+- gl
+- gu
+- ha
+- he
+- hi
+- hr
+- ht
+- hu
+- hy
+- id
+- ig
+- ilo
+- is
+- it
+- ja
+- jv
+- ka
+- kk
+- km
+- kn
+- ko
+- lb
+- lg
+- ln
+- lo
+- lt
+- lv
+- mg
+- mk
+- ml
+- mn
+- mr
+- ms
+- my
+- ne
+- nl
+- 'no'
+- ns
+- oc
+- or
+- pa
+- pl
+- ps
+- pt
+- ro
+- ru
+- sd
+- si
+- sk
+- sl
+- so
+- sq
+- sr
+- ss
+- su
+- sv
+- sw
+- ta
+- th
+- tl
+- tn
+- tr
+- uk
+- ur
+- uz
+- vi
+- wo
+- xh
+- yi
+- yo
+- zh
+- zu
+license: mit
+tags:
+- small100
+- translation
+- flores101
+- gsarti/flores_101
+- tico19
+- gmnlp/tico19
+- tatoeba
+datasets:
+- tico19
+- flores101
+- tatoeba
+---
+From: https://huggingface.co/alirezamsh/small100
+# SMALL-100 Model
+SMaLL-100 is a compact and fast massively multilingual machine translation model covering more than 10K language pairs, that achieves competitive results with M2M-100 while being much smaller and faster. It is introduced in [this paper](https://arxiv.org/abs/2210.11621)(accepted to EMNLP2022), and initially released in [this repository](https://github.com/alirezamshi/small100).
+The model architecture and config are the same as [M2M-100](https://huggingface.co/facebook/m2m100_418M/tree/main) implementation, but the tokenizer is modified to adjust language codes. So, you should load the tokenizer locally from [tokenization_small100.py](https://huggingface.co/alirezamsh/small100/blob/main/tokenization_small100.py) file for the moment.
+**Demo**: https://huggingface.co/spaces/alirezamsh/small100
+**Note**: SMALL100Tokenizer requires sentencepiece, so make sure to install it by:
+```pip install sentencepiece```
+- **Supervised Training**
+SMaLL-100 is a seq-to-seq model for the translation task. The input to the model is ```source:[tgt_lang_code] + src_tokens + [EOS]``` and ```target: tgt_tokens + [EOS]```.
+An example of supervised training is shown below:
+```
+from transformers import M2M100ForConditionalGeneration
+from tokenization_small100 import SMALL100Tokenizer
+model = M2M100ForConditionalGeneration.from_pretrained("alirezamsh/small100")
+tokenizer = SMALL100Tokenizer.from_pretrained("alirezamsh/small100", tgt_lang="fr")
+src_text = "Life is like a box of chocolates."
+tgt_text = "La vie est comme une boîte de chocolat."
+model_inputs = tokenizer(src_text, text_target=tgt_text, return_tensors="pt")
+loss = model(**model_inputs).loss  # forward pass
+```
+Training data can be provided upon request.
+- **Generation**
+Beam size of 5, and maximum target length of 256 is used for the generation.
+- **Evaluation**
+Please refer to [original repository](https://github.com/alirezamshi/small100) for spBLEU computation.
+- **Languages Covered**
+Afrikaans (af), Amharic (am), Arabic (ar), Asturian (ast), Azerbaijani (az), Bashkir (ba), Belarusian (be), Bulgarian (bg), Bengali (bn), Breton (br), Bosnian (bs), Catalan; Valencian (ca), Cebuano (ceb), Czech (cs), Welsh (cy), Danish (da), German (de), Greeek (el), English (en), Spanish (es), Estonian (et), Persian (fa), Fulah (ff), Finnish (fi), French (fr), Western Frisian (fy), Irish (ga), Gaelic; Scottish Gaelic (gd), Galician (gl), Gujarati (gu), Hausa (ha), Hebrew (he), Hindi (hi), Croatian (hr), Haitian; Haitian Creole (ht), Hungarian (hu), Armenian (hy), Indonesian (id), Igbo (ig), Iloko (ilo), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Kazakh (kk), Central Khmer (km), Kannada (kn), Korean (ko), Luxembourgish; Letzeburgesch (lb), Ganda (lg), Lingala (ln), Lao (lo), Lithuanian (lt), Latvian (lv), Malagasy (mg), Macedonian (mk), Malayalam (ml), Mongolian (mn), Marathi (mr), Malay (ms), Burmese (my), Nepali (ne), Dutch; Flemish (nl), Norwegian (no), Northern Sotho (ns), Occitan (post 1500) (oc), Oriya (or), Panjabi; Punjabi (pa), Polish (pl), Pushto; Pashto (ps), Portuguese (pt), Romanian; Moldavian; Moldovan (ro), Russian (ru), Sindhi (sd), Sinhala; Sinhalese (si), Slovak (sk), Slovenian (sl), Somali (so), Albanian (sq), Serbian (sr), Swati (ss), Sundanese (su), Swedish (sv), Swahili (sw), Tamil (ta), Thai (th), Tagalog (tl), Tswana (tn), Turkish (tr), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Wolof (wo), Xhosa (xh), Yiddish (yi), Yoruba (yo), Chinese (zh), Zulu (zu)
+# Citation
+If you use this model for your research, please cite the following work:
+```
+@inproceedings{mohammadshahi-etal-2022-small,
+    title = "{SM}a{LL}-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource Languages",
+    author = "Mohammadshahi, Alireza  and
+      Nikoulina, Vassilina  and
+      Berard, Alexandre  and
+      Brun, Caroline  and
+      Henderson, James  and
+      Besacier, Laurent",
+    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
+    month = dec,
+    year = "2022",
+    address = "Abu Dhabi, United Arab Emirates",
+    publisher = "Association for Computational Linguistics",
+    url = "https://aclanthology.org/2022.emnlp-main.571",
+    pages = "8348--8359",
+    abstract = "In recent years, multilingual machine translation models have achieved promising performance on low-resource language pairs by sharing information between similar languages, thus enabling zero-shot translation. To overcome the {``}curse of multilinguality{''}, these models often opt for scaling up the number of parameters, which makes their use in resource-constrained environments challenging. We introduce SMaLL-100, a distilled version of the M2M-100(12B) model, a massively multilingual machine translation model covering 100 languages. We train SMaLL-100 with uniform sampling across all language pairs and therefore focus on preserving the performance of low-resource languages. We evaluate SMaLL-100 on different low-resource benchmarks: FLORES-101, Tatoeba, and TICO-19 and demonstrate that it outperforms previous massively multilingual models of comparable sizes (200-600M) while improving inference latency and memory usage. Additionally, our model achieves comparable results to M2M-100 (1.2B), while being 3.6x smaller and 4.3x faster at inference.",
+}
+@inproceedings{mohammadshahi-etal-2022-compressed,
+    title = "What Do Compressed Multilingual Machine Translation Models Forget?",
+    author = "Mohammadshahi, Alireza  and
+      Nikoulina, Vassilina  and
+      Berard, Alexandre  and
+      Brun, Caroline  and
+      Henderson, James  and
+      Besacier, Laurent",
+    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2022",
+    month = dec,
+    year = "2022",
+    address = "Abu Dhabi, United Arab Emirates",
+    publisher = "Association for Computational Linguistics",
+    url = "https://aclanthology.org/2022.findings-emnlp.317",
+    pages = "4308--4329",
+    abstract = "Recently, very large pre-trained models achieve state-of-the-art results in various natural language processing (NLP) tasks, but their size makes it more challenging to apply them in resource-constrained environments. Compression techniques allow to drastically reduce the size of the models and therefore their inference time with negligible impact on top-tier metrics. However, the general performance averaged across multiple tasks and/or languages may hide a drastic performance drop on under-represented features, which could result in the amplification of biases encoded by the models. In this work, we assess the impact of compression methods on Multilingual Neural Machine Translation models (MNMT) for various language groups, gender, and semantic biases by extensive analysis of compressed models on different machine translation benchmarks, i.e. FLORES-101, MT-Gender, and DiBiMT. We show that the performance of under-represented languages drops significantly, while the average BLEU metric only slightly decreases. Interestingly, the removal of noisy memorization with compression leads to a significant improvement for some medium-resource languages. Finally, we demonstrate that compression amplifies intrinsic gender and semantic biases, even in high-resource languages.",
+}
+```
+## How to download this model using python
+- Install Python https://www.python.org/downloads/
+- cmd
+- python --version
+- python -m pip install huggingface_hub
+- python
+```
+import huggingface_hub
+huggingface_hub.download_snapshot('entai2965/small100-ctranslate2',local_dir='small100-ctranslate2')
+```
+## How to run this model
+- https://opennmt.net/CTranslate2/guides/transformers.html#m2m-100
+- https://huggingface.co/alirezamsh/small100
+- cmd
+- python -m pip install ctranslate2 transformers
+- python
+```
+import sys
+import ctranslate2
+#model_path=r'Downloads\models\small100-ctranslate2'
+model_path='Downloads/models/small100-ctranslate2'
+sys.path.insert(1,model_path)
+from tokenization_small100 import SMALL100Tokenizer
+string1='जीवन एक चॉकलेट बॉक्स की तरह है।'
+translator=ctranslate2.Translator(model_path,device='cpu')
+tokenizer=SMALL100Tokenizer.from_pretrained(model_path, clean_up_tokenization_spaces=True)
+tokenizer.src_lang='hi'
+tokenizer.tgt_lang='es'
+target_language_token=[tokenizer.lang_code_to_token['es']]
+encoded_string=tokenizer.convert_ids_to_tokens(tokenizer.encode(string1))
+output=translator.translate_batch([encoded_string], target_prefix=[target_language_token])
+output=tokenizer.decode(tokenizer.convert_tokens_to_ids(output[0].hypotheses[0][1:]))
+print(output)
+```
+## How to run this model (batch syntax)
+```
+import sys
+import os
+import ctranslate2
+#set defaults
+model_name='alirezamsh/small100'
+home_path=os.path.expanduser('~')
+model_path=home_path+'/Downloads/models/small100-ctranslate2'
+source_language_code='hi'
+#target_language_code='ar'
+#target_language_code='fr'
+#target_language_code='en'
+target_language_code='es'
+device='cpu'
+#device=gpu
+#import tokenizer.py library
+#https://stackoverflow.com/questions/16114391/adding-directory-to-sys-path-pythonpath
+sys.path.insert(1,model_path)
+from tokenization_small100 import SMALL100Tokenizer
+#load data, languages list ->   https://huggingface.co/alirezamsh/small100   <-
+string1='जीवन एक चॉकलेट बॉक्स की तरह है।'
+string2='生活就像一盒巧克力。'
+string3="You never know what you are going to get."
+raw_list=[string1,string2,string3]
+#load models
+translator=ctranslate2.Translator(model_path,device='cpu')
+tokenizer=SMALL100Tokenizer.from_pretrained(model_path, clean_up_tokenization_spaces=True)
+#configure languages
+tokenizer.src_lang=source_language_code #this tokenizer seems to completely ignore this setting
+tokenizer.tgt_lang=target_language_code
+target_language_token=[tokenizer.lang_code_to_token[target_language_code]]
+#encode
+encoded_list=[]
+for text in raw_list:
+    encoded_list.append(tokenizer.convert_ids_to_tokens(tokenizer.encode(text)))
+# translate
+translated_list=translator.translate_batch(encoded_list,target_prefix=[target_language_token]*len(raw_list))
+#decode
+for counter,token in enumerate(translated_list):
+    translated_list[counter]=tokenizer.decode(tokenizer.convert_tokens_to_ids(token.hypotheses[0][1:]))
+#output
+for text in translated_list:
+    print(text)
+```
+[Functional programming](https://docs.python.org/3/howto/functional.html) version
+```
+import sys
+import os
+import ctranslate2
+#set defaults
+model_name='alirezamsh/small100'
+home_path=os.path.expanduser('~')
+model_path=home_path+'/Downloads/models/models--alirezamsh--small100-ctranslate2'
+source_language_code='hi'
+#target_language_code='ar'
+#target_language_code='fr'
+#target_language_code='en'
+target_language_code='es'
+device='cpu'
+#device=gpu
+#import tokenizer.py library
+#https://stackoverflow.com/questions/16114391/adding-directory-to-sys-path-pythonpath
+sys.path.insert(1,model_path)
+from tokenization_small100 import SMALL100Tokenizer
+#load data, languages list ->   https://huggingface.co/alirezamsh/small100   <-
+string1='जीवन एक चॉकलेट बॉक्स की तरह है।'
+string2='生活就像一盒巧克力。'
+string3="You never know what you are going to get."
+raw_list=[string1,string2,string3]
+#load models
+translator=ctranslate2.Translator(model_path,device='cpu')
+tokenizer=SMALL100Tokenizer.from_pretrained(model_path, clean_up_tokenization_spaces=True)
+tokenizer.tgt_lang=target_language_code
+#invoke witchcraft
+translated_list=[tokenizer.decode(tokenizer.convert_tokens_to_ids(token.hypotheses[0][1:])) for token in translator.translate_batch([tokenizer.convert_ids_to_tokens(tokenizer.encode(text)) for text in raw_list],target_prefix=[[tokenizer.lang_code_to_token[target_language_code]]]*len(raw_list))]
+#output
+for text in translated_list:
+    print(text)
+```

config.json ADDED Viewed

	@@ -0,0 +1,9 @@

+{
+  "add_source_bos": false,
+  "add_source_eos": false,
+  "bos_token": "<s>",
+  "decoder_start_token": "</s>",
+  "eos_token": "</s>",
+  "layer_norm_epsilon": null,
+  "unk_token": "<unk>"
+}

model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4fae54c20aa744b25ecee3ec84b0c2ffa5338179d4cb2dd4e03783f5cc7740d5
+size 1335148325

sentencepiece.bpe.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d8f7c76ed2a5e0822be39f0a4f95a55eb19c78f4593ce609e2edbc2aea4d380a
+size 2423393

shared_vocabulary.json ADDED Viewed

The diff for this file is too large to render. See raw diff

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,109 @@

+{
+  "additional_special_tokens": [
+    "__af__",
+    "__am__",
+    "__ar__",
+    "__ast__",
+    "__az__",
+    "__ba__",
+    "__be__",
+    "__bg__",
+    "__bn__",
+    "__br__",
+    "__bs__",
+    "__ca__",
+    "__ceb__",
+    "__cs__",
+    "__cy__",
+    "__da__",
+    "__de__",
+    "__el__",
+    "__en__",
+    "__es__",
+    "__et__",
+    "__fa__",
+    "__ff__",
+    "__fi__",
+    "__fr__",
+    "__fy__",
+    "__ga__",
+    "__gd__",
+    "__gl__",
+    "__gu__",
+    "__ha__",
+    "__he__",
+    "__hi__",
+    "__hr__",
+    "__ht__",
+    "__hu__",
+    "__hy__",
+    "__id__",
+    "__ig__",
+    "__ilo__",
+    "__is__",
+    "__it__",
+    "__ja__",
+    "__jv__",
+    "__ka__",
+    "__kk__",
+    "__km__",
+    "__kn__",
+    "__ko__",
+    "__lb__",
+    "__lg__",
+    "__ln__",
+    "__lo__",
+    "__lt__",
+    "__lv__",
+    "__mg__",
+    "__mk__",
+    "__ml__",
+    "__mn__",
+    "__mr__",
+    "__ms__",
+    "__my__",
+    "__ne__",
+    "__nl__",
+    "__no__",
+    "__ns__",
+    "__oc__",
+    "__or__",
+    "__pa__",
+    "__pl__",
+    "__ps__",
+    "__pt__",
+    "__ro__",
+    "__ru__",
+    "__sd__",
+    "__si__",
+    "__sk__",
+    "__sl__",
+    "__so__",
+    "__sq__",
+    "__sr__",
+    "__ss__",
+    "__su__",
+    "__sv__",
+    "__sw__",
+    "__ta__",
+    "__th__",
+    "__tl__",
+    "__tn__",
+    "__tr__",
+    "__uk__",
+    "__ur__",
+    "__uz__",
+    "__vi__",
+    "__wo__",
+    "__xh__",
+    "__yi__",
+    "__yo__",
+    "__zh__",
+    "__zu__"
+  ],
+  "bos_token": "<s>",
+  "eos_token": "</s>",
+  "pad_token": "<pad>",
+  "sep_token": "</s>",
+  "unk_token": "<unk>"
+}

tokenization_small100.py ADDED Viewed

	@@ -0,0 +1,365 @@

+# Copyright (c) 2022 Idiap Research Institute, http://www.idiap.ch/
+# Written by Alireza Mohammadshahi <[email protected]>
+# This is a modified version of https://github.com/huggingface/transformers/blob/main/src/transformers/models/m2m_100/tokenization_m2m_100.py
+# which owns by Fariseq Authors and The HuggingFace Inc. team.
+#
+#
+# Licensed under the Apache License, Version 2.0 (the "License");
+# you may not use this file except in compliance with the License.
+# You may obtain a copy of the License at
+#
+#     http://www.apache.org/licenses/LICENSE-2.0
+#
+# Unless required by applicable law or agreed to in writing, software
+# distributed under the License is distributed on an "AS IS" BASIS,
+# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+# See the License for the specific language governing permissions and
+# limitations under the License.
+"""Tokenization classes for SMALL100."""
+import json
+import os
+from pathlib import Path
+from shutil import copyfile
+from typing import Any, Dict, List, Optional, Tuple, Union
+import sentencepiece
+from transformers.tokenization_utils import BatchEncoding, PreTrainedTokenizer
+from transformers.utils import logging
+logger = logging.get_logger(__name__)
+SPIECE_UNDERLINE = "▁"
+VOCAB_FILES_NAMES = {
+    "vocab_file": "vocab.json",
+    "spm_file": "sentencepiece.bpe.model",
+    "tokenizer_config_file": "tokenizer_config.json",
+}
+PRETRAINED_VOCAB_FILES_MAP = {
+    "vocab_file": {
+        "alirezamsh/small100": "https://huggingface.co/alirezamsh/small100/resolve/main/vocab.json",
+    },
+    "spm_file": {
+        "alirezamsh/small100": "https://huggingface.co/alirezamsh/small100/resolve/main/sentencepiece.bpe.model",
+    },
+    "tokenizer_config_file": {
+        "alirezamsh/small100": "https://huggingface.co/alirezamsh/small100/resolve/main/tokenizer_config.json",
+    },
+}
+PRETRAINED_POSITIONAL_EMBEDDINGS_SIZES = {
+    "alirezamsh/small100": 1024,
+}
+# fmt: off
+FAIRSEQ_LANGUAGE_CODES = {
+    "m2m100": ["af", "am", "ar", "ast", "az", "ba", "be", "bg", "bn", "br", "bs", "ca", "ceb", "cs", "cy", "da", "de", "el", "en", "es", "et", "fa", "ff", "fi", "fr", "fy", "ga", "gd", "gl", "gu", "ha", "he", "hi", "hr", "ht", "hu", "hy", "id", "ig", "ilo", "is", "it", "ja", "jv", "ka", "kk", "km", "kn", "ko", "lb", "lg", "ln", "lo", "lt", "lv", "mg", "mk", "ml", "mn", "mr", "ms", "my", "ne", "nl", "no", "ns", "oc", "or", "pa", "pl", "ps", "pt", "ro", "ru", "sd", "si", "sk", "sl", "so", "sq", "sr", "ss", "su", "sv", "sw", "ta", "th", "tl", "tn", "tr", "uk", "ur", "uz", "vi", "wo", "xh", "yi", "yo", "zh", "zu"]
+}
+# fmt: on
+class SMALL100Tokenizer(PreTrainedTokenizer):
+    """
+    Construct an SMALL100 tokenizer. Based on [SentencePiece](https://github.com/google/sentencepiece).
+    This tokenizer inherits from [`PreTrainedTokenizer`] which contains most of the main methods. Users should refer to
+    this superclass for more information regarding those methods.
+    Args:
+        vocab_file (`str`):
+            Path to the vocabulary file.
+        spm_file (`str`):
+            Path to [SentencePiece](https://github.com/google/sentencepiece) file (generally has a .spm extension) that
+            contains the vocabulary.
+        tgt_lang (`str`, *optional*):
+            A string representing the target language.
+        eos_token (`str`, *optional*, defaults to `"</s>"`):
+            The end of sequence token.
+        sep_token (`str`, *optional*, defaults to `"</s>"`):
+            The separator token, which is used when building a sequence from multiple sequences, e.g. two sequences for
+            sequence classification or for a text and a question for question answering. It is also used as the last
+            token of a sequence built with special tokens.
+        unk_token (`str`, *optional*, defaults to `"<unk>"`):
+            The unknown token. A token that is not in the vocabulary cannot be converted to an ID and is set to be this
+            token instead.
+        pad_token (`str`, *optional*, defaults to `"<pad>"`):
+            The token used for padding, for example when batching sequences of different lengths.
+        language_codes (`str`, *optional*):
+            What language codes to use. Should be `"m2m100"`.
+        sp_model_kwargs (`dict`, *optional*):
+            Will be passed to the `SentencePieceProcessor.__init__()` method. The [Python wrapper for
+            SentencePiece](https://github.com/google/sentencepiece/tree/master/python) can be used, among other things,
+            to set:
+            - `enable_sampling`: Enable subword regularization.
+            - `nbest_size`: Sampling parameters for unigram. Invalid for BPE-Dropout.
+              - `nbest_size = {0,1}`: No sampling is performed.
+              - `nbest_size > 1`: samples from the nbest_size results.
+              - `nbest_size < 0`: assuming that nbest_size is infinite and samples from the all hypothesis (lattice)
+                using forward-filtering-and-backward-sampling algorithm.
+            - `alpha`: Smoothing parameter for unigram sampling, and dropout probability of merge operations for
+              BPE-dropout.
+    Examples:
+    ```python
+    >>> from tokenization_small100 import SMALL100Tokenizer
+    >>> tokenizer = SMALL100Tokenizer.from_pretrained("alirezamsh/small100", tgt_lang="ro")
+    >>> src_text = " UN Chief Says There Is No Military Solution in Syria"
+    >>> tgt_text = "Şeful ONU declară că nu există o soluţie militară în Siria"
+    >>> model_inputs = tokenizer(src_text, text_target=tgt_text, return_tensors="pt")
+    >>> model(**model_inputs)  # should work
+    ```"""
+    vocab_files_names = VOCAB_FILES_NAMES
+    max_model_input_sizes = PRETRAINED_POSITIONAL_EMBEDDINGS_SIZES
+    pretrained_vocab_files_map = PRETRAINED_VOCAB_FILES_MAP
+    model_input_names = ["input_ids", "attention_mask"]
+    prefix_tokens: List[int] = []
+    suffix_tokens: List[int] = []
+    def __init__(
+        self,
+        vocab_file,
+        spm_file,
+        tgt_lang=None,
+        bos_token="<s>",
+        eos_token="</s>",
+        sep_token="</s>",
+        pad_token="<pad>",
+        unk_token="<unk>",
+        language_codes="m2m100",
+        sp_model_kwargs: Optional[Dict[str, Any]] = None,
+        num_madeup_words=8,
+        **kwargs,
+    ) -> None:
+        self.sp_model_kwargs = {} if sp_model_kwargs is None else sp_model_kwargs
+        self.language_codes = language_codes
+        fairseq_language_code = FAIRSEQ_LANGUAGE_CODES[language_codes]
+        self.lang_code_to_token = {lang_code: f"__{lang_code}__" for lang_code in fairseq_language_code}
+        kwargs["additional_special_tokens"] = kwargs.get("additional_special_tokens", [])
+        kwargs["additional_special_tokens"] += [
+            self.get_lang_token(lang_code)
+            for lang_code in fairseq_language_code
+            if self.get_lang_token(lang_code) not in kwargs["additional_special_tokens"]
+        ]
+        self.vocab_file = vocab_file
+        self.encoder = load_json(vocab_file)
+        self.decoder = {v: k for k, v in self.encoder.items()}
+        self.spm_file = spm_file
+        self.sp_model = load_spm(spm_file, self.sp_model_kwargs)
+        self.encoder_size = len(self.encoder)
+        self.lang_token_to_id = {
+            self.get_lang_token(lang_code): self.encoder_size + i for i, lang_code in enumerate(fairseq_language_code)
+        }
+        self.lang_code_to_id = {lang_code: self.encoder_size + i for i, lang_code in enumerate(fairseq_language_code)}
+        self.id_to_lang_token = {v: k for k, v in self.lang_token_to_id.items()}
+        self._tgt_lang = tgt_lang if tgt_lang is not None else "en"
+        self.cur_lang_id = self.get_lang_id(self._tgt_lang)
+        self.num_madeup_words = num_madeup_words
+        super().__init__(
+            tgt_lang=tgt_lang,
+            bos_token=bos_token,
+            eos_token=eos_token,
+            sep_token=sep_token,
+            unk_token=unk_token,
+            pad_token=pad_token,
+            language_codes=language_codes,
+            sp_model_kwargs=self.sp_model_kwargs,
+            num_madeup_words=num_madeup_words,
+            **kwargs,
+        )
+        self.set_lang_special_tokens(self._tgt_lang)
+    @property
+    def vocab_size(self) -> int:
+        return len(self.encoder) + len(self.lang_token_to_id) + self.num_madeup_words
+    @property
+    def tgt_lang(self) -> str:
+        return self._tgt_lang
+    @tgt_lang.setter
+    def tgt_lang(self, new_tgt_lang: str) -> None:
+        self._tgt_lang = new_tgt_lang
+        self.set_lang_special_tokens(self._tgt_lang)
+    def _tokenize(self, text: str) -> List[str]:
+        return self.sp_model.encode(text, out_type=str)
+    def _convert_token_to_id(self, token):
+        if token in self.lang_token_to_id:
+            return self.lang_token_to_id[token]
+        return self.encoder.get(token, self.encoder[self.unk_token])
+    def _convert_id_to_token(self, index: int) -> str:
+        """Converts an index (integer) in a token (str) using the decoder."""
+        if index in self.id_to_lang_token:
+            return self.id_to_lang_token[index]
+        return self.decoder.get(index, self.unk_token)
+    def convert_tokens_to_string(self, tokens: List[str]) -> str:
+        """Converts a sequence of tokens (strings for sub-words) in a single string."""
+        return self.sp_model.decode(tokens)
+    def get_special_tokens_mask(
+        self, token_ids_0: List[int], token_ids_1: Optional[List[int]] = None, already_has_special_tokens: bool = False
+    ) -> List[int]:
+        """
+        Retrieve sequence ids from a token list that has no special tokens added. This method is called when adding
+        special tokens using the tokenizer `prepare_for_model` method.
+        Args:
+            token_ids_0 (`List[int]`):
+                List of IDs.
+            token_ids_1 (`List[int]`, *optional*):
+                Optional second list of IDs for sequence pairs.
+            already_has_special_tokens (`bool`, *optional*, defaults to `False`):
+                Whether or not the token list is already formatted with special tokens for the model.
+        Returns:
+            `List[int]`: A list of integers in the range [0, 1]: 1 for a special token, 0 for a sequence token.
+        """
+        if already_has_special_tokens:
+            return super().get_special_tokens_mask(
+                token_ids_0=token_ids_0, token_ids_1=token_ids_1, already_has_special_tokens=True
+            )
+        prefix_ones = [1] * len(self.prefix_tokens)
+        suffix_ones = [1] * len(self.suffix_tokens)
+        if token_ids_1 is None:
+            return prefix_ones + ([0] * len(token_ids_0)) + suffix_ones
+        return prefix_ones + ([0] * len(token_ids_0)) + ([0] * len(token_ids_1)) + suffix_ones
+    def build_inputs_with_special_tokens(
+        self, token_ids_0: List[int], token_ids_1: Optional[List[int]] = None
+    ) -> List[int]:
+        """
+        Build model inputs from a sequence or a pair of sequence for sequence classification tasks by concatenating and
+        adding special tokens. An MBART sequence has the following format, where `X` represents the sequence:
+        - `input_ids` (for encoder) `X [eos, src_lang_code]`
+        - `decoder_input_ids`: (for decoder) `X [eos, tgt_lang_code]`
+        BOS is never used. Pairs of sequences are not the expected use case, but they will be handled without a
+        separator.
+        Args:
+            token_ids_0 (`List[int]`):
+                List of IDs to which the special tokens will be added.
+            token_ids_1 (`List[int]`, *optional*):
+                Optional second list of IDs for sequence pairs.
+        Returns:
+            `List[int]`: List of [input IDs](../glossary#input-ids) with the appropriate special tokens.
+        """
+        if token_ids_1 is None:
+            if self.prefix_tokens is None:
+                return token_ids_0 + self.suffix_tokens
+            else:
+                return self.prefix_tokens + token_ids_0 + self.suffix_tokens
+        # We don't expect to process pairs, but leave the pair logic for API consistency
+        if self.prefix_tokens is None:
+            return token_ids_0 + token_ids_1 + self.suffix_tokens
+        else:
+            return self.prefix_tokens + token_ids_0 + token_ids_1 + self.suffix_tokens
+    def get_vocab(self) -> Dict:
+        vocab = {self.convert_ids_to_tokens(i): i for i in range(self.vocab_size)}
+        vocab.update(self.added_tokens_encoder)
+        return vocab
+    def __getstate__(self) -> Dict:
+        state = self.__dict__.copy()
+        state["sp_model"] = None
+        return state
+    def __setstate__(self, d: Dict) -> None:
+        self.__dict__ = d
+        # for backward compatibility
+        if not hasattr(self, "sp_model_kwargs"):
+            self.sp_model_kwargs = {}
+        self.sp_model = load_spm(self.spm_file, self.sp_model_kwargs)
+    def save_vocabulary(self, save_directory: str, filename_prefix: Optional[str] = None) -> Tuple[str]:
+        save_dir = Path(save_directory)
+        if not save_dir.is_dir():
+            raise OSError(f"{save_directory} should be a directory")
+        vocab_save_path = save_dir / (
+            (filename_prefix + "-" if filename_prefix else "") + self.vocab_files_names["vocab_file"]
+        )
+        spm_save_path = save_dir / (
+            (filename_prefix + "-" if filename_prefix else "") + self.vocab_files_names["spm_file"]
+        )
+        save_json(self.encoder, vocab_save_path)
+        if os.path.abspath(self.spm_file) != os.path.abspath(spm_save_path) and os.path.isfile(self.spm_file):
+            copyfile(self.spm_file, spm_save_path)
+        elif not os.path.isfile(self.spm_file):
+            with open(spm_save_path, "wb") as fi:
+                content_spiece_model = self.sp_model.serialized_model_proto()
+                fi.write(content_spiece_model)
+        return (str(vocab_save_path), str(spm_save_path))
+    def prepare_seq2seq_batch(
+        self,
+        src_texts: List[str],
+        tgt_texts: Optional[List[str]] = None,
+        tgt_lang: str = "ro",
+        **kwargs,
+    ) -> BatchEncoding:
+        self.tgt_lang = tgt_lang
+        self.set_lang_special_tokens(self.tgt_lang)
+        return super().prepare_seq2seq_batch(src_texts, tgt_texts, **kwargs)
+    def _build_translation_inputs(self, raw_inputs, tgt_lang: Optional[str], **extra_kwargs):
+        """Used by translation pipeline, to prepare inputs for the generate function"""
+        if tgt_lang is None:
+            raise ValueError("Translation requires a `tgt_lang` for this model")
+        self.tgt_lang = tgt_lang
+        inputs = self(raw_inputs, add_special_tokens=True, **extra_kwargs)
+        return inputs
+    def _switch_to_input_mode(self):
+        self.set_lang_special_tokens(self.tgt_lang)
+    def _switch_to_target_mode(self):
+        self.prefix_tokens = None
+        self.suffix_tokens = [self.eos_token_id]
+    def set_lang_special_tokens(self, src_lang: str) -> None:
+        """Reset the special tokens to the tgt lang setting. No prefix and suffix=[eos, tgt_lang_code]."""
+        lang_token = self.get_lang_token(src_lang)
+        self.cur_lang_id = self.lang_token_to_id[lang_token]
+        self.prefix_tokens = [self.cur_lang_id]
+        self.suffix_tokens = [self.eos_token_id]
+    def get_lang_token(self, lang: str) -> str:
+        return self.lang_code_to_token[lang]
+    def get_lang_id(self, lang: str) -> int:
+        lang_token = self.get_lang_token(lang)
+        return self.lang_token_to_id[lang_token]
+def load_spm(path: str, sp_model_kwargs: Dict[str, Any]) -> sentencepiece.SentencePieceProcessor:
+    spm = sentencepiece.SentencePieceProcessor(**sp_model_kwargs)
+    spm.Load(str(path))
+    return spm
+def load_json(path: str) -> Union[Dict, List]:
+    with open(path, "r") as f:
+        return json.load(f)
+def save_json(data, path: str) -> None:
+    with open(path, "w") as f:
+        json.dump(data, f, indent=2)

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,118 @@

+{
+  "additional_special_tokens": [
+    "__af__",
+    "__am__",
+    "__ar__",
+    "__ast__",
+    "__az__",
+    "__ba__",
+    "__be__",
+    "__bg__",
+    "__bn__",
+    "__br__",
+    "__bs__",
+    "__ca__",
+    "__ceb__",
+    "__cs__",
+    "__cy__",
+    "__da__",
+    "__de__",
+    "__el__",
+    "__en__",
+    "__es__",
+    "__et__",
+    "__fa__",
+    "__ff__",
+    "__fi__",
+    "__fr__",
+    "__fy__",
+    "__ga__",
+    "__gd__",
+    "__gl__",
+    "__gu__",
+    "__ha__",
+    "__he__",
+    "__hi__",
+    "__hr__",
+    "__ht__",
+    "__hu__",
+    "__hy__",
+    "__id__",
+    "__ig__",
+    "__ilo__",
+    "__is__",
+    "__it__",
+    "__ja__",
+    "__jv__",
+    "__ka__",
+    "__kk__",
+    "__km__",
+    "__kn__",
+    "__ko__",
+    "__lb__",
+    "__lg__",
+    "__ln__",
+    "__lo__",
+    "__lt__",
+    "__lv__",
+    "__mg__",
+    "__mk__",
+    "__ml__",
+    "__mn__",
+    "__mr__",
+    "__ms__",
+    "__my__",
+    "__ne__",
+    "__nl__",
+    "__no__",
+    "__ns__",
+    "__oc__",
+    "__or__",
+    "__pa__",
+    "__pl__",
+    "__ps__",
+    "__pt__",
+    "__ro__",
+    "__ru__",
+    "__sd__",
+    "__si__",
+    "__sk__",
+    "__sl__",
+    "__so__",
+    "__sq__",
+    "__sr__",
+    "__ss__",
+    "__su__",
+    "__sv__",
+    "__sw__",
+    "__ta__",
+    "__th__",
+    "__tl__",
+    "__tn__",
+    "__tr__",
+    "__uk__",
+    "__ur__",
+    "__uz__",
+    "__vi__",
+    "__wo__",
+    "__xh__",
+    "__yi__",
+    "__yo__",
+    "__zh__",
+    "__zu__"
+  ],
+  "bos_token": "<s>",
+  "eos_token": "</s>",
+  "language_codes": "m2m100",
+  "model_max_length": 1024,
+  "name_or_path": "facebook/m2m100_418M",
+  "num_madeup_words": 8,
+  "pad_token": "<pad>",
+  "sep_token": "</s>",
+  "sp_model_kwargs": {},
+  "special_tokens_map_file": "m2m_100_1.2B_v2/special_tokens_map.json",
+  "tgt_lang": null,
+  "tokenizer_class": "M2M100Tokenizer",
+  "tokenizer_file": null,
+  "unk_token": "<unk>"
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff