adding Yoruba BERT

Browse files

Files changed (7) hide show

README.md +47 -0
config.json +30 -0
pytorch_model.bin +3 -0
special_tokens_map.json +1 -0
tokenizer_config.json +1 -0
training_args.bin +3 -0
vocab.txt +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,47 @@

+Hugging Face's logo
+---
+language: yo
+datasets:
+- Bible, JW300, [Menyo-20k](https://huggingface.co/datasets/menyo20k_mt), [Yoruba Embedding corpus](https://huggingface.co/datasets/yoruba_text_c3) and [CC-Aligned](https://opus.nlpl.eu/), Wikipedia, news corpora (BBC Yoruba, VON Yoruba, Asejere, Alaroye), and other small datasets curated from friends.
+---
+# bert-base-multilingual-cased-finetuned-yoruba
+## Model description
+**bert-base-multilingual-cased-finetuned-yoruba** is a **Yoruba BERT** model obtained by fine-tuning **bert-base-multilingual-cased** model on Yorùbá language texts.  It provides **better performance** than the multilingual BERT on text classification and named entity recognition datasets.
+Specifically, this model is a *bert-base-multilingual-cased* model that was fine-tuned on Yorùbá corpus.
+## Intended uses & limitations
+#### How to use
+You can use this model with Transformers *pipeline* for masked token prediction.
+```python
+from transformers import AutoTokenizer, AutoModelForTokenClassification
+from transformers import pipeline
+tokenizer = AutoTokenizer.from_pretrained("")
+model = AutoModelForTokenClassification.from_pretrained("")
+nlp = pipeline("", model=model, tokenizer=tokenizer)
+example = "Emir of Kano turban Zhang wey don spend 18 years for Nigeria"
+ner_results = nlp(example)
+print(ner_results)
+```
+#### Limitations and bias
+This model is limited by its training dataset of entity-annotated news articles from a specific span of time. This may not generalize well for all use cases in different domains.
+## Training data
+This model was fine-tuned on on  JW300 Yorùbá corpus and [Menyo-20k](https://huggingface.co/datasets/menyo20k_mt) dataset
+## Training procedure
+This model was trained on a single NVIDIA V100 GPU
+## Eval results on Test set (F-score)
+Dataset|F1-score
+-|-
+Yoruba GV NER |86.26
+MasakhaNER |75.76
+BBC Yoruba |91.75
+### BibTeX entry and citation info
+By David Adelani
+```
+```

config.json ADDED Viewed

	@@ -0,0 +1,30 @@

+{
+  "_name_or_path": "bert-base-multilingual-cased",
+  "architectures": [
+    "BertForMaskedLM"
+  ],
+  "attention_probs_dropout_prob": 0.1,
+  "directionality": "bidi",
+  "gradient_checkpointing": false,
+  "hidden_act": "gelu",
+  "hidden_dropout_prob": 0.1,
+  "hidden_size": 768,
+  "initializer_range": 0.02,
+  "intermediate_size": 3072,
+  "layer_norm_eps": 1e-12,
+  "max_position_embeddings": 512,
+  "model_type": "bert",
+  "num_attention_heads": 12,
+  "num_hidden_layers": 12,
+  "pad_token_id": 0,
+  "pooler_fc_size": 768,
+  "pooler_num_attention_heads": 12,
+  "pooler_num_fc_layers": 3,
+  "pooler_size_per_head": 128,
+  "pooler_type": "first_token_transform",
+  "position_embedding_type": "absolute",
+  "transformers_version": "4.3.2",
+  "type_vocab_size": 2,
+  "use_cache": true,
+  "vocab_size": 119547
+}

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:1fd8d904fbd91faefa02c18ee1c36564e2a7332889d4437b462b0ad03910a99b
+size 711988242

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1 @@


1	+ {"unk_token": "[UNK]", "sep_token": "[SEP]", "pad_token": "[PAD]", "cls_token": "[CLS]", "mask_token": "[MASK]"}

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1 @@


1	+ {"do_lower_case": false, "unk_token": "[UNK]", "sep_token": "[SEP]", "pad_token": "[PAD]", "cls_token": "[CLS]", "mask_token": "[MASK]", "tokenize_chinese_chars": true, "strip_accents": null, "model_max_length": 512, "name_or_path": "bert-base-multilingual-cased"}

training_args.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f7208c95e83a8c6c168de1505adabce7f51277e3067207441ffbb842e12312fd
+size 2095

vocab.txt ADDED Viewed

The diff for this file is too large to render. See raw diff