ccasimiro commited on
Commit
bf095a4
1 Parent(s): b95b8de

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +14 -12
README.md CHANGED
@@ -40,30 +40,32 @@ used in the original [RoBERTA](https://github.com/pytorch/fairseq/tree/master/ex
40
 
41
  ## Training corpora and preprocessing
42
 
43
- The training corpus is composed of several biomedical corpora in Spanish, collected from publicly available corpora and crawlers:
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
  | Name | No. tokens | Description |
46
  |-----------------------------------------------------------------------------------------|-------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
47
  | [Medical crawler](https://zenodo.org/record/4561971#.YTtwM32xXbQ) | 745,705,946 | Crawler of more than 3,000 URLs belonging to Spanish biomedical and health domains. |
48
  | [Scielo](https://github.com/PlanTL-SANIDAD/SciELO-Spain-Crawler) | 60,007,289 | Publications written in Spanish crawled from the Spanish SciELO server in 2017. |
49
  | [BARR2_background](https://temu.bsc.es/BARR2/downloads/background_set.raw_text.tar.bz2) | 24,516,442 | Biomedical Abbreviation Recognition and Resolution (BARR2) containing Spanish clinical case study sections from a variety of clinical disciplines. |
50
- | Wikipedia_life_sciences | 13,890,501 | Wikipedia articles belonging to the Life Sciences category crawled on 04/01/2021 |
51
  | Patents | 13,463,387 | Google Patent in Medical Domain for Spain (Spanish). The accepted codes (Medical Domain) for Json files of patents are: "A61B", "A61C","A61F", "A61H", "A61K", "A61L","A61M", "A61B", "A61P". |
52
  | [EMEA](http://opus.nlpl.eu/download.php?f=EMEA/v3/moses/en-es.txt.zip) | 5,377,448 | Spanish-side documents extracted from parallel corpora made out of PDF documents from the European Medicines Agency. |
53
  | [mespen_Medline](https://zenodo.org/record/3562536#.YTt1fH2xXbR) | 4,166,077 | Spanish-side articles extracted from a collection of Spanish-English parallel corpus consisting of biomedical scientific literature. The collection of parallel resources are aggregated from the MedlinePlus source. |
54
  | PubMed | 1,858,966 | Open-access articles from the PubMed repository crawled in 2017. |
55
 
56
- To obtain a high-quality training corpus, a cleaning pipeline with the following operations has been applied:
57
-
58
- - data parsing in different formats
59
- - sentence splitting
60
- - language detection
61
- - filtering of ill-formed sentences
62
- - deduplication of repetitive contents
63
- - keep the original document boundaries
64
 
65
- Finally, the corpora are concatenated and further global deduplication among the corpora have been applied.
66
- The result is a medium-size biomedical corpus for Spanish composed of about 860M tokens.
67
 
68
  ## Evaluation and results
69
 
 
40
 
41
  ## Training corpora and preprocessing
42
 
43
+ The training corpus is composed of several biomedical corpora in Spanish, collected from publicly available corpora and crawlers.
44
+ To obtain a high-quality training corpus, a cleaning pipeline with the following operations has been applied:
45
+
46
+ - data parsing in different formats
47
+ - sentence splitting
48
+ - language detection
49
+ - filtering of ill-formed sentences
50
+ - deduplication of repetitive contents
51
+ - keep the original document boundaries
52
+
53
+ Finally, the corpora are concatenated and further global deduplication among the corpora have been applied.
54
+ The result is a medium-size biomedical corpus for Spanish composed of about 860M tokens. The table below shows some basic statistics of the individual cleaned corpora:
55
+
56
 
57
  | Name | No. tokens | Description |
58
  |-----------------------------------------------------------------------------------------|-------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
59
  | [Medical crawler](https://zenodo.org/record/4561971#.YTtwM32xXbQ) | 745,705,946 | Crawler of more than 3,000 URLs belonging to Spanish biomedical and health domains. |
60
  | [Scielo](https://github.com/PlanTL-SANIDAD/SciELO-Spain-Crawler) | 60,007,289 | Publications written in Spanish crawled from the Spanish SciELO server in 2017. |
61
  | [BARR2_background](https://temu.bsc.es/BARR2/downloads/background_set.raw_text.tar.bz2) | 24,516,442 | Biomedical Abbreviation Recognition and Resolution (BARR2) containing Spanish clinical case study sections from a variety of clinical disciplines. |
62
+ | Wikipedia_life_sciences | 13,890,501 | Wikipedia articles crawled 04/01/2021 with the [Wikipedia API python library](https://pypi.org/project/Wikipedia-API/) starting from the "Ciencias\_de\_la\_vida" category up to a maximum of 5 subcategories. Multiple links to the same articles are then discarded to avoid repeating content. |
63
  | Patents | 13,463,387 | Google Patent in Medical Domain for Spain (Spanish). The accepted codes (Medical Domain) for Json files of patents are: "A61B", "A61C","A61F", "A61H", "A61K", "A61L","A61M", "A61B", "A61P". |
64
  | [EMEA](http://opus.nlpl.eu/download.php?f=EMEA/v3/moses/en-es.txt.zip) | 5,377,448 | Spanish-side documents extracted from parallel corpora made out of PDF documents from the European Medicines Agency. |
65
  | [mespen_Medline](https://zenodo.org/record/3562536#.YTt1fH2xXbR) | 4,166,077 | Spanish-side articles extracted from a collection of Spanish-English parallel corpus consisting of biomedical scientific literature. The collection of parallel resources are aggregated from the MedlinePlus source. |
66
  | PubMed | 1,858,966 | Open-access articles from the PubMed repository crawled in 2017. |
67
 
 
 
 
 
 
 
 
 
68
 
 
 
69
 
70
  ## Evaluation and results
71