SBB
/

Text Classification
French
annif
glam
lam
File size: 13,988 Bytes
5031e0b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
---
language:
- fr
tags:
- annif
- glam
- lam
pipeline_tag: text-classification
license: apache-2.0
dataset: "[published on Zenodo](https://zenodo.org/doi/10.5281/zenodo.13301020)"
library_name: annif

---

# Model Card for ark-omikuji-fre-title-content


An [Annif](https://annif.org/) model, trained on historical titles and additional catalogue metadata for automatic subject indexing tasks. It classifies a given text into one or multiple subjects from the “Alter Realkatalog” ([ARK](https://staatsbibliothek-berlin.de/recherche/kataloge-der-staatsbibliothek/alter-realkatalog-und-historische-systematik)) classification system. The model was developed in the research project [Human.Machine.Culture](https://mmk.sbb.berlin/?lang=en) at Staatsbibliothek zu Berlin – Berlin State Library (SBB).

# Table of Contents

* [Model Card for ark-omikuji-fre-title-content](#model-card-for-ark-omikuji-fre-title-content)  
* [Table of Contents](#table-of-contents)  
* [Model Details](#model-details)  
  * [Model Description](#model-description)  
* [Uses](#uses)  
  * [Direct Use](#direct-use)  
  * [Downstream Use](#downstream-use)  
  * [Out-of-Scope Use](#out-of-scope-use)  
* [Bias, Risks, and Limitations](#bias-risks-and-limitations)  
  * [Recommendations](#recommendations)  
* [Training Details](#training-details)  
  * [Training Data](#training-data)  
  * [Training Procedure](#training-procedure)  
    * [Preprocessing](#preprocessing)  
    * [Speeds, Sizes, Times](#speeds-sizes-times)  
    * [Training hyperparameters](#training-hyperparameters)  
    * [Training results](#training-results)  
* [Evaluation](#evaluation)  
  * [Testing Data, Factors and Metrics](#testing-data-factors-and-metrics)  
    * [Testing Data](#testing-data)  
    * [Metrics](#metrics)  
* [Environmental Impact](#environmental-impact)  
* [Technical Specifications](#technical-specifications)  
  * [Model Architecture and Objective](#model-architecture-and-objective)  
    * [Software](#software)  
* [Model Card Authors](#model-card-authors)  
* [Model Card Contact](#model-card-contact)  
* [How to Get Started with the Model](#how-to-get-started-with-the-model)

# Model Details

## Model Description

An [Annif](https://annif.org/) model, trained on historical titles and additional catalogue metadata for automatic subject indexing tasks. Subject indexing is a classical library task, aiming at describing the content of a resource. The model is intended to be used to automatically classify historical texts with a historical classification system developed in the 19th century to enrich those texts that have not been classified manually so far. For each input text, the model outputs one or multiple subjects from the [ARK](https://staatsbibliothek-berlin.de/recherche/kataloge-der-staatsbibliothek/alter-realkatalog-und-historische-systematik) classification system. It is part of a collection of 5 models, created with the help of the Annif toolkit which addresses this task of automated subject indexing.

* **Developed by:** [Sophie Schneider](mailto:[email protected])  
* **Shared by:** [Staatsbibliothek zu Berlin – Berlin State Library](https://huggingface.co/SBB)  
* **Model type:** tree-based  
* **Language(s) (NLP):** fr (French)  
* **License:** apache-2.0

# Uses

## Direct Use

This model can directly be used to automatically classify historical texts with the ARK classification scheme. It is intended to be used together with the Annif automated subject indexing toolkit version 0.60.0-1.1.0.

## Downstream Use

Other/downstream uses outside of the Annif setting described above are not intended but also not excluded.

## Out-of-Scope Use

The model is not intended for use on contemporary texts, as language and concept drifts will probably influence the results negatively and some terms from the vocabulary are not appropriate for more recent publications.

# Bias, Risks, and Limitations

Since we are dealing with historical texts and especially with a historical classification system such as the ARK, the classes suggested for an input text might not be suitable for today’s understanding or might even be of an unethical nature (for more information, see also [the datasheet accompanying the Metadata of the “Alter Realkatalog” (ARK of Berlin State Library)](https://zenodo.org/doi/10.5281/zenodo.12783813) and the [Datasheet for Machine-Readable Vocabulary Files of the ARK (Alter Realkatalog)](https://zenodo.org/doi/10.5281/zenodo.13301020)).

Another limitation when using the ARK as a vocabulary arises from its hierarchical structure: the system contains multiple classes that do not describe the same content (e.g. they have different IDs) but are labeled identical (same name). This is due to the fact that the manual inspection of the whole path to a class, including all the upper level classes leading to it, delivers additional information that allows for contextualization. As duplicate label names seem to be \- as expected \- a challenge for lexical methods, we decided to focus on statistical rather than lexical algorithms.

## Recommendations

Considering that the ARK classification scheme consists of 225.691 classes in total and that there is only limited training material at hand plus an overall unbalanced distribution of classes, we might describe this task as an Extreme Multi-Label Classification (XMC) problem. We recommend being aware of this limitation and, if available, use additional training data to improve the current model’s performance (e.g. by running `annif learn`, see [CLI commands documentation](https://annif.readthedocs.io/en/v1.1.0/source/commands.html\#annif-learn)).

# Training Details

## Training Data

Training data include a selection of metadata fields that were obtained via CBS export: 

* Lehmann, J., & Schneider, S. (2024). Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB) (Version 1\) \[Data set\]. Zenodo. [https://doi.org/10.5281/zenodo.12783813](https://doi.org/10.5281/zenodo.12783813) 

The following title and content data fields have been extracted and combined from this dataset: 

* "Abweichender Titel" ([4212](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4212\&katalog=Standard))  
* "Abweichender Titel (Sucheinstieg)" ([3260](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=3260\&katalog=Standard))  
* "Ansetzungssachtitel" ([3220](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=3220\&katalog=Standard))  
* "Einheitssachtitel des beigefügten oder kommentierten Werkes" ([4210](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4210\&katalog=Standard))  
* "Frühere/frühester Haupttitel (nur für fortlaufende und integrierende Ressourcen)" ([4213](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4213\&katalog=Standard))  
* "Gesamttitel der Reproduktion" ([4110](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4110\&katalog=Standard))  
* "Gesamttitel der fortlaufenden Ressource" ([4170](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4170\&katalog=Standard))  
* "Gesamttitel der mehrteiligen Monografie" ([4150](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4150\&katalog=Standard))  
* "Haupttitel, Titelzusatz, Verantwortlichkeitsangabe" ([4000](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4000\&katalog=Standard))  
* "Normierter Zeitschriftenkurztitel" ([3232](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=3232\&katalog=Standard))  
* "Paralleltitel, paralleler Titelzusatz, parallele Verantwortlichkeitsangabe" ([4002](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4002\&katalog=Standard))  
* "Titelkonkordanzen" ([4245](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4245\&katalog=Standard))  
* "Titelzusätze und Verantwortlichkeitsangabe zur gesamten Vorlage" ([4011](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4011\&katalog=Standard))  
* "Weitere Titel etc. bei Zusammenstellungen" ([4010](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4010\&katalog=Standard))  
* "Weiterer Werktitel und sonstige unterscheidende Merkmale" ([3211](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=3211\&katalog=Standard))  
* "Werktitel und sonstige unterscheidende Merkmale des Werks" ([3210](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=3210\&katalog=Standard))  
* "Zusätzliche Sucheinstiege" ([4200](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4200\&katalog=Standard))  
* "Veröffentlichungsart und Inhalt" ([1140](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=1140\&katalog=Standard))  
* "Sonstige Anmerkungen" ([4201](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4201\&katalog=Standard))  
* "Zusammenfassende Register" ([4203](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4203\&katalog=Standard))  
* "Inhaltliche Zusammenfassung" ([4207](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=4207\&katalog=Standard) bzw. [9000](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=9000\&katalog=Standard))  
* "Einleitender Text" ([7124](https://swbtools.bsz-bw.de/cgi-bin/k10plushelp.pl?cmd=kat\&val=7124\&katalog=Standard))

The vocabulary files themselves can be found here:
* Schneider, S., & Lehmann, J. (2024). Machine-Readable Vocabulary Files of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13301020

## Training Procedure

Training procedure includes loading the ARK vocabulary (see [Datasheet for Machine-Readable Vocabulary Files of the ARK (Alter Realkatalog)](https://zenodo.org/doi/10.5281/zenodo.13301020)) into Annif and training the [Omikuji backend](https://github.com/NatLibFi/Annif/wiki/Backend%3A-Omikuji) with the help of our training data. Further aspects on technical specifications can be found in the section [Training hyperparameters](#training-hyperparameters).

### Preprocessing

Besides merging and transforming the data described under [Training Data](\#training-data) to fit the [corpus formats](https://github.com/NatLibFi/Annif/wiki/Corpus-formats) accepted by Annif, no further preprocessing of natural language or similar has been performed. 

### Speeds, Sizes, Times

Training takes from several minutes to a few hours on a V100, depending on the choice of dataset and algorithm as well as hyperparameter settings.

### Training hyperparameters

For some of the ARK Annif models, a slight hyperparameter optimization has been conducted to identify the final hyperparameter settings stated below.

hyperparameter configuration (as it needs to be stated in the Annif `projects.cfg` file):
```
[ark-omikuji-fre-title-content]
name=ARK-DE-18 Omikuji
language=fr
backend=omikuji
analyzer=snowball(french)
vocab=arktsv
cluster_k=2
collapse_every_n_layers=5
```

### Training results

* Precision (`--limit` 1, `--threshold` 0): 0.4228
* Recall (`--limit` 1, `--threshold` 0): 0.4044
* F1 (`--limit` 1, `--threshold` 0): 0.4103
* NDCG (`--limit` 1, `--threshold` 0): 0.4169
* F1@5: 0.1885
* NDCG@5: 0.4939

# Evaluation

## Testing Data, Factors and Metrics

### Testing Data

The dataset is described under [Training Data](#training-data). It was split into smaller subsets used for training, testing and validating (80%/10%/10% split). 

### Metrics

Model performance has been evaluated based on the following metrics: Precision, Recall, F1 and NDCG. These are standard metrics for machine learning and more specifically automatic subject indexing tasks and are directly provided in Annif by calling the `annif eval` statement. Evaluation parameters (`--limit` = maximum number of results to return; `--threshold` = minimum confidence for a suggestion to be considered) have been optimized before using the validation dataset and affect the results accordingly. We also state F1@5 and NDCG@5 scores reached without any evaluation parameters.

# Environmental Impact

Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact\#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).

* **Hardware Type:** V100  
* **Hours used:** 0,5-5 hours  
* **Cloud Provider:** No cloud.  
* **Compute Region:** Germany.

# Technical Specifications

## Model Architecture and Objective

See [Annif](https://github.com/NatLibFi/Annif) and [Omikuji](https://github.com/tomtung/omikuji) repositories on Github. Omikuji is an implementation of Partitioned Label Trees (Prabhu et al., 2018): 
* Y. Prabhu, A. Kag, S. Harsola, R. Agrawal, and M. Varma, “Parabel: Partitioned Label Trees for Extreme Classification with Application to Dynamic Search Advertising,” in Proceedings of the 2018 World Wide Web Conference, 2018, pp. 993–1002.

### Software

To run this model, Annif version 0.60.0 or higher (min. up to 1.1.0) must be installed.

# Model Card Authors

[Sophie Schneider](mailto:[email protected]) and [Jörg Lehmann](mailto:[email protected])

# Model Card Contact

Questions and comments about the model can be directed to Sophie Schneider at [email protected], questions and comments about the model card can be directed to Jörg Lehmann at [email protected].

# How to Get Started with the Model

Follow the Annif [Getting Started](https://github.com/NatLibFi/Annif/wiki/Getting-started) page to set up and run Annif. Create a projects.cfg file (see section [Training hyperparameters](#training-hyperparameters) for details on the specific project configuration), load the ARK vocabulary (see [Datasheet for Machine-Readable Vocabulary Files of the ARK (Alter Realkatalog)](https://zenodo.org/doi/10.5281/zenodo.13301020)) via `annif load-vocab` command and copy the model folder over to `data/projects`.