README.md
3.5 KB · 95 lines · markdown Raw
1 ---
2 license: mit
3 tags:
4 - generated_from_trainer
5 metrics:
6 - precision
7 - recall
8 - f1
9 - accuracy
10 model_index:
11 - name: bert-portuguese-ner-archive
12 results:
13 - task:
14 name: Token Classification
15 type: token-classification
16 metric:
17 name: Accuracy
18 type: accuracy
19 value: 0.9700325118974698
20 ---
21
22 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
23 should probably proofread and complete it, then remove this comment. -->
24
25 # bert-portuguese-ner
26
27 This model is a fine-tuned version of [neuralmind/bert-base-portuguese-cased](https://huggingface.co/neuralmind/bert-base-portuguese-cased)
28 It achieves the following results on the evaluation set:
29 - Loss: 0.1140
30 - Precision: 0.9147
31 - Recall: 0.9483
32 - F1: 0.9312
33 - Accuracy: 0.9700
34
35 ## Model description
36
37 This model was fine-tunned on token classification task (NER) on Portuguese archival documents. The annotated labels are: Date, Profession, Person, Place, Organization
38
39 ### Datasets
40
41 All the training and evaluation data is available at: http://ner.epl.di.uminho.pt/
42
43
44 ### Training hyperparameters
45
46 The following hyperparameters were used during training:
47 - learning_rate: 2e-05
48 - train_batch_size: 16
49 - eval_batch_size: 16
50 - seed: 42
51 - optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
52 - lr_scheduler_type: linear
53 - num_epochs: 4
54
55 ### Training results
56
57 | Training Loss | Epoch | Step | Validation Loss | Precision | Recall | F1 | Accuracy |
58 |:-------------:|:-----:|:----:|:---------------:|:---------:|:------:|:------:|:--------:|
59 | No log | 1.0 | 192 | 0.1438 | 0.8917 | 0.9392 | 0.9148 | 0.9633 |
60 | 0.2454 | 2.0 | 384 | 0.1222 | 0.8985 | 0.9417 | 0.9196 | 0.9671 |
61 | 0.0526 | 3.0 | 576 | 0.1098 | 0.9150 | 0.9481 | 0.9312 | 0.9698 |
62 | 0.0372 | 4.0 | 768 | 0.1140 | 0.9147 | 0.9483 | 0.9312 | 0.9700 |
63
64
65 ### Framework versions
66
67 - Transformers 4.10.0.dev0
68 - Pytorch 1.9.0+cu111
69 - Datasets 1.10.2
70 - Tokenizers 0.10.3
71 ### Citation
72
73 ```bibtex
74
75 @Article{make4010003,
76 AUTHOR = {Cunha, Luís Filipe and Ramalho, José Carlos},
77 TITLE = {NER in Archival Finding Aids: Extended},
78 JOURNAL = {Machine Learning and Knowledge Extraction},
79 VOLUME = {4},
80 YEAR = {2022},
81 NUMBER = {1},
82 PAGES = {42--65},
83 URL = {https://www.mdpi.com/2504-4990/4/1/3},
84 ISSN = {2504-4990},
85 ABSTRACT = {The amount of information preserved in Portuguese archives has increased over the years. These documents represent a national heritage of high importance, as they portray the country&rsquo;s history. Currently, most Portuguese archives have made their finding aids available to the public in digital format, however, these data do not have any annotation, so it is not always easy to analyze their content. In this work, Named Entity Recognition solutions were created that allow the identification and classification of several named entities from the archival finding aids. These named entities translate into crucial information about their context and, with high confidence results, they can be used for several purposes, for example, the creation of smart browsing tools by using entity linking and record linking techniques. In order to achieve high result scores, we annotated several corpora to train our own Machine Learning algorithms in this context domain. We also used different architectures, such as CNNs, LSTMs, and Maximum Entropy models. Finally, all the created datasets and ML models were made available to the public with a developed web platform, NER@DI.},
86 DOI = {10.3390/make4010003}
87 }
88
89
90
91
92 ```
93
94
95