Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as
part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science,
Unit 2, Universidad Politécnica de Yucatán).
Model description
bert-base-uncased with a token-level linear classification head over the last
hidden state of every token, predicting one of 9 BIO-scheme entity tags (person,
organization, location, miscellaneous, or none).
Training data
- Dataset: CoNLL-2003
(
lhoestq/conll2003), using its own official train/validation/test splits.
- Subword alignment: since the dataset labels whole words but the tokenizer can
split a word into multiple subwords, the label is placed on the first subword of
each word only; all other positions (continuation subwords,
[CLS], [SEP],
padding) get label -100, which CrossEntropyLoss ignores.
Training procedure
Two adaptation methods were trained and compared; full fine-tuning is the delivered
model.
| Hyperparameter |
Value |
| Base model |
bert-base-uncased (110M params) |
| Method |
Full fine-tuning (BERT body + head, jointly) |
| Learning rate (head) |
1e-3 |
| Learning rate (BERT body) |
2e-5 |
| Epochs |
4 |
| Batch size |
32 (train) / 64 (eval) |
| Seed |
42 |
| Trainable parameters |
108,898,569 |
| Training time |
4.4 min (single T4 GPU) |
Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT
body except its last 2 encoder layers (14,182,665 trainable params, 1.9 min).
Evaluation results
Metric: seqeval F1 (evaluated at the entity level, e.g. "New York" counts as one
LOC entity rather than two tokens — important because the dominant "O" class would
make per-token accuracy misleadingly high).
| Method |
Val F1 |
Test F1 |
Test Precision |
Test Recall |
| Partial fine-tuning (last 2 layers + head) |
88.75% |
85.37% |
83.58% |
87.23% |
| Full fine-tuning (delivered) |
94.27% |
90.16% |
89.36% |
90.97% |
The 4.79-point F1 gap on test is above the ±1–3 point noise margin. Note the
consistent drop from validation to test in both methods — a known property of the
CoNLL-2003 test split being harder than its validation split, not a sign of
overfitting.
Intended uses & limitations
- Intended use: named entity recognition (person, organization, location, misc)
on English news-style text similar to CoNLL-2003, for coursework/research.
- Limitations: trained only on CoNLL-2003 (Reuters newswire from the 1990s); may
not generalize well to informal text, social media, or domains with different
entity distributions (e.g. biomedical, legal). Single seed, 4 epochs, no extensive
hyperparameter search. Inherits biases from the source corpus and from
bert-base-uncased's pretraining data.
References
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of
Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
https://arxiv.org/abs/1810.04805
- Hugging Face. Fine-tune a pretrained model.
https://huggingface.co/docs/transformers/training
- Dataset: lhoestq/conll2003