Fine-tuning of bert-base-uncased for extractive Question Answering, delivered as
part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science,
Unit 2, Universidad Politécnica de Yucatán).
Model description
bert-base-uncased with a qa_outputs head that predicts, per token, the
probability of being the start and the end of the answer span within the given
context.
Training data
- Dataset: SQuAD v1.1
(
rajpurkar/squad), subsampled to 15,000 examples from the official train
split (as required by the assignment), with a 90/10 split used for training/
validation. SQuAD's official validation set (10,570 questions, never seen during
training) was reserved as the final test set.
- Preprocessing uses a sliding window (stride) since contexts can exceed the maximum
input length; each window is labeled with the start/end position of the answer
inside that window, or
(0, 0) if the answer does not fit.
Training procedure
Two adaptation methods were trained and compared; full fine-tuning is the delivered
model, since it is necessary (not optional) for this task — partial fine-tuning
underperforms by a wide margin.
| Hyperparameter |
Value |
| Base model |
bert-base-uncased (110M params) |
| Method |
Full fine-tuning (BERT body + head, jointly) |
| Learning rate (head) |
1e-3 |
| Learning rate (BERT body) |
2e-5 |
| Epochs |
3 |
| Batch size |
16 (train) / 32 (eval) |
| Seed |
42 |
| Trainable parameters |
108,893,186 |
| Training time |
13.5 min (single T4 GPU) |
Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT
body except its last 2 encoder layers (14,177,282 trainable params, 5.9 min).
Evaluation results
Metrics: official SQuAD Exact Match (EM) and word-level F1.
| Method |
Val EM |
Val F1 |
Test EM |
Test F1 |
| Partial fine-tuning (last 2 layers + head) |
39.53% |
55.05% |
45.91% |
59.45% |
| Full fine-tuning (delivered) |
59.27% |
73.79% |
71.16% |
81.06% |
The 21.61-point F1 gap on test is by far the largest gap of the four tasks in this
project (roughly 5x the gap seen in classification, NER, or POS), reflecting that
extractive QA requires jointly reasoning over the question and locating a span in
the context — something only 2 unfrozen layers cannot capture well.
Note on scale: this model was trained on 15k examples rather than the full 87.6k
SQuAD train set, so its 81.06 test F1 is below the ~88 F1 reported in the literature
for BERT-base trained on the full dataset — a result consistent with the reduced
amount of training data.
Intended uses & limitations
- Intended use: extractive question answering over short English passages, in
the style of SQuAD v1.1, for coursework/research.
- Limitations: trained on only 15k (of 87.6k) SQuAD examples, so accuracy is
meaningfully below full-dataset BERT-base benchmarks. Only answers questions whose
answer is a contiguous span present in the given context (no "no answer" case, as
in SQuAD v1.1, and no closed-book/generative QA). Single seed, 3 epochs, no
extensive hyperparameter search.
References
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of
Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805.
https://arxiv.org/abs/1810.04805
- Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+
Questions for Machine Comprehension of Text.
- Hugging Face. Fine-tune a pretrained model.
https://huggingface.co/docs/transformers/training
- Dataset: rajpurkar/squad