BERT-base-uncased Fine-Tuned for Extractive Question Answering (SQuAD v1.1)
This repository contains the weights and tokenizer for bert-base-uncased fully fine-tuned on SQuAD v1.1 (Stanford Question Answering Dataset) for extractive question answering.
The model was adapted using end-to-end full fine-tuning with a differential two-learning-rate optimizer scheme as part of an academic comparative study on BERT adaptation paradigms.
Model Details
- Base Architecture: BERT-base (
bert-base-uncased)
- Model Type:
BertForQuestionAnswering
- Layers: 12 Transformer encoder blocks
- Hidden Dimension: 768
- Attention Heads: 12
- Parameters: 108,893,186 trainable parameters (~110M total)
- Framework: PyTorch & Hugging Face
transformers
Training Data & Preprocessing
The model was trained on a deterministic subsample of SQuAD v1.1 (rajpurkar/squad):
- Subsample Size: 15,000 question-context pairs sampled deterministically with seed
42 from the training split.
- Sliding Window Strategy: Long contexts exceeding BERT's sequence limit were tokenized using a sliding window:
- Maximum sequence length:
384 tokens
- Document stride:
128 tokens
- Truncation:
only_second (context tokens truncated, questions kept intact)
- Unanswerable / out-of-slice spans: Routed to index
[CLS] (position 0)
- Windowed Features: The 15,000 source examples expanded to 15,141 training feature windows.
Hyperparameters & Training Configuration
- Epochs: 2
- Per-Device Batch Size: 4 (to ensure memory stability on 384-token sequences without CUDA OOM)
- Total Training Steps: 7,570 optimizer steps
- Optimizer: Two-learning-rate AdamW (
build_two_lr_adamw):
- Task Head (QA outputs, 1,538 params): Learning rate =
1e-3
- Transformer Body (108,891,648 params): Learning rate =
2e-5
- Precision: FP16 mixed precision
- Random Seed: 42 (deterministic CuDNN backend)
- Wall-Clock Training Time: ~704.1 seconds (~11.7 minutes)
Evaluation Results
Evaluation was performed on the complete official SQuAD v1.1 validation split (10,570 examples, spanning 10,753 feature windows) using an optimized post-processing pipeline (top-15 start/end logit pruning, maximum answer span length of 30 tokens):
| Adaptation Method |
Trainable Parameters |
Parameter Ratio |
Training Time (s) |
Exact Match (%) |
F1 Score (%) |
| Feature-based (frozen body) |
1,538 |
0.001% |
171.5s |
16.23% |
25.26 |
| Partial fine-tuning (top-2 layers) |
14,177,282 |
13.02% |
235.2s |
49.60% |
62.80 |
| Full fine-tuning (this model) |
108,893,186 |
100.00% |
704.1s |
73.59% |
82.70% |
Sample Predictions
Below are sample outputs produced on the validation split:
[
{
"id": "56be4db0acb8001400a502ed",
"prediction": "Carolina Panthers",
"gold": ["Carolina Panthers", "Carolina Panthers", "Carolina Panthers"]
},
{
"id": "56be4db0acb8001400a502ee",
"prediction": "Santa Clara, California",
"gold": ["Santa Clara, California", "Levi's Stadium", "Levi's Stadium in the San Francisco Bay Area at Santa Clara, California."]
},
{
"id": "56be4db0acb8001400a502f0",
"prediction": "gold",
"gold": ["gold", "gold", "gold"]
},
{
"id": "56be8e613aeaaa14008c90d2",
"prediction": "February 7, 2016",
"gold": ["February 7, 2016", "February 7", "February 7, 2016"]
}
]
Intended Use
- Task: Extractive question answering in English. Given a question and a reference passage, the model extracts the continuous sub-string (span) that answers the question.
- Input: A query string (question) and a text passage (context).
- Target Audience: Researchers, students, and practitioners exploring transformer adaptation methods, resource-efficient training, or lightweight extractive QA deployments.
Quickstart / How to Use
from transformers import pipeline
qa_pipeline = pipeline(
"question-answering",
model="path_to_model/qa_full",
tokenizer="path_to_model/qa_full"
)
question = "Where was Super Bowl 50 played?"
context = "Super Bowl 50 took place at Levi's Stadium in Santa Clara, California."
result = qa_pipeline(question=question, context=context)
print(f"Answer: {result['answer']}")
print(f"Confidence score: {result['score']:.4f}")
# Output:
# Answer: Santa Clara, California
Limitations
- Extractive Only: The model identifies spans verbatim within the context passage. It cannot synthesize, paraphrase, or generate abstractive responses.
- SQuAD v1.1 Constraint (Answer-Present): SQuAD v1.1 assumes all questions have an answer in the provided text. The model does not support unanswerable question detection (as found in SQuAD v2.0).
- Subsampled Training: The model was trained on a 15,000-example subset of SQuAD v1.1 rather than the full ~87.6k training set. Consequently, while it achieves a strong 82.70% F1, training on the complete dataset would yield higher metrics (~88% F1 for BERT-base).
- Context Length Truncation: While sliding window inference allows covering long documents, answer spans that cross window seams or exceed 384 tokens without stride coverage may suffer degradation.
- Language & Casing: The model was pretrained on lowercased English text (
uncased) and is unsuitable for multilingual texts or applications sensitive to letter casing.
References
- Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 §5.3.
- Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. arXiv:1606.05250.
- Hugging Face Documentation. Fine-tuning a pretrained model. Hugging Face Transformers Training Documentation.