token-classification - bert - ner - lhoestq/conll2003 - precision - recall pipelinetag: token-classification basemodel: bert-base-uncased Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 9 BIO-scheme entity tags (person, organization, location, miscellaneous, or none). (lhoestq/conll2003), using its own official train/validation/test splits. split a word into multiple subwords, the label is placed on the first…
Open-weight model · Token classification
bert-base-ud-ewt-pos
by Dalila Ku Dalila-Ku/bert-base-ud-ewt-pos
bert-base-ud-ewt-pos is an open-weight model for token classification from Dalila Ku, released under Apache License 2.0. It has 109M parameters and a 512-token context. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán).
Runs On
What it takes to serve bert-base-ud-ewt-pos (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.
bert-base-ud-ewt-pos on every accelerator the SAVRN Index prices, at every precision
Model Card
By Dalila Ku, published under apache-2.0, revision 5ad4de477100.
Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 17 Universal POS (UPOS) tags (NOUN, VERB, ADJ, DET, PUNCT, etc.). (universal-dependencies/universaldependencies, config enewt). (universaldependencies/universaldependencies) has a typo; the correct organization name on the Hub uses a hyphen (universal-dependencies). - Same per-token label shape as NER, so the same subword-alignment convention is used: the label goes…
Read Dalila Ku's full model card
Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as
part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science,
Unit 2, Universidad Politécnica de Yucatán).
Model description
bert-base-uncased with a token-level linear classification head over the last
hidden state of every token, predicting one of 17 Universal POS (UPOS) tags (NOUN,
VERB, ADJ, DET, PUNCT, etc.).
Training data
- Dataset: UD English EWT
(
universal-dependencies/universal_dependencies, configen_ewt). - Note: the dataset ID given in the original assignment prompt
(
universaldependencies/universal_dependencies) has a typo; the correct organization name on the Hub uses a hyphen (universal-dependencies). - Same per-token label shape as NER, so the same subword-alignment convention is
used: the label goes on the first subword of each word,
-100everywhere else.
Training procedure
Two adaptation methods were trained and compared; full fine-tuning is the delivered model.
| Hyperparameter | Value |
|---|---|
| Base model | bert-base-uncased (110M params) |
| Method | Full fine-tuning (BERT body + head, jointly) |
| Learning rate (head) | 1e-3 |
| Learning rate (BERT body) | 2e-5 |
| Epochs | 4 |
| Batch size | 32 (train) / 64 (eval) |
| Seed | 42 |
| Trainable parameters | ~108,900,000 |
| Training time | 6.3 min (single T4 GPU) |
Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT body except its last 2 encoder layers (~14,190,000 trainable params, 2.1 min).
Evaluation results
Accuracy is the primary metric (POS has no multi-token "entity" concept, so per-token accuracy is directly meaningful); macro F1 is reported as a secondary metric so rare tags (INTJ, SYM, X) are not hidden behind frequent ones (NOUN, PUNCT).
| Method | Test Accuracy | Test Macro F1 |
|---|---|---|
| Partial fine-tuning (last 2 layers + head) | 95.59% | 89.19% |
| Full fine-tuning (delivered) | 97.62% | 93.81% |
The gap between methods depends heavily on the metric: only 2.03 points in accuracy, but 4.62 points in macro F1 — partial fine-tuning already handles frequent tags well but falls behind specifically on rare ones.
Intended uses & limitations
- Intended use: part-of-speech tagging of English text following the Universal Dependencies UPOS tagset, for coursework/research.
- Limitations: trained only on UD English EWT (web text domain); performance on rare tags (INTJ, SYM, X) is weaker than on frequent ones, and this may be more pronounced on out-of-domain text. Single seed, 4 epochs, no extensive hyperparameter search.
References
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805. https://arxiv.org/abs/1810.04805
- Hugging Face. Fine-tune a pretrained model. https://huggingface.co/docs/transformers/training
- Dataset: universal-dependencies/universal_dependencies
(config
en_ewt)
Configuration
- Architecture
- BertForTokenClassification
- Context length (tokens)
- 512
- Layers
- 12
- Hidden size
- 768
- Feed-forward size
- 3,072
- Attention heads
- 12
- Vocabulary size
- 30,522
- Model type
- bert
Identity and Version
- Repository
- Dalila-Ku/bert-base-ud-ewt-pos
- Publisher
- Dalila Ku
- Task
- Token classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- 109M parameters
- Languages
- en
- Revision
- 5ad4de477100f134e12334075a94f17d5716e964
- First published
- 2026-09-27
- Last updated
- 2026-09-27
Files and Weights
6 files, 436.4 MB in total. The weights are 1 file totalling 435.6 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 435.6 MB | 253a811851b4 |
| config.json | Configuration | 1.4 KB | — |
| README.md | Documentation | 3.5 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 711.5 KB | — |
| tokenizer_config.json | Tokenizer | 351 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 435.6 MB
Released by Dalila Ku through its official repository on Hugging Face. Read the license.
Built From
- Derived from google-bert/bert-base-uncased
- Described by arXiv:1810.04805
- Trained on (disclosed) universal-dependencies/universal_dependencies
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 435.6 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About bert-base-ud-ewt-pos
How much GPU memory does bert-base-ud-ewt-pos need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run bert-base-ud-ewt-pos on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use bert-base-ud-ewt-pos commercially?
Yes. bert-base-ud-ewt-pos is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is bert-base-ud-ewt-pos's context length?
512 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Specialized model for Anatomical Entity Recognition - Anatomical structures and body parts This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for anatomical entity recognition - anatomical structures and body parts. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications.…
Specialized model for Species Entity Recognition - Species and organism names This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for species entity recognition - species and organism names. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications. This model can identify and…
Specialized model for Gene Entity Recognition - Gene-related entities This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for gene entity recognition - gene-related entities. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications. This model can identify and classify the…
Specialized model for Disease Entity Recognition - Disease entities from the BC5CDR dataset This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for disease entity recognition - disease entities from the bc5cdr dataset. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications.…
Specialized model for Species Entity Recognition - Species names from the Species-800 dataset This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for species entity recognition - species names from the species-800 dataset. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research…