SAVRN
Search Contact SAVRN

Open-weight model · Token classification

bert-base-ud-ewt-pos

by Dalila Ku Dalila-Ku/bert-base-ud-ewt-pos

bert-base-ud-ewt-pos is an open-weight model for token classification from Dalila Ku, released under Apache License 2.0. It has 109M parameters and a 512-token context. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán).

Parameters109M
Context512
Weights435.6 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve bert-base-ud-ewt-pos (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.

bert-base-ud-ewt-pos on every accelerator the SAVRN Index prices, at every precision

Model Card

By Dalila Ku, published under apache-2.0, revision 5ad4de477100.

Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 17 Universal POS (UPOS) tags (NOUN, VERB, ADJ, DET, PUNCT, etc.). (universal-dependencies/universaldependencies, config enewt). (universaldependencies/universaldependencies) has a typo; the correct organization name on the Hub uses a hyphen (universal-dependencies). - Same per-token label shape as NER, so the same subword-alignment convention is used: the label goes…

Read Dalila Ku's full model card

Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán).

Model description

bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 17 Universal POS (UPOS) tags (NOUN, VERB, ADJ, DET, PUNCT, etc.).

Training data

  • Dataset: UD English EWT (universal-dependencies/universal_dependencies, config en_ewt).
  • Note: the dataset ID given in the original assignment prompt (universaldependencies/universal_dependencies) has a typo; the correct organization name on the Hub uses a hyphen (universal-dependencies).
  • Same per-token label shape as NER, so the same subword-alignment convention is used: the label goes on the first subword of each word, -100 everywhere else.

Training procedure

Two adaptation methods were trained and compared; full fine-tuning is the delivered model.

Hyperparameter Value
Base model bert-base-uncased (110M params)
Method Full fine-tuning (BERT body + head, jointly)
Learning rate (head) 1e-3
Learning rate (BERT body) 2e-5
Epochs 4
Batch size 32 (train) / 64 (eval)
Seed 42
Trainable parameters ~108,900,000
Training time 6.3 min (single T4 GPU)

Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT body except its last 2 encoder layers (~14,190,000 trainable params, 2.1 min).

Evaluation results

Accuracy is the primary metric (POS has no multi-token "entity" concept, so per-token accuracy is directly meaningful); macro F1 is reported as a secondary metric so rare tags (INTJ, SYM, X) are not hidden behind frequent ones (NOUN, PUNCT).

Method Test Accuracy Test Macro F1
Partial fine-tuning (last 2 layers + head) 95.59% 89.19%
Full fine-tuning (delivered) 97.62% 93.81%

The gap between methods depends heavily on the metric: only 2.03 points in accuracy, but 4.62 points in macro F1 — partial fine-tuning already handles frequent tags well but falls behind specifically on rare ones.

Intended uses & limitations

  • Intended use: part-of-speech tagging of English text following the Universal Dependencies UPOS tagset, for coursework/research.
  • Limitations: trained only on UD English EWT (web text domain); performance on rare tags (INTJ, SYM, X) is weaker than on frequent ones, and this may be more pronounced on out-of-domain text. Single seed, 4 epochs, no extensive hyperparameter search.

References

  • Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805. https://arxiv.org/abs/1810.04805
  • Hugging Face. Fine-tune a pretrained model. https://huggingface.co/docs/transformers/training
  • Dataset: universal-dependencies/universal_dependencies (config en_ewt)

Configuration

Architecture
BertForTokenClassification
Context length (tokens)
512
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
30,522
Model type
bert

Identity and Version

Repository
Dalila-Ku/bert-base-ud-ewt-pos
Publisher
Dalila Ku
Task
Token classification
Modality
Text
Library
Not stated by the source
Parameters
109M parameters
Languages
en
Revision
5ad4de477100f134e12334075a94f17d5716e964
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

6 files, 436.4 MB in total. The weights are 1 file totalling 435.6 MB in safetensors.

Weights1 file · 435.6 MB
Configuration1 file · 1.4 KB
Tokenizer2 files · 711.8 KB
Documentation1 file · 3.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights435.6 MB 253a811851b4
config.jsonConfiguration1.4 KB —
README.mdDocumentation3.5 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer711.5 KB —
tokenizer_config.jsonTokenizer351 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
435.6 MB
Download from Dalila Ku

Released by Dalila Ku through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published435.6 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About bert-base-ud-ewt-pos

How much GPU memory does bert-base-ud-ewt-pos need?

About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run bert-base-ud-ewt-pos on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use bert-base-ud-ewt-pos commercially?

Yes. bert-base-ud-ewt-pos is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is bert-base-ud-ewt-pos's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Token classification

bert-base-conll03-ner

Dalila Ku

token-classification - bert - ner - lhoestq/conll2003 - precision - recall pipelinetag: token-classification basemodel: bert-base-uncased Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 9 BIO-scheme entity tags (person, organization, location, miscellaneous, or none). (lhoestq/conll2003), using its own official train/validation/test splits. split a word into multiple subwords, the label is placed on the first…

Open weights apache-2.0 109M parameters 512 tokens

Model · Token classification

OpenMed-NER-AnatomyDetect-ElectraMed-109M

OpenMed

Specialized model for Anatomical Entity Recognition - Anatomical structures and body parts This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for anatomical entity recognition - anatomical structures and body parts. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications.…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-SpeciesDetect-ElectraMed-109M

OpenMed

Specialized model for Species Entity Recognition - Species and organism names This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for species entity recognition - species and organism names. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications. This model can identify and…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-GenomicDetect-PubMed-109M

OpenMed

Specialized model for Gene Entity Recognition - Gene-related entities This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for gene entity recognition - gene-related entities. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications. This model can identify and classify the…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-DiseaseDetect-ElectraMed-109M

OpenMed

Specialized model for Disease Entity Recognition - Disease entities from the BC5CDR dataset This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for disease entity recognition - disease entities from the bc5cdr dataset. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications.…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-OrganismDetect-BioMed-109M

OpenMed

Specialized model for Species Entity Recognition - Species names from the Species-800 dataset This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for species entity recognition - species names from the species-800 dataset. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research…

Open weights apache-2.0 109M parameters 512 tokens transformers