SAVRN
Search Contact SAVRN

Open-weight model · Token classification

bert-base-conll03-ner

by Dalila Ku Dalila-Ku/bert-base-conll03-ner

bert-base-conll03-ner is an open-weight model for token classification from Dalila Ku, released under Apache License 2.0. It has 109M parameters and a 512-token context. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

token-classification - bert - ner - lhoestq/conll2003 - precision - recall pipelinetag: token-classification basemodel: bert-base-uncased Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as part of assignment U2T01 (Adapting BERT…

Parameters109M
Context512
Weights435.6 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve bert-base-conll03-ner (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.

bert-base-conll03-ner on every accelerator the SAVRN Index prices, at every precision

Model Card

By Dalila Ku, published under apache-2.0, revision 0a95ad7b6d15.

token-classification - bert - ner - lhoestq/conll2003 - precision - recall pipelinetag: token-classification basemodel: bert-base-uncased Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 9 BIO-scheme entity tags (person, organization, location, miscellaneous, or none). (lhoestq/conll2003), using its own official train/validation/test splits. split a word into multiple subwords, the label is placed on the first…

Read Dalila Ku's full model card

Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán).

Model description

bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 9 BIO-scheme entity tags (person, organization, location, miscellaneous, or none).

Training data

  • Dataset: CoNLL-2003 (lhoestq/conll2003), using its own official train/validation/test splits.
  • Subword alignment: since the dataset labels whole words but the tokenizer can split a word into multiple subwords, the label is placed on the first subword of each word only; all other positions (continuation subwords, [CLS], [SEP], padding) get label -100, which CrossEntropyLoss ignores.

Training procedure

Two adaptation methods were trained and compared; full fine-tuning is the delivered model.

Hyperparameter Value
Base model bert-base-uncased (110M params)
Method Full fine-tuning (BERT body + head, jointly)
Learning rate (head) 1e-3
Learning rate (BERT body) 2e-5
Epochs 4
Batch size 32 (train) / 64 (eval)
Seed 42
Trainable parameters 108,898,569
Training time 4.4 min (single T4 GPU)

Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT body except its last 2 encoder layers (14,182,665 trainable params, 1.9 min).

Evaluation results

Metric: seqeval F1 (evaluated at the entity level, e.g. "New York" counts as one LOC entity rather than two tokens — important because the dominant "O" class would make per-token accuracy misleadingly high).

Method Val F1 Test F1 Test Precision Test Recall
Partial fine-tuning (last 2 layers + head) 88.75% 85.37% 83.58% 87.23%
Full fine-tuning (delivered) 94.27% 90.16% 89.36% 90.97%

The 4.79-point F1 gap on test is above the ±1–3 point noise margin. Note the consistent drop from validation to test in both methods — a known property of the CoNLL-2003 test split being harder than its validation split, not a sign of overfitting.

Intended uses & limitations

  • Intended use: named entity recognition (person, organization, location, misc) on English news-style text similar to CoNLL-2003, for coursework/research.
  • Limitations: trained only on CoNLL-2003 (Reuters newswire from the 1990s); may not generalize well to informal text, social media, or domains with different entity distributions (e.g. biomedical, legal). Single seed, 4 epochs, no extensive hyperparameter search. Inherits biases from the source corpus and from bert-base-uncased's pretraining data.

References

  • Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805. https://arxiv.org/abs/1810.04805
  • Hugging Face. Fine-tune a pretrained model. https://huggingface.co/docs/transformers/training
  • Dataset: lhoestq/conll2003

Configuration

Architecture
BertForTokenClassification
Context length (tokens)
512
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
30,522
Model type
bert

Identity and Version

Repository
Dalila-Ku/bert-base-conll03-ner
Publisher
Dalila Ku
Task
Token classification
Modality
Text
Library
Not stated by the source
Parameters
109M parameters
Languages
en
Revision
0a95ad7b6d1538887ec6c6014211edf47119ddd8
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

6 files, 436.3 MB in total. The weights are 1 file totalling 435.6 MB in safetensors.

Weights1 file · 435.6 MB
Configuration1 file · 1.1 KB
Tokenizer2 files · 711.8 KB
Documentation1 file · 3.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights435.6 MB 8a3b048266d1
config.jsonConfiguration1.1 KB —
README.mdDocumentation3.5 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer711.5 KB —
tokenizer_config.jsonTokenizer351 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
435.6 MB
Download from Dalila Ku

Released by Dalila Ku through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published435.6 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About bert-base-conll03-ner

How much GPU memory does bert-base-conll03-ner need?

About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run bert-base-conll03-ner on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use bert-base-conll03-ner commercially?

Yes. bert-base-conll03-ner is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is bert-base-conll03-ner's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Token classification

OpenMed-NER-AnatomyDetect-ElectraMed-109M

OpenMed

Specialized model for Anatomical Entity Recognition - Anatomical structures and body parts This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for anatomical entity recognition - anatomical structures and body parts. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications.…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-SpeciesDetect-ElectraMed-109M

OpenMed

Specialized model for Species Entity Recognition - Species and organism names This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for species entity recognition - species and organism names. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications. This model can identify and…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-GenomicDetect-PubMed-109M

OpenMed

Specialized model for Gene Entity Recognition - Gene-related entities This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for gene entity recognition - gene-related entities. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications. This model can identify and classify the…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-DiseaseDetect-ElectraMed-109M

OpenMed

Specialized model for Disease Entity Recognition - Disease entities from the BC5CDR dataset This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for disease entity recognition - disease entities from the bc5cdr dataset. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research applications.…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

OpenMed-NER-OrganismDetect-BioMed-109M

OpenMed

Specialized model for Species Entity Recognition - Species names from the Species-800 dataset This model is a state-of-the-art fine-tuned transformer engineered to deliver enterprise-grade accuracy for species entity recognition - species names from the species-800 dataset. This specialized model excels at identifying and extracting biomedical entities from clinical texts, research papers, and healthcare documents, enabling applications such as drug interaction detection, medication extraction from patient records, adverse event monitoring, literature mining for drug discovery, and biomedical knowledge graph construction with production-ready reliability for clinical and research…

Open weights apache-2.0 109M parameters 512 tokens transformers

Model · Token classification

bert-base-ud-ewt-pos

Dalila Ku

Fine-tuning of bert-base-uncased for Part-of-Speech (POS) tagging, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a token-level linear classification head over the last hidden state of every token, predicting one of 17 Universal POS (UPOS) tags (NOUN, VERB, ADJ, DET, PUNCT, etc.). (universal-dependencies/universaldependencies, config enewt). (universaldependencies/universaldependencies) has a typo; the correct organization name on the Hub uses a hyphen (universal-dependencies). - Same per-token label shape as NER, so the same subword-alignment convention is used: the label goes…

Open weights apache-2.0 109M parameters 512 tokens