This model was fine-tuned on the AG News dataset (fancyzhx/agnews) for four-class news topic classification: The dataset was divided into 108,000 training examples, 12,000 validation examples, and 7,600 test examples. A random seed of 42 was used. This model is intended for English news topic classification into the four AG News categories: World, Sports, Business, and Sci/Tech. It was developed for educational purposes and experimentation with BERT adaptation methods. The model is trained on English news data and may not generalize well to other domains or languages. It only supports the four categories present in AG News. Performance on real-world data may differ from the reported…
Open-weight model · Text classification
bert-base-agnews-topic-classification
by Dalila Ku Dalila-Ku/bert-base-agnews-topic-classification
bert-base-agnews-topic-classification is an open-weight model for text classification from Dalila Ku, released under Apache License 2.0. It has 109M parameters and a 512-token context. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
Full fine-tuning of bert-base-uncased for 4-class news topic classification, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán).
Runs On
What it takes to serve bert-base-agnews-topic-classification (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 9, 2026.
Model Card
By Dalila Ku, published under apache-2.0, revision 7b63bc6101d8.
Full fine-tuning of bert-base-uncased for 4-class news topic classification, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán). bert-base-uncased with a linear classification head on top of the [CLS] token's last hidden state, mapping to 4 topic classes: World, Sports, Business, Sci/Tech. - 10% of the official train split was held out as validation; the official test split was left untouched. - 4 balanced classes. Two adaptation methods were trained and compared; full fine-tuning is the delivered model, since it beat the feature-based baseline by a margin well above the run-to-run noise floor (±1–3 points…
Read Dalila Ku's full model card
Full fine-tuning of bert-base-uncased for 4-class news topic classification, delivered
as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science,
Unit 2, Universidad Politécnica de Yucatán).
Model description
bert-base-uncased with a linear classification head on top of the [CLS] token's
last hidden state, mapping to 4 topic classes: World, Sports, Business, Sci/Tech.
Training data
- Dataset: AG News (
fancyzhx/ag_news) - 10% of the official train split was held out as validation; the official test split was left untouched.
- 4 balanced classes.
Training procedure
Two adaptation methods were trained and compared; full fine-tuning is the delivered model, since it beat the feature-based baseline by a margin well above the run-to-run noise floor (±1–3 points from seed variance).
| Hyperparameter | Value |
|---|---|
| Base model | bert-base-uncased (110M params) |
| Method | Full fine-tuning (BERT body + head, jointly) |
| Learning rate (head) | 1e-3 |
| Learning rate (BERT body) | 2e-5 |
| Epochs | 2 |
| Batch size | 32 (train) / 64 (eval) |
| Seed | 42 |
| Trainable parameters | 109,485,316 |
| Training time | 76.4 min (single T4 GPU) |
Compared alternative (not delivered): feature-based adaptation — BERT body fully
frozen, with a scikit-learn logistic regression trained on the frozen [CLS]
embeddings (0 trainable BERT parameters, 17.7 min).
Evaluation results
| Method | Val Accuracy | Val F1 | Test Accuracy | Test F1 |
|---|---|---|---|---|
| Feature-based (LogReg on frozen [CLS]) | 90.60% | 90.56% | 90.24% | 90.23% |
| Full fine-tuning (delivered) | 94.92% | 94.89% | 94.59% | 94.60% |
The 4.36-point F1 gap on test is above the ±1–3 point noise margin observed for this setup, so it is treated as a real effect of the adaptation method rather than seed variance.
Intended uses & limitations
- Intended use: topic classification of short English news text into the 4 AG News categories, for coursework/research and as a fine-tuning reference example.
- Limitations: trained and evaluated only on AG News; it will not generalize to
topics or writing styles outside that domain (e.g. non-English text, other news
taxonomies, social media text). Trained for 2 epochs on a single seed — not
hyperparameter-tuned for production use. Inherits any biases present in the AG
News corpus and in
bert-base-uncased's pretraining data.
References
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805. https://arxiv.org/abs/1810.04805
- Hugging Face. Fine-tune a pretrained model. https://huggingface.co/docs/transformers/training
- Dataset: fancyzhx/ag_news
Configuration
- Architecture
- BertForSequenceClassification
- Context length (tokens)
- 512
- Layers
- 12
- Hidden size
- 768
- Feed-forward size
- 3,072
- Attention heads
- 12
- Vocabulary size
- 30,522
- Model type
- bert
Identity and Version
- Repository
- Dalila-Ku/bert-base-agnews-topic-classification
- Publisher
- Dalila Ku
- Task
- Text classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- 109M parameters
- Languages
- en
- Revision
- 7b63bc6101d80dde9652a4ee4d0562f84d72ccb9
- First published
- 2026-09-27
- Last updated
- 2026-09-27
Files and Weights
6 files, 438.7 MB in total. The weights are 1 file totalling 438.0 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 438.0 MB | 84a2a4765bcc |
| config.json | Configuration | 1.0 KB | — |
| README.md | Documentation | 3.2 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 711.7 KB | — |
| tokenizer_config.json | Tokenizer | 351 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 438.0 MB
Released by Dalila Ku through its official repository on Hugging Face. Read the license.
Built From
- Derived from google-bert/bert-base-uncased
- Described by arXiv:1810.04805
- Trained on (disclosed) fancyzhx/ag_news
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 438.0 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About bert-base-agnews-topic-classification
How much GPU memory does bert-base-agnews-topic-classification need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run bert-base-agnews-topic-classification on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use bert-base-agnews-topic-classification commercially?
Yes. bert-base-agnews-topic-classification is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is bert-base-agnews-topic-classification's context length?
512 tokens, from the maximum position embeddings in its published configuration.
Similar Models
This model is a fine-tuned version of bert-base-uncased on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 32 - evalbatchsize: 64 - lrschedulertype: linear - numepochs: 2 - Transformers 5.16.1 - Pytorch 2.11.0+cu128 - Datasets 4.8.5 - Tokenizers 0.23.1
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
libraryname: transformers - autotrain - text-classification basemodel: google-bert/bert-base-uncased f1macro: 0.7533020080884588 f1micro: 0.7533333333333333 f1weighted: 0.7533020080884587 precisionmacro: 0.7551310982162045 precisionmicro: 0.7533333333333333 precisionweighted: 0.7551310982162046 recallmacro: 0.7533333333333333 recallmicro: 0.7533333333333333 recallweighted: 0.7533333333333333
1. bert-log-anomaly-detection is a BERT-based NLP model fine-tuned for single SQL transaction log anomaly detection. 2. The model classifies each database transaction log as either Normal or Anomaly, with the goal of supporting AI-powered fraud detection and cybersecurity monitoring systems. 3. This model was developed as part of the Samsung × KBTG Digital Fraud Cybersecurity Hackathon (Thailand) under the AI-Powered Fraud Detection & Prevention track. This model analyzes individual SQL database transaction logs and detects abnormal patterns that may indicate fraudulent, malicious, or suspicious behavior. - Developed by Waris Sripatoomrak, this model integrates with an n8n workflow to…
This model is a fine-tuned version of google/electra-base-discriminator on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 2e-05 - trainbatchsize: 8 - evalbatchsize: 8 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 16 - lrschedulertype: linear - lrschedulerwarmupsteps: 50 - numepochs: 10 - Transformers 5.2.0 - Pytorch 2.10.0+cu128 - Datasets 4.5.0 - Tokenizers 0.22.2