SAVRN
Search Contact SAVRN

Open-weight model · Text classification

ai-agent-sme-sentiment-mbert

by Atsushi Hatakeyama AtsushiHatake/ai-agent-sme-sentiment-mbert

ai-agent-sme-sentiment-mbert is an open-weight model for text classification from Atsushi Hatakeyama, released under Apache License 2.0. It has 178M parameters and a 512-token context. At 16-bit it needs about 0.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

This is a three-class Japanese/English sentiment classifier for short hospitality reviews, fine-tuned from google-bert/bert-base-multilingual-cased. It categorizes each review as negative, neutral, or positive.

Parameters178M
Context512
Weights711.4 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve ai-agent-sme-sentiment-mbert (178M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.4 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

ai-agent-sme-sentiment-mbert on every accelerator the SAVRN Index prices, at every precision

Model Card

By Atsushi Hatakeyama, published under apache-2.0, revision 5c7888401e6b.

This is a three-class Japanese/English sentiment classifier for short hospitality reviews, fine-tuned from google-bert/bert-base-multilingual-cased. It categorizes each review as negative, neutral, or positive. The model predicts one overall label for a short Japanese or English hospitality It is intended for research, prototyping, and human-reviewed analytics. It must not be used as the sole basis for consequential business, employment, moderation, or customer-service decisions. Applications using this model should allow operators to review and correct its predictions. The model was fine-tuned on 480 synthetic bilingual reviews: 240 Japanese and 240 English, balanced across the three…

Read Atsushi Hatakeyama's full model card

Hospitality Review Sentiment mBERT

This is a three-class Japanese/English sentiment classifier for short hospitality reviews, fine-tuned from google-bert/bert-base-multilingual-cased. It categorizes each review as negative, neutral, or positive.

Intended use

The model predicts one overall label for a short Japanese or English hospitality review:

ID Label
0 negative
1 neutral
2 positive

It is intended for research, prototyping, and human-reviewed analytics. It must not be used as the sole basis for consequential business, employment, moderation, or customer-service decisions. Applications using this model should allow operators to review and correct its predictions.

Training data

The model was fine-tuned on 480 synthetic bilingual reviews: 240 Japanese and 240 English, balanced across the three labels. No real customer reviews were collected or used.

The data was produced by a deterministic, template-based generator. Training and test splits use disjoint sentence templates, but they share the same small set of polarity-bearing fragments. This shared vocabulary can substantially inflate performance compared with genuinely novel reviews.

Evaluation

The exact weights in this directory were re-evaluated on the 240-item synthetic test split (120 Japanese and 120 English; 80 items per class):

Metric Result
Macro-F1 1.000
Japanese macro-F1 1.000
English macro-F1 1.000
Negative F1 1.000
Neutral F1 1.000
Positive F1 1.000

The published weights achieved macro-F1 1.000 on a 240-item synthetic test split. Training and test templates were disjoint, but polarity-bearing fragment vocabulary was shared. The result therefore measures performance within the generator's narrow distribution and must not be interpreted as real-world deployment accuracy.

The zero-shot comparison baseline achieved macro-F1 0.664 on the same split.

Known limitations

  • The model has not been evaluated on an independently annotated set of genuine restaurant reviews.
  • Synthetic training and test text has narrow vocabulary and simple sentence structure.
  • Japanese/English code-switching and aspect-level sentiment were not evaluated.
  • A single overall label can conceal mixed opinions such as positive food but negative service.
  • Softmax confidence is not a calibrated probability of correctness on real reviews.
  • The unseen short review ramen was good is incorrectly classified as negative with approximately 0.966 confidence by these weights. In the synthetic training vocabulary, good occurs inside the neutral phrase neither good nor bad, illustrating generator-specific lexical learning.

Example

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="AtsushiHatake/ai-agent-sme-sentiment-mbert",
)

print(classifier("料理がおいしく、スタッフも親切でした。"))  # Japanese
print(classifier("The food was delicious, and the staff were very friendly."))  # English

# Response examples:
# [{'label': 'positive', 'score': 0.9935460686683655}]
# [{'label': 'positive', 'score': 0.993598461151123}]

The score may differ slightly depending on the hardware and software environment.

Reproducibility

The data generator, training and evaluation scripts, synthetic data, and recorded results are available in the source repository:

https://github.com/atsushi729/ai-agent-sme

Consumers should pin a Hugging Face commit revision rather than loading an unversioned main branch.

Licence and attribution

The base BERT model and the fine-tuned distribution are provided under the Apache License 2.0. See LICENSE and NOTICE. This is a modified model: multilingual BERT was fine-tuned for three-class hospitality-review sentiment on the synthetic dataset described above. The release is not an official Google product.

Configuration

Architecture
BertForSequenceClassification
Context length (tokens)
512
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
119,547
Stored precision
float32
Model type
bert

Identity and Version

Repository
AtsushiHatake/ai-agent-sme-sentiment-mbert
Publisher
Atsushi Hatakeyama
Task
Text classification
Modality
Text
Library
transformers
Parameters
178M parameters
Languages
ja, en
Revision
5c7888401e6bd8b1b616817f211de6348e70165d
First published
2026-09-20
Last updated
2026-09-20

Files and Weights

10 files, 715.4 MB in total. The weights are 1 file totalling 711.4 MB in safetensors.

Weights1 file · 711.4 MB
Configuration2 files · 1.2 KB
Tokenizer3 files · 3.9 MB
Documentation3 files · 16.0 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights711.4 MB 4f731bdebcd0
config.jsonConfiguration1.1 KB —
special_tokens_map.jsonConfiguration125 B —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation450 B —
README.mdDocumentation4.2 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer2.9 MB —
tokenizer_config.jsonTokenizer1.2 KB —
vocab.txtTokenizer995.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
711.4 MB
Download from Atsushi Hatakeyama

Released by Atsushi Hatakeyama through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published711.4 MB
16-bit0.4 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ai-agent-sme-sentiment-mbert

How much GPU memory does ai-agent-sme-sentiment-mbert need?

About 0.4 GB at 16-bit and 0.1 GB at 4-bit: the weights (178M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run ai-agent-sme-sentiment-mbert on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use ai-agent-sme-sentiment-mbert commercially?

Yes. ai-agent-sme-sentiment-mbert is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is ai-agent-sme-sentiment-mbert's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

deberta-v3-base-prompt-injection-v2

Protect AI

This model is a fine-tuned version of microsoft/deberta-v3-base specifically developed to detect and classify prompt injection attacks which can manipulate language models into producing unintended outputs. Prompt injection attacks manipulate language models by inserting or altering prompts to trigger harmful or unintended responses. The deberta-v3-base-prompt-injection-v2 model is designed to enhance security in language model applications by detecting these malicious interventions. This model classifies inputs into benign (0) and injection-detected (1). deberta-v3-base-prompt-injection-v2 is highly accurate in identifying prompt injections in English. It does not detect jailbreak attacks…

Open weights apache-2.0 184M parameters 512 tokens transformers

Model · Text classification

DEBATE-kor-base

Jong Rock Jeong

DEBATE-kor-base is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTabasev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-base → PolNLI-kor → DEBATE-kor-base Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…

Open weights 184M parameters 512 tokens transformers

This model was trained on 1.279.665 hypothesis-premise pairs from 8 NLI datasets: MultiNLI, Fever-NLI, LingNLI and DocNLI (which includes ANLI, QNLI, DUC, CNN/DailyMail, Curation). It is the only model in the model hub trained on 8 NLI datasets, including DocNLI with very long texts to learn long range reasoning. Note that the model was trained on binary NLI to predict either "entailment" or "not-entailment". The DocNLI merges the classes "neural" and "contradiction" into "not-entailment" to enable the inclusion of the DocNLI dataset. The base model is DeBERTa-v3-base from Microsoft. The v3 variant of DeBERTa substantially outperforms previous versions of the model by including a different…

Open weights mit 184M parameters 512 tokens transformers

Model · Text classification

bert-base-multilingual-uncased-sentiment

NLP Town

Visit the NLP Town website for an updated version of this model, with a 40% error reduction on product reviews. This is a bert-base-multilingual-uncased model finetuned for sentiment analysis on product reviews in six languages: English, Dutch, German, French, Spanish, and Italian. It predicts the sentiment of the review as a number of stars (between 1 and 5). This model is intended for direct use as a sentiment analysis model for product reviews in any of the six languages above or for further finetuning on related sentiment analysis tasks. Here is the number of product reviews we used for finetuning the model: The fine-tuned model obtained the following accuracy on 5,000 held-out product…

Open weights mit 167M parameters 512 tokens transformers

Model · Text classification

pyrrho-v2-nano-g1

Yan Fitzner

Pyrrho is a CPU-runnable co-processor for retrieval-augmented generation. It reads a question before retrieval to suggest what evidence to seek, then reads the question with retrieved passages to assess whether those passages support an answer. A surrounding RAG runtime decides whether to answer, retrieve again, or surface a conflict. The broader project is described in the The evidence verdict is SUFFICIENT, DISPUTED, or INSUFFICIENT. These are predictions about the supplied passages. The model does not retrieve sources, generate answers or citations, check external facts, or prove that a corpus has been searched completely. The two passes use the same encoder with different input…

Open weights cc-by-nc-4.0 150M parameters 8,192 tokens transformers

Model · Text classification

Firebird-ModernBERT-512-RW

Noumenon, Inc.

Firebird-ModernBERT-512-RW is an experimental post-trained variant of It is a ~149M parameter ModernBERT binary classifier for distinguishing: - 0 — HUMAN The maximum sequence length is 512 tokens. This checkpoint was produced through reward-weighted classifier post-training. The original Firebird checkpoint was kept frozen as a reference model. Training examples were scored by the original classifier, difficult examples received larger loss weights, and the post-trained model was constrained against the frozen reference using a KL penalty. L = weightedcrossentropy + beta KL(reference || policy) Hard human examples received greater weighting than ordinary examples because one goal of the…

Open weights 150M parameters 8,192 tokens transformers