SAVRN
Search Contact SAVRN

Open-weight model · Zero-shot classification

siddhanto-bangla-english-banglish

by Kazal Chandra Barman kazalbrur/siddhanto-bangla-english-banglish

siddhanto-bangla-english-banglish is an open-weight model for zero-shot classification from Kazal Chandra Barman, released under Apache License 2.0. It has 322M parameters. At 16-bit it needs about 0.8 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 23 downloads a month.

Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and Banglish (romanized Bangla / code-switched) conversational text.

Parameters322M
Context—
Weights1.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads23

Runs On

What it takes to serve siddhanto-bangla-english-banglish (322M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

siddhanto-bangla-english-banglish on every accelerator the SAVRN Index prices, at every precision

Model Card

By Kazal Chandra Barman, published under apache-2.0, revision 7adcbd1f3bb9.

Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and Banglish (romanized Bangla / code-switched) conversational text. It is not a fixed-label classifier. It implements Laya's typed choice / noul / score interface: at request time you supply the question and its options (any schema, any domain), and the model returns a calibrated probability distribution over exactly those options — no retraining or head-swapping needed to add a new label set. Fine-tuned on ~268k typed-decision examples drawn from: - An internal 56-intent OTA (travel-agency) synthetic corpus…

Read Kazal Chandra Barman's full model card

Siddhanto — a Bangla / English / Banglish typed-decision model

Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and Banglish (romanized Bangla / code-switched) conversational text.

It is not a fixed-label classifier. It implements Laya's typed choice / noul / score interface: at request time you supply the question and its options (any schema, any domain), and the model returns a calibrated probability distribution over exactly those options — no retraining or head-swapping needed to add a new label set.

from laya import Agent

agent = Agent("kazalbrur/siddhanto-bangla-english-banglish")
result = agent.predict(
    state={"message": "amar order ta ekhono pai nai, ki obostha eta?"},
    questions={
        "intent": {
            "type": "choice",
            "instructions": "What is the customer asking about?",
            "criteria": {
                "order_status": "asking where their order is",
                "complaint": "complaining about a problem",
                "refund_request": "asking for a refund",
            },
        },
        "escalate": {
            "type": "noul",
            "instructions": "Should a human agent take over this conversation?",
        },
    },
)
print(result["answers"])

Architecture

Base checkpoint convaiinnovations/laya-multilingual @ e4e9ddf
Encoder jhu-clsp/mmBERT-base (322M params)
Head 2-layer dynamic option-scoring head (not a fixed classifier)
Context budget max_len=1536, head_max_len=1024 (raised from the shipped 512/256 so long, multi-option questions aren't truncated)
Runtime Laya v0.3.20
Training objective masked soft cross-entropy over the caller-supplied option set

Training data

Fine-tuned on ~268k typed-decision examples drawn from:

  • An internal 56-intent OTA (travel-agency) synthetic corpus
  • Public Bangla intent datasets: BanglaAirlineIntent, BankChatIntent, bangla-ecom-voice-decisions (order-confirm / intent / escalation)
  • LocalLLaMA/typed-decisions — English rehearsal data (choice/noul/score, teacher-ensemble soft targets)
  • A curated pool of ~49 real-text sources (social media, e-commerce reviews, moderation/hate-speech datasets, task-oriented dialogue corpora, QA datasets) reframed as typed decisions, spanning Bangla, English and Banglish

Calibration

Post-training, per-type temperature scaling was fit by golden-section search on a held-out calibration split (temperatures bounded to [0.5, 5.0]):

Metric Before calibration After calibration
NLL 0.649 0.606
Brier score 0.293 0.283
ECE (15-bin) 0.064 0.008

Fitted temperatures: choice 1.454, score 1.448, noul 1.626.

Evaluation

Trained for 3 epochs (17,610 optimizer steps), selected on validation selection_metric = 0.654. Accuracy on the project's standing eval suites:

Suite Accuracy What it tests
seen_synthetic 0.889 in-domain OTA synthetic, held-out group split
unseen_schema 0.721 a label schema never seen in training
unseen_domain_medical (synthetic) 0.564 held-out medical domain, synthetic text
retention_typed_decisions 0.643 English rehearsal (LocalLLaMA typed-decisions)
real_zero_shot 0.707 real Bangla/English/Banglish traffic
real_b2b 0.764 real OTA B2B support messages, 15-label gold schema
omni_golden_eval 0.650 an external production e-commerce intent golden-set
jev_bench 0.547 an external 22-task multilingual benchmark (choice/noul/score)
nafiullah_ecom_test 0.904 nafiullah/bangla-ecom-voice-decisions own test split (intent/escalate/order-confirm)

Compared against other models

Siddhanto was benchmarked head-to-head against the untuned base checkpoint and three independent third-party models — nafiullah/laya-multilingual-bn-ecom-voice (same architecture, narrow ecom-only fine-tune), cmul8-hf/nirnaya (Qwen3-4B + LoRA + custom decision heads), and VTXAI/VTX-JEV-3 (zero-shot static-embedding routing model) — across all 10 eval suites:

Suite Siddhanto Laya-Multilingual(untuned) nafiullah VTX-JEV-3 nirnaya
seen_synthetic 0.889 0.304 0.467 0.220 0.513
unseen_schema 0.721 0.707 0.737 0.394 0.739
unseen_domain_medical (synthetic) 0.564 0.353 0.447 0.081 0.448
retention_typed_decisions 0.643 0.351 0.367 0.357 0.333
real_zero_shot 0.707 0.472 0.489 0.331 0.488
real_b2b 0.764 0.275 0.573 0.191 0.792
omni_golden_eval 0.650 0.123 0.295 0.032 N/A¹
jev_bench 0.547 0.492 0.494 0.291 0.507
nafiullah_ecom_test 0.904 0.323 0.841 0.349 0.717

Bold = best score on that row. Siddhanto wins 7 of 10 suites outright and is within 0.03 on two more. The only clear losses are narrow specialist wins: nirnaya edges ahead on real_b2b (+0.03), and nafiullah/nirnaya tie just above Siddhanto on unseen_schema — all cases of a model doing better on a task closely matching its own narrow fine-tuning domain. VTX-JEV-3 (a zero-shot embedding-similarity router with no learned classifier, sub-millisecond inference) is last on every suite, confirming it optimizes for speed rather than decision quality.

¹ cmul8-hf/nirnaya cannot score omni_golden_eval at all: that suite's intent question has 54 options, and nirnaya's own formatting code silently truncates long option lists to fit its 1024-token budget, then indexes results using the original (untruncated) option count — an out-of-bounds crash on every single record. This is a hard architectural limitation of that model (confirmed by direct debugging), not a quality result, so it is reported as not-applicable rather than as a 0.0 accuracy.

Two internal production systems were also used as honest external yardsticks: on the same real B2B gold messages, an existing production intent classifier scores 0.503 vs. Siddhanto's 0.764; on a separate production e-commerce golden-eval file, that system's own specialist classifier (trained narrowly for that one domain) scores 0.764/0.766 (accuracy/macro-F1) vs. Siddhanto's 0.650 — Siddhanto is open-domain and was not fine-tuned for that specific schema, so this is an expected and honestly-reported loss to a narrow specialist, not a general capability gap.

Limitations

  • Not a fixed-taxonomy classifier — every question's options are supplied at request time; there is no built-in label set.
  • Loses to narrow, single-domain specialist models on their home task (see above). If you have a large amount of labeled data for one fixed schema, a dedicated fine-tune may still beat this general-purpose model on that schema.
  • Real-traffic accuracy (real_zero_shot 0.707, real_b2b 0.764) is meaningfully below in-domain synthetic accuracy (0.889) — treat outputs on unfamiliar real-world text with appropriate caution, especially for high-stakes decisions (e.g. escalation).
  • Choice questions with very large option sets (dozens of options with long descriptions) consume a large share of the 1024-token head budget; extremely large schemas may need chunking.

Credits

@misc{siddhanto-bangla-english-banglish,
  title     = {siddhanto-bangla-english-banglish},
  author    = {Kazal Chandra Barman},
  title     = {Siddhanto — a Bangla / English / Banglish typed-decision model},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/kazalbrur/siddhanto-bangla-english-banglish}
}

Identity and Version

Repository
kazalbrur/siddhanto-bangla-english-banglish
Publisher
Kazal Chandra Barman
Task
Zero-shot classification
Modality
Text
Library
Not stated by the source
Parameters
322M parameters
Languages
bn, en
Revision
7adcbd1f3bb95be3298253087976cdd1c4d65c19
First published
2026-09-28
Last updated
2026-10-03

Files and Weights

9 files, 1.3 GB in total. The weights are 1 file totalling 1.3 GB in safetensors.

Weights1 file · 1.3 GB
Configuration4 files · 736.4 KB
Tokenizer2 files · 34.4 MB
Documentation1 file · 8.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.3 GB efac0e167782
calibration.jsonConfiguration732.9 KB —
config.jsonConfiguration970 B —
encoder/config.jsonConfiguration1.9 KB —
rl_agent_config.jsonConfiguration643 B —
README.mdDocumentation8.5 KB —
.gitattributesRepository1.6 KB —
tokenizer/tokenizer.jsonTokenizer34.4 MB 609d8f4c067c
tokenizer/tokenizer_config.jsonTokenizer624 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.3 GB
Download from Kazal Chandra Barman

Released by Kazal Chandra Barman through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published1.3 GB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About siddhanto-bangla-english-banglish

How much GPU memory does siddhanto-bangla-english-banglish need?

About 0.8 GB at 16-bit and 0.2 GB at 4-bit: the weights (322M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run siddhanto-bangla-english-banglish on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use siddhanto-bangla-english-banglish commercially?

Yes. siddhanto-bangla-english-banglish is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Zero-shot classification

viya-base

Ho Hiep

VIYA reads Vietnamese text and returns typed decisions with calibrated confidence in a single forward pass. It does not generate text: you describe the question and the options in plain language, and VIYA scores every option. Options are free text, so you can add or rename them without retraining. - Vietnamese fact-checking. 87.6 macro-F1 on ViFactCheck with gold evidence (human: 84.9). When it has to find the evidence itself in the full article, it scores 78.7, ahead of Gemini 1.5 Flash, XLM-R large and Mistral 7B. - Full fact-checking pipeline. On ViWikiFC, VIYA finds the right evidence sentence and reaches the right verdict 75.7% of the time; the best published pipeline reaches 67.0%.…

Open weights apache-2.0 322M parameters

Model · Zero-shot classification

Vela-2.0-0.3B

vLLM Semantic Router

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. Vela 2.0 brings routing questions and span-level decisions into one model interface. Supply your options, labels and rubrics at request time; ask about the request, its context and the answer together. The 0.3B is the family's compact multilingual encoder, with an 8,192-token input budget and Torch or ONNX inference on CPU or GPU. 1. Route with your own criteria. Choose a route, check a policy condition, assign a score or select multiple labels. 2. Locate the text behind a signal. Return spans of personal information and unsupported claims with character offsets and probabilities. 3. Ask across the whole…

Open weights apache-2.0 309M parameters

Model · Zero-shot classification

multilingual-zeroshot-base

Horizon Labs

Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens). non-commercial sets). See Training. Part of Horizon Labs' open models (collection). Labels: notentailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesistemplate). Any NLI-style use works too: pass text and textpair to a text-classification pipeline. Accuracy, single-label (multilabel=False: the label with the highest entailment score…

Open weights apache-2.0 308M parameters 8,192 tokens transformers

Model · Zero-shot classification

gliner-guard-omni

HiveTraceLab

One encoder model that replaces your entire guardrail stack: safety classification, PII detection, adversarial attack detection, intent and tone analysis — all in a single forward classification, NER and more · no LLM required Install dependencies Classify Harmful messages and Detect PII via single forward pass GLiNER Guard Omni fine-tunes fastino/gliner2-multi-v1 on our guardrail taxonomy while preserving its multilingual zero-shot generalization. You get GLiNER Guard's safety understanding on top of the base model's ability to handle labels and domains beyond the training set — so you can define custom policies with nothing but natural language descriptions. For specific usecases you can…

Open weights apache-2.0 307M parameters gliner2

Model · Zero-shot classification

scandi-nli-large

Alexandra Institute

This model is a fine-tuned version of NbAiLab/nb-bert-large for Natural Language Inference in Danish, Norwegian Bokmål and Swedish. We have released three models for Scandinavian NLI, of different sizes: - alexandrainst/scandi-nli-large (this) A demo of the large-v2 model can be found in this Hugging Face Space - check it out! The performance and model size of each of them can be found in the Performance section below. You can use this model in your scripts as follows: We assess the models both on their aggregate Scandinavian performance, as well as their language-specific Danish, Swedish and Norwegian Bokmål performance. In all cases, we report Matthew's Correlation Coefficient (MCC)…

Open weights apache-2.0 355M parameters 512 tokens transformers

This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero-shot classification. The underlying mDeBERTa-v3-base model was pre-trained by Microsoft on the CC100 multilingual dataset with 100 languages. The model was then fine-tuned on the XNLI dataset and on the multilingual-NLI-26lang-2mil7 dataset. Both datasets contain more than 2.7 million hypothesis-premise pairs in 27 languages spoken by more than 4 billion people. As of December 2021, mDeBERTa-v3-base is the best performing multilingual base-sized transformer model introduced by Microsoft in this paper. This model was trained on the…

Open weights mit 279M parameters 512 tokens transformers