VIYA reads Vietnamese text and returns typed decisions with calibrated confidence in a single forward pass. It does not generate text: you describe the question and the options in plain language, and VIYA scores every option. Options are free text, so you can add or rename them without retraining. - Vietnamese fact-checking. 87.6 macro-F1 on ViFactCheck with gold evidence (human: 84.9). When it has to find the evidence itself in the full article, it scores 78.7, ahead of Gemini 1.5 Flash, XLM-R large and Mistral 7B. - Full fact-checking pipeline. On ViWikiFC, VIYA finds the right evidence sentence and reaches the right verdict 75.7% of the time; the best published pipeline reaches 67.0%.…
Open-weight model · Zero-shot classification
siddhanto-bangla-english-banglish
by Kazal Chandra Barman kazalbrur/siddhanto-bangla-english-banglish
siddhanto-bangla-english-banglish is an open-weight model for zero-shot classification from Kazal Chandra Barman, released under Apache License 2.0. It has 322M parameters. At 16-bit it needs about 0.8 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 23 downloads a month.
Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and Banglish (romanized Bangla / code-switched) conversational text.
Runs On
What it takes to serve siddhanto-bangla-english-banglish (322M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.6 GB | 0.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.3 GB | 0.4 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
siddhanto-bangla-english-banglish on every accelerator the SAVRN Index prices, at every precision
Model Card
By Kazal Chandra Barman, published under apache-2.0, revision 7adcbd1f3bb9.
Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and Banglish (romanized Bangla / code-switched) conversational text. It is not a fixed-label classifier. It implements Laya's typed choice / noul / score interface: at request time you supply the question and its options (any schema, any domain), and the model returns a calibrated probability distribution over exactly those options — no retraining or head-swapping needed to add a new label set. Fine-tuned on ~268k typed-decision examples drawn from: - An internal 56-intent OTA (travel-agency) synthetic corpus…
Read Kazal Chandra Barman's full model card
Siddhanto — a Bangla / English / Banglish typed-decision model
Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual
(mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and
Banglish (romanized Bangla / code-switched) conversational text.
It is not a fixed-label classifier. It implements Laya's typed choice / noul / score
interface: at request time you supply the question and its options (any schema, any domain),
and the model returns a calibrated probability distribution over exactly those options —
no retraining or head-swapping needed to add a new label set.
from laya import Agent
agent = Agent("kazalbrur/siddhanto-bangla-english-banglish")
result = agent.predict(
state={"message": "amar order ta ekhono pai nai, ki obostha eta?"},
questions={
"intent": {
"type": "choice",
"instructions": "What is the customer asking about?",
"criteria": {
"order_status": "asking where their order is",
"complaint": "complaining about a problem",
"refund_request": "asking for a refund",
},
},
"escalate": {
"type": "noul",
"instructions": "Should a human agent take over this conversation?",
},
},
)
print(result["answers"])
Architecture
| Base checkpoint | convaiinnovations/laya-multilingual @ e4e9ddf |
| Encoder | jhu-clsp/mmBERT-base (322M params) |
| Head | 2-layer dynamic option-scoring head (not a fixed classifier) |
| Context budget | max_len=1536, head_max_len=1024 (raised from the shipped 512/256 so long, multi-option questions aren't truncated) |
| Runtime | Laya v0.3.20 |
| Training objective | masked soft cross-entropy over the caller-supplied option set |
Training data
Fine-tuned on ~268k typed-decision examples drawn from:
- An internal 56-intent OTA (travel-agency) synthetic corpus
- Public Bangla intent datasets:
BanglaAirlineIntent,BankChatIntent,bangla-ecom-voice-decisions(order-confirm / intent / escalation) LocalLLaMA/typed-decisions— English rehearsal data (choice/noul/score, teacher-ensemble soft targets)- A curated pool of ~49 real-text sources (social media, e-commerce reviews, moderation/hate-speech datasets, task-oriented dialogue corpora, QA datasets) reframed as typed decisions, spanning Bangla, English and Banglish
Calibration
Post-training, per-type temperature scaling was fit by golden-section search on a held-out
calibration split (temperatures bounded to [0.5, 5.0]):
| Metric | Before calibration | After calibration |
|---|---|---|
| NLL | 0.649 | 0.606 |
| Brier score | 0.293 | 0.283 |
| ECE (15-bin) | 0.064 | 0.008 |
Fitted temperatures: choice 1.454, score 1.448, noul 1.626.
Evaluation
Trained for 3 epochs (17,610 optimizer steps), selected on validation selection_metric = 0.654.
Accuracy on the project's standing eval suites:
| Suite | Accuracy | What it tests |
|---|---|---|
| seen_synthetic | 0.889 | in-domain OTA synthetic, held-out group split |
| unseen_schema | 0.721 | a label schema never seen in training |
| unseen_domain_medical (synthetic) | 0.564 | held-out medical domain, synthetic text |
| retention_typed_decisions | 0.643 | English rehearsal (LocalLLaMA typed-decisions) |
| real_zero_shot | 0.707 | real Bangla/English/Banglish traffic |
| real_b2b | 0.764 | real OTA B2B support messages, 15-label gold schema |
| omni_golden_eval | 0.650 | an external production e-commerce intent golden-set |
| jev_bench | 0.547 | an external 22-task multilingual benchmark (choice/noul/score) |
| nafiullah_ecom_test | 0.904 | nafiullah/bangla-ecom-voice-decisions own test split (intent/escalate/order-confirm) |
Compared against other models
Siddhanto was benchmarked head-to-head against the untuned base checkpoint and three independent
third-party models — nafiullah/laya-multilingual-bn-ecom-voice (same architecture, narrow
ecom-only fine-tune), cmul8-hf/nirnaya (Qwen3-4B + LoRA + custom decision heads), and
VTXAI/VTX-JEV-3 (zero-shot static-embedding routing model) — across all 10 eval suites:
| Suite | Siddhanto | Laya-Multilingual(untuned) | nafiullah | VTX-JEV-3 | nirnaya |
|---|---|---|---|---|---|
| seen_synthetic | 0.889 | 0.304 | 0.467 | 0.220 | 0.513 |
| unseen_schema | 0.721 | 0.707 | 0.737 | 0.394 | 0.739 |
| unseen_domain_medical (synthetic) | 0.564 | 0.353 | 0.447 | 0.081 | 0.448 |
| retention_typed_decisions | 0.643 | 0.351 | 0.367 | 0.357 | 0.333 |
| real_zero_shot | 0.707 | 0.472 | 0.489 | 0.331 | 0.488 |
| real_b2b | 0.764 | 0.275 | 0.573 | 0.191 | 0.792 |
| omni_golden_eval | 0.650 | 0.123 | 0.295 | 0.032 | N/A¹ |
| jev_bench | 0.547 | 0.492 | 0.494 | 0.291 | 0.507 |
| nafiullah_ecom_test | 0.904 | 0.323 | 0.841 | 0.349 | 0.717 |
Bold = best score on that row. Siddhanto wins 7 of 10 suites outright and is within 0.03 on two
more. The only clear losses are narrow specialist wins: nirnaya edges ahead on
real_b2b (+0.03), and nafiullah/nirnaya tie just above
Siddhanto on unseen_schema — all cases of a model doing better on a task closely matching its
own narrow fine-tuning domain. VTX-JEV-3 (a zero-shot embedding-similarity router with no
learned classifier, sub-millisecond inference) is last on every suite, confirming it optimizes for
speed rather than decision quality.
¹ cmul8-hf/nirnaya cannot score omni_golden_eval at all: that suite's intent question has 54
options, and nirnaya's own formatting code silently truncates long option lists to fit its
1024-token budget, then indexes results using the original (untruncated) option count — an
out-of-bounds crash on every single record. This is a hard architectural limitation of that
model (confirmed by direct debugging), not a quality result, so it is reported as not-applicable
rather than as a 0.0 accuracy.
Two internal production systems were also used as honest external yardsticks: on the same real B2B gold messages, an existing production intent classifier scores 0.503 vs. Siddhanto's 0.764; on a separate production e-commerce golden-eval file, that system's own specialist classifier (trained narrowly for that one domain) scores 0.764/0.766 (accuracy/macro-F1) vs. Siddhanto's 0.650 — Siddhanto is open-domain and was not fine-tuned for that specific schema, so this is an expected and honestly-reported loss to a narrow specialist, not a general capability gap.
Limitations
- Not a fixed-taxonomy classifier — every question's options are supplied at request time; there is no built-in label set.
- Loses to narrow, single-domain specialist models on their home task (see above). If you have a large amount of labeled data for one fixed schema, a dedicated fine-tune may still beat this general-purpose model on that schema.
- Real-traffic accuracy (
real_zero_shot0.707,real_b2b0.764) is meaningfully below in-domain synthetic accuracy (0.889) — treat outputs on unfamiliar real-world text with appropriate caution, especially for high-stakes decisions (e.g. escalation). - Choice questions with very large option sets (dozens of options with long descriptions) consume a large share of the 1024-token head budget; extremely large schemas may need chunking.
Credits
- Base architecture and runtime: Laya by NandhaKishorM,
and the base checkpoint
convaiinnovations/laya-multilingual. LocalLLaMA/typed-decisions(Apache-2.0) for English rehearsal data.
@misc{siddhanto-bangla-english-banglish,
title = {siddhanto-bangla-english-banglish},
author = {Kazal Chandra Barman},
title = {Siddhanto — a Bangla / English / Banglish typed-decision model},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/kazalbrur/siddhanto-bangla-english-banglish}
}
Identity and Version
- Repository
- kazalbrur/siddhanto-bangla-english-banglish
- Publisher
- Kazal Chandra Barman
- Task
- Zero-shot classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- 322M parameters
- Languages
- bn, en
- Revision
- 7adcbd1f3bb95be3298253087976cdd1c4d65c19
- First published
- 2026-09-28
- Last updated
- 2026-10-03
Files and Weights
9 files, 1.3 GB in total. The weights are 1 file totalling 1.3 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 1.3 GB | efac0e167782 |
| calibration.json | Configuration | 732.9 KB | — |
| config.json | Configuration | 970 B | — |
| encoder/config.json | Configuration | 1.9 KB | — |
| rl_agent_config.json | Configuration | 643 B | — |
| README.md | Documentation | 8.5 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer/tokenizer.json | Tokenizer | 34.4 MB | 609d8f4c067c |
| tokenizer/tokenizer_config.json | Tokenizer | 624 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 1.3 GB
Released by Kazal Chandra Barman through its official repository on Hugging Face. Read the license.
Built From
- Derived from convaiinnovations/laya-multilingual
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 1.3 GB |
| 16-bit | 0.6 GB |
| 8-bit | 0.3 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About siddhanto-bangla-english-banglish
How much GPU memory does siddhanto-bangla-english-banglish need?
About 0.8 GB at 16-bit and 0.2 GB at 4-bit: the weights (322M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run siddhanto-bangla-english-banglish on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use siddhanto-bangla-english-banglish commercially?
Yes. siddhanto-bangla-english-banglish is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. Vela 2.0 brings routing questions and span-level decisions into one model interface. Supply your options, labels and rubrics at request time; ask about the request, its context and the answer together. The 0.3B is the family's compact multilingual encoder, with an 8,192-token input budget and Torch or ONNX inference on CPU or GPU. 1. Route with your own criteria. Choose a route, check a policy condition, assign a score or select multiple labels. 2. Locate the text behind a signal. Return spans of personal information and unsupported claims with character offsets and probabilities. 3. Ask across the whole…
Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens). non-commercial sets). See Training. Part of Horizon Labs' open models (collection). Labels: notentailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesistemplate). Any NLI-style use works too: pass text and textpair to a text-classification pipeline. Accuracy, single-label (multilabel=False: the label with the highest entailment score…
One encoder model that replaces your entire guardrail stack: safety classification, PII detection, adversarial attack detection, intent and tone analysis — all in a single forward classification, NER and more · no LLM required Install dependencies Classify Harmful messages and Detect PII via single forward pass GLiNER Guard Omni fine-tunes fastino/gliner2-multi-v1 on our guardrail taxonomy while preserving its multilingual zero-shot generalization. You get GLiNER Guard's safety understanding on top of the base model's ability to handle labels and domains beyond the training set — so you can define custom policies with nothing but natural language descriptions. For specific usecases you can…
This model is a fine-tuned version of NbAiLab/nb-bert-large for Natural Language Inference in Danish, Norwegian Bokmål and Swedish. We have released three models for Scandinavian NLI, of different sizes: - alexandrainst/scandi-nli-large (this) A demo of the large-v2 model can be found in this Hugging Face Space - check it out! The performance and model size of each of them can be found in the Performance section below. You can use this model in your scripts as follows: We assess the models both on their aggregate Scandinavian performance, as well as their language-specific Danish, Swedish and Norwegian Bokmål performance. In all cases, we report Matthew's Correlation Coefficient (MCC)…
Model · Zero-shot classification
mDeBERTa-v3-base-xnli-multilingual-nli-2mil7
This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero-shot classification. The underlying mDeBERTa-v3-base model was pre-trained by Microsoft on the CC100 multilingual dataset with 100 languages. The model was then fine-tuned on the XNLI dataset and on the multilingual-NLI-26lang-2mil7 dataset. Both datasets contain more than 2.7 million hypothesis-premise pairs in 27 languages spoken by more than 4 billion people. As of December 2021, mDeBERTa-v3-base is the best performing multilingual base-sized transformer model introduced by Microsoft in this paper. This model was trained on the…