SAVRN
Search Contact SAVRN

Open-weight model · Fill mask

LAMB

by Aidan aimgo/LAMB

LAMB is a model for fill mask from Aidan, released under Creative Commons Attribution-NonCommercial 4.0 (access requested at publisher). It has 164M parameters. At 16-bit it needs about 0.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 22 downloads a month.

LAMB (LAtin ModernBERT) is a Latin encoder-only model based on the ModernBERT architecture, pre-trained on nearly 24B Latin tokens, and ready for use with any Latin orthography. If you use this in your work, please cite: Paper pending shortly.

Parameters164M
Context—
Weights1.3 GB
Licensecc-by-nc-4.0
AccessAccess requested at publisher
Monthly Downloads22

Runs On

What it takes to serve LAMB (164M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

LAMB on every accelerator the SAVRN Index prices, at every precision

Model Card

LAMB (LAtin ModernBERT) is a Latin encoder-only model based on the ModernBERT architecture, pre-trained on nearly 24B Latin tokens, and ready for use with any Latin orthography. If you use this in your work, please cite: Paper pending shortly.

Excerpt from the card by Aidan, licensed cc-by-nc-4.0.

Identity and Version

Repository
aimgo/LAMB
Publisher
Aidan
Task
Fill mask
Modality
Text
Library
Not stated by the source
Parameters
164M parameters
Languages
la
Revision
759d4d214e8034473167a2032673357cfed403f0
First published
2025-12-16
Last updated
2026-09-27

Files and Weights

8 files, 1.3 GB in total. The weights are 2 files totalling 1.3 GB in bin, safetensors.

Weights2 files · 1.3 GB
Configuration2 files · 1.8 KB
Tokenizer2 files · 669.9 KB
Documentation1 file · 1.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights656.4 MB —
pytorch_model.binWeights656.4 MB —
config.jsonConfiguration1.1 KB —
special_tokens_map.jsonConfiguration693 B —
README.mdDocumentation1.5 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer668.7 KB —
tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
cc-by-nc-4.0
Access
Access requested at publisher
Download size
1.3 GB
Request access from Aidan

Aidan grants access through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published1.3 GB
16-bit0.3 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About LAMB

How much GPU memory does LAMB need?

About 0.4 GB at 16-bit and 0.1 GB at 4-bit: the weights (164M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run LAMB on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use LAMB commercially?

Not without separate permission. LAMB is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

Similar Models

Pretrained model on the top 102 languages with the largest Wikipedia using a masked language modeling (MLM) objective. It was introduced in this paper and first released in this repository. This model is uncased: it does not make a difference between english and English. Disclaimer: The team releasing BERT did not write a model card for this model so this model card has been written by the Hugging Face team. BERT is a transformers model pretrained on a large corpus of multilingual data in a self-supervised fashion. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an automatic process…

Open weights apache-2.0 168M parameters 512 tokens transformers

Model · Fill mask

ModernBERT-base

Answer.AI

ModernBERT is a modernized bidirectional encoder-only Transformer model (BERT-style) pre-trained on 2 trillion tokens of English and code data with a native context length of up to 8,192 tokens. ModernBERT leverages recent architectural improvements such as: - Rotary Positional Embeddings (RoPE) for long-context support. - Local-Global Alternating Attention for efficiency on long inputs. - Unpadding and Flash Attention for efficient inference. ModernBERT’s native long context length makes it ideal for tasks that require processing long documents, such as retrieval, classification, and semantic search within large corpora. The model was trained on a large corpus of text and code, making it…

Open weights apache-2.0 150M parameters 8,192 tokens transformers

Model · Fill mask

logun-base

Heitorrosa

This model is a fine-tuned version of Itau-Unibanco/NorBERTo-base on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 2 - evalbatchsize: 4 - gradientaccumulationsteps: 16 - totaltrainbatchsize: 32 - lrschedulertype: linear - lrschedulerwarmupsteps: 500 - numepochs: 1.5 - mixedprecisiontraining: Native AMP - PEFT 0.20.0 - Transformers 5.16.1 - Pytorch 2.11.0+cu128 - Datasets 4.0.0 - Tokenizers 0.23.1

Open weights cc-by-nc-sa-4.0 150M parameters 8,192 tokens peft

Pretrained model on the top 104 languages with the largest Wikipedia using a masked language modeling (MLM) objective. It was introduced in this paper and first released in this repository. This model is case sensitive: it makes a difference between english and English. Disclaimer: The team releasing BERT did not write a model card for this model so this model card has been written by the Hugging Face team. BERT is a transformers model pretrained on a large corpus of multilingual data in a self-supervised fashion. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an automatic process to…

Open weights apache-2.0 179M parameters 512 tokens transformers

Model · Fill mask

bert-base-arabertv02

AUB MIND LAB

AraBERT is an Arabic pretrained language model based on Google's BERT architechture. AraBERT uses the same BERT-Base config. More details are available in the AraBERT Paper and in the AraBERT Meetup There are two versions of the model, AraBERTv0.1 and AraBERTv1, with the difference being that AraBERTv1 uses pre-segmented text where prefixes and suffixes were split using the Farasa Segmenter. We evaluate AraBERT models on different downstream tasks and compare them to mBERT), and other state of the art models (To the extent of our knowledge). The Tasks were Sentiment Analysis on 6 different datasets (HARD, ASTD-Balanced, ArsenTD-Lev, LABR), Named Entity Recognition with the ANERcorp, and…

Open weights 136M parameters 512 tokens transformers

This model is a distilled version of the BERT base multilingual model. The code for the distillation process can be found here. This model is cased: it does make a difference between english and English. The model is trained on the concatenation of Wikipedia in 104 different languages listed here. The model has 6 layers, 768 dimension and 12 heads, totalizing 134M parameters (compared to 177M parameters for mBERT-base). On average, this model, referred to as DistilmBERT, is twice as fast as mBERT-base. We encourage potential users of this model to check out the BERT base multilingual model card to learn more about usage, limitations and potential biases. You can use the raw model for either…

Open weights apache-2.0 135M parameters 512 tokens transformers