This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-shuff-dyck-100mbseed10. It has been trained using TRL. This model was trained with SFT.
Open-weight model · Text generation
arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10
by Francesca Padovani fpadovani/arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10
This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-shuff-dyck-10mbseed10. It has been trained using TRL. This model was trained with SFT.
Runs On
What it takes to serve arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10 (125M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-shuff-dyck-10mbseed10. It has been trained using TRL. This model was trained with SFT.
Excerpt from the card by Francesca Padovani.
Configuration
- Architecture
- GPT2LMHeadModel
- Vocabulary size
- 51,200
- Model type
- gpt2
Identity and Version
- Repository
- fpadovani/arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10
- Publisher
- Francesca Padovani
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 125M parameters
- Languages
- sft, trl
- Revision
- 39e253d1c682a95512eb5d0a875375d0bcb7e809
- First published
- 2026-09-18
- Last updated
- 2026-09-18
Files and Weights
120 files, 3.0 GB in total. The weights are 35 files totalling 3.0 GB in bin, pth, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| checkpoint-1000/model.safetensors | Weights | 249.6 MB | 60b0ddcce024 |
| checkpoint-1000/rng_state.pth | Weights | 14.6 KB | 21d9e5c98678 |
| checkpoint-1000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-10000/model.safetensors | Weights | 249.6 MB | 3d4d2a113cc3 |
| checkpoint-10000/rng_state.pth | Weights | 14.6 KB | 2066141fe56d |
| checkpoint-10000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-10840/model.safetensors | Weights | 249.6 MB | 4fa563d72985 |
| checkpoint-10840/rng_state.pth | Weights | 14.6 KB | 950377194415 |
| checkpoint-10840/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-2000/model.safetensors | Weights | 249.6 MB | c9f5faf76c88 |
| checkpoint-2000/rng_state.pth | Weights | 14.6 KB | 23af0cfd8903 |
| checkpoint-2000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-3000/model.safetensors | Weights | 249.6 MB | 3be47cb61985 |
| checkpoint-3000/rng_state.pth | Weights | 14.6 KB | 7d60b7213ad7 |
| checkpoint-3000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-4000/model.safetensors | Weights | 249.6 MB | 0d90db8f5807 |
| checkpoint-4000/rng_state.pth | Weights | 14.6 KB | d31fe4298fa4 |
| checkpoint-4000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-5000/model.safetensors | Weights | 249.6 MB | 34ea79bc723f |
| checkpoint-5000/rng_state.pth | Weights | 14.6 KB | 265a2239fce3 |
| checkpoint-5000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-6000/model.safetensors | Weights | 249.6 MB | 47f8f9f626b2 |
| checkpoint-6000/rng_state.pth | Weights | 14.6 KB | 9083f287e0bb |
| checkpoint-6000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-7000/model.safetensors | Weights | 249.6 MB | 6ebf5b8a93a7 |
| checkpoint-7000/rng_state.pth | Weights | 14.6 KB | 1c040b34d441 |
| checkpoint-7000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-8000/model.safetensors | Weights | 249.6 MB | 98725fc48915 |
| checkpoint-8000/rng_state.pth | Weights | 14.6 KB | 94168933c95c |
| checkpoint-8000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| checkpoint-9000/model.safetensors | Weights | 249.6 MB | 416df7b9a93d |
| checkpoint-9000/rng_state.pth | Weights | 14.6 KB | 158309b5e7e2 |
| checkpoint-9000/training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| model.safetensors | Weights | 249.6 MB | 4fa563d72985 |
| training_args.bin | Weights | 6.4 KB | 19f1238ebf83 |
| added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-1000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-1000/config.json | Configuration | 803 B | — |
| checkpoint-1000/generation_config.json | Configuration | 154 B | — |
| checkpoint-1000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-1000/trainer_state.json | Configuration | 55.3 KB | — |
| checkpoint-10000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-10000/config.json | Configuration | 803 B | — |
| checkpoint-10000/generation_config.json | Configuration | 154 B | — |
| checkpoint-10000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-10000/trainer_state.json | Configuration | 563.7 KB | — |
| checkpoint-10840/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-10840/config.json | Configuration | 803 B | — |
| checkpoint-10840/generation_config.json | Configuration | 154 B | — |
| checkpoint-10840/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-10840/trainer_state.json | Configuration | 611.0 KB | — |
| checkpoint-2000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-2000/config.json | Configuration | 803 B | — |
| checkpoint-2000/generation_config.json | Configuration | 154 B | — |
| checkpoint-2000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-2000/trainer_state.json | Configuration | 111.7 KB | — |
| checkpoint-3000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-3000/config.json | Configuration | 803 B | — |
| checkpoint-3000/generation_config.json | Configuration | 154 B | — |
| checkpoint-3000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-3000/trainer_state.json | Configuration | 168.2 KB | — |
| checkpoint-4000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-4000/config.json | Configuration | 803 B | — |
| checkpoint-4000/generation_config.json | Configuration | 154 B | — |
| checkpoint-4000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-4000/trainer_state.json | Configuration | 224.7 KB | — |
| checkpoint-5000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-5000/config.json | Configuration | 803 B | — |
| checkpoint-5000/generation_config.json | Configuration | 154 B | — |
| checkpoint-5000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-5000/trainer_state.json | Configuration | 281.0 KB | — |
| checkpoint-6000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-6000/config.json | Configuration | 803 B | — |
| checkpoint-6000/generation_config.json | Configuration | 154 B | — |
| checkpoint-6000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-6000/trainer_state.json | Configuration | 337.6 KB | — |
| checkpoint-7000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-7000/config.json | Configuration | 803 B | — |
| checkpoint-7000/generation_config.json | Configuration | 154 B | — |
| checkpoint-7000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-7000/trainer_state.json | Configuration | 394.1 KB | — |
| checkpoint-8000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-8000/config.json | Configuration | 803 B | — |
| checkpoint-8000/generation_config.json | Configuration | 154 B | — |
| checkpoint-8000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-8000/trainer_state.json | Configuration | 450.8 KB | — |
| checkpoint-9000/added_tokens.json | Configuration | 27.7 KB | — |
| checkpoint-9000/config.json | Configuration | 803 B | — |
| checkpoint-9000/generation_config.json | Configuration | 154 B | — |
| checkpoint-9000/special_tokens_map.json | Configuration | 22.6 KB | — |
| checkpoint-9000/trainer_state.json | Configuration | 507.3 KB | — |
| config.json | Configuration | 803 B | — |
| generation_config.json | Configuration | 154 B | — |
| special_tokens_map.json | Configuration | 22.6 KB | — |
| README.md | Documentation | 1.9 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| checkpoint-1000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-1000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-10000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-10000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-10840/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-10840/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-2000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-2000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-3000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-3000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-4000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-4000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-5000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-5000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-6000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-6000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-7000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-7000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-8000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-8000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| checkpoint-9000/spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| checkpoint-9000/tokenizer_config.json | Tokenizer | 233.9 KB | — |
| spiece.model | Tokenizer | 1.3 MB | 40e1337dd115 |
| tokenizer_config.json | Tokenizer | 233.9 KB | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
- Download size
- 3.0 GB
Released by Francesca Padovani through its official repository on Hugging Face.
Built From
- Derived from fpadovani/arb-arab-100mb-ppt-shuff-dyck-10mb_seed10
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 3.0 GB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10
How much GPU memory does arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10 need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (125M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run arb-arab-100mb-after-ppt-shuff-dyck-10mb-ckpt500_seed10 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Similar Models
This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-Dp-100mbseed10. It has been trained using TRL. This model was trained with SFT.
This model is a fine-tuned version of GeorgeUwaifo/iviegpt2new01cresults on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 4 - evalbatchsize: 8 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 16 - lrschedulertype: linear - lrschedulerwarmupsteps: 387 - numepochs: 5 - mixedprecisiontraining: Native AMP - Transformers 5.16.1 - Pytorch 2.11.0+cu128 - Datasets 4.8.5 - Tokenizers 0.23.1
This model is a custom-code derivative of AxiomicLabs/GPT-X2-125M, adapted for experimental long-context causal language modeling and architecture research. The repository includes a Hugging Face Transformers-compatible GPT-X2 implementation with optional Symplectic Metric-RoPE Governor support and training utilities built around CIxOpt, a heterogeneous optimizer developed for efficient parameter routing across large projection matrices, sensitive normalization parameters, and optional governor modules. The model is intended as a research checkpoint for compact long-context generation, positional encoding experiments, optimizer testing, and continued fine-tuning. This implementation uses a…
SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device. More details in our paper: https://arxiv.org/abs/2502.02737 SmolLM2 demonstrates significant advances over its predecessor SmolLM1, particularly in instruction following, knowledge, reasoning. The 135M model was trained on 2 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new filtered datasets we curated and will release soon. We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our…
SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device. More details in our paper https://arxiv.org/abs/2502.02737 SmolLM2 demonstrates significant advances over its predecessor SmolLM1, particularly in instruction following, knowledge, reasoning. The 135M model was trained on 2 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new filtered datasets we curated and will release soon. We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our own…