Open-weight model
led-large-16384-BASE3ep-ASPRPreACE3ep
by Rosa Rodriguez-Sánchez rosadecsai/led-large-16384-BASE3ep-ASPRPreACE3ep
led-large-16384-BASE3ep-ASPRPreACE3ep is an open-weight model from Rosa Rodriguez-Sánchez, released under Apache License 2.0. It has 460M parameters. At 16-bit it needs about 1.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 53 downloads a month.
This model is a fine-tuned version of rosadecsai/led-large-16384-BASE3ep on the None dataset.
Runs On
What it takes to serve led-large-16384-BASE3ep-ASPRPreACE3ep (460M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.9 GB | 1.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.5 GB | 0.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.
Model Card
By Rosa Rodriguez-Sánchez, published under apache-2.0, revision 90f199862c6c.
This model is a fine-tuned version of rosadecsai/led-large-16384-BASE3ep on the None dataset. It achieves the following results on the evaluation set: - Loss: 2.0896 - Rouge1: 46.0449 - Rouge2: 15.3846 - Rougel: 20.0708 - Rougelsum: 44.1558 - Gen Len: 1.0
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training: - learning_rate: 5e-05 - train_batch_size: 8 - eval_batch_size: 8 - seed: 42 - gradient_accumulation_steps: 2 - total_train_batch_size: 16 - optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments - lr_scheduler_type: linear - num_epochs: 3 - mixed_precision_training: Native AMP
Training results
| Training Loss | Epoch | Step | Validation Loss | Rouge1 | Rouge2 | Rougel | Rougelsum | Gen Len |
|---|---|---|---|---|---|---|---|---|
| 2.305 | 1.0 | 1132 | 2.0878 | 33.1015 | 12.5523 | 16.968 | 31.4325 | 1.0 |
| 2.3229 | 2.0 | 2264 | 2.0782 | 39.2573 | 16.2234 | 18.5676 | 38.1963 | 1.0 |
| 2.2209 | 2.9978 | 3393 | 2.0896 | 46.0449 | 15.3846 | 20.0708 | 44.1558 | 1.0 |
Framework versions
Configuration
- Architecture
- MultiTask_LED
- Layers
- 12
- Vocabulary size
- 50,265
- Stored precision
- float32
- Model type
- led
Identity and Version
- Repository
- rosadecsai/led-large-16384-BASE3ep-ASPRPreACE3ep
- Publisher
- Rosa Rodriguez-Sánchez
- Task
- Not stated by the source
- Modality
- Other
- Library
- transformers
- Parameters
- 460M parameters
- Languages
- led
- Revision
- 90f199862c6cc7f2233456c0b169014e7989f46a
- First published
- 2026-09-23
- Last updated
- 2026-09-29
Files and Weights
14 files, 1.8 GB in total. The weights are 2 files totalling 1.8 GB in bin, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 1.8 GB | 17161ab1c946 |
| training_args.bin | Weights | 8.1 KB | 868cb6cb5d2e |
| config.json | Configuration | 1.4 KB | — |
| generation_config.json | Configuration | 303 B | — |
| special_tokens_map.json | Configuration | 957 B | — |
| README.md | Documentation | 2.1 KB | — |
| runs/Sep23_17-00-06_ea84f59ae7ce/events.out.tfevents.1790182814.ea84f59ae7ce.1911.0 | Other | 11.2 KB | 4d05261122c3 |
| runs/Sep24_08-57-51_d6031ce138c3/events.out.tfevents.1790240288.d6031ce138c3.668.0 | Other | 11.2 KB | e31d7f37bffc |
| runs/Sep28_16-09-40_1e1d46e6fde9/events.out.tfevents.1790611796.1e1d46e6fde9.1591.0 | Other | 10.1 KB | 64cf4a9ff6f4 |
| .gitattributes | Repository | 1.5 KB | — |
| merges.txt | Tokenizer | 456.3 KB | — |
| tokenizer.json | Tokenizer | 3.6 MB | — |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |
| vocab.json | Tokenizer | 798.3 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 1.8 GB
Released by Rosa Rodriguez-Sánchez through its official repository on Hugging Face. Read the license.
Built From
- Derived from rosadecsai/led-large-16384-BASE3ep
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 1.8 GB |
| 16-bit | 0.9 GB |
| 8-bit | 0.5 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About led-large-16384-BASE3ep-ASPRPreACE3ep
How much GPU memory does led-large-16384-BASE3ep-ASPRPreACE3ep need?
About 1.1 GB at 16-bit and 0.3 GB at 4-bit: the weights (460M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run led-large-16384-BASE3ep-ASPRPreACE3ep on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use led-large-16384-BASE3ep-ASPRPreACE3ep commercially?
Yes. led-large-16384-BASE3ep-ASPRPreACE3ep is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.