# my_awesome_eli5_clm-model by Bornil Phukon: Open Model
Source: https://savrn.com/models/my-awesome-eli5-clm-model
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve my_awesome_eli5_clm-model (82M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 0.2 GB | 0.2 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.0 GB | 0.0 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[my_awesome_eli5_clm-model on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/my-awesome-eli5-clm-model/gpus)

## Model Card

By Bornil Phukon, published under apache-2.0, revision 3ef84cbc7418.

This model is a fine-tuned version of distilbert/distilgpt2 on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 2e-05 - trainbatchsize: 8 - evalbatchsize: 8 - lrschedulertype: linear - numepochs: 3.0 - Transformers 5.18.0 - Pytorch 2.11.0+cu130 - Datasets 4.8.5 - Tokenizers 0.23.2

Read Bornil Phukon's full model card

This model is a fine-tuned version of [distilbert/distilgpt2](https://savrn.com/models/distilgpt2) on an unknown dataset. It achieves the following results on the evaluation set: - Loss: 3.7819

### Model description

More information needed

### Intended uses & limitations

More information needed

### Training and evaluation data

More information needed

### Training procedure

#### Training hyperparameters

The following hyperparameters were used during training: - learning_rate: 2e-05 - train_batch_size: 8 - eval_batch_size: 8 - seed: 42 - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments - lr_scheduler_type: linear - num_epochs: 3.0

#### Training results

| Training Loss | Epoch | Step | Validation Loss |
| --- | --- | --- | --- |
| 3.9166 | 1.0 | 1300 | 3.7924 |
| 3.8251 | 2.0 | 2600 | 3.7837 |
| 3.7804 | 3.0 | 3900 | 3.7819 |

#### Framework versions

- Transformers 5.18.0
- Pytorch 2.11.0+cu130
- Datasets 4.8.5
- Tokenizers 0.23.2

## Configuration

Architecture

GPT2LMHeadModel

Vocabulary size

50,257

Model type

gpt2

## Identity and Version

Repository

bornil20005/my_awesome_eli5_clm-model

Publisher

Bornil Phukon

Task

Text generation

Modality

Text

Library

transformers

Parameters

82M parameters

Languages

Not stated by the source

Revision

3ef84cbc741887d37bd98982416a6733f6d5ffea

First published

2026-09-22

Last updated

2026-10-06

## Files and Weights

8 files, 331.2 MB in total. The weights are 2 files totalling 327.7 MB in bin, safetensors.

Weights2 files · 327.7 MB

Configuration2 files · 1.2 KB

Tokenizer2 files · 3.6 MB

Documentation1 file · 1.5 KB

Repository1 file · 1.5 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 327.7 MB | f843e1771c1a |
| training_args.bin | Weights | 5.2 KB | 99a47435f33c |
| config.json | Configuration | 1.1 KB | — |
| generation_config.json | Configuration | 144 B | — |
| README.md | Documentation | 1.5 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 3.6 MB | — |
| tokenizer_config.json | Tokenizer | 326 B | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

327.7 MB

[Download from Bornil Phukon](https://huggingface.co/bornil20005/my_awesome_eli5_clm-model)

Released by Bornil Phukon through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from [distilbert/distilgpt2](https://savrn.com/models/distilgpt2)

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 327.7 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.0 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About my_awesome_eli5_clm-model

### How much GPU memory does my_awesome_eli5_clm-model need?

About 0.2 GB at 16-bit and 0 GB at 4-bit: the weights (82M parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run my_awesome_eli5_clm-model on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use my_awesome_eli5_clm-model commercially?

Yes. my_awesome_eli5_clm-model is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

## Similar Models

Model · Text generation

### [rick-morty-distilgpt2](https://savrn.com/models/rick-morty-distilgpt2)

[Zune toka](https://savrn.com/model-publishers/hfzune)

This model is a fine-tuned version of distilbert/distilgpt2 on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 4 - evalbatchsize: 4 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 16 - lrschedulertype: linear - numepochs: 1 - mixedprecisiontraining: Native AMP - Transformers 5.17.0 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.23.1

Open weights apache-2.0 82M parameters transformers

[View model](https://savrn.com/models/rick-morty-distilgpt2)

Model · Text generation

### [ppt-pythia-160m-uniform250-previous_mse-seed208-stage1](https://savrn.com/models/ppt-pythia-160m-uniform250-previous-mse-seed208-stage1)

[Qing Yao](https://savrn.com/model-publishers/qing-yao)

This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset. The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 13 - trainingsteps: 250 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Open weights apache-2.0 85M parameters 2,048 tokens transformers

[View model](https://savrn.com/models/ppt-pythia-160m-uniform250-previous-mse-seed208-stage1)

Model · Text generation

### [ppt-pythia-160m-uniform250-previous_mse-seed1024-stage1](https://savrn.com/models/ppt-pythia-160m-uniform250-previous-mse-seed1024-stage1)

[Qing Yao](https://savrn.com/model-publishers/qing-yao)

This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset. The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 13 - trainingsteps: 250 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Open weights apache-2.0 85M parameters 2,048 tokens transformers

[View model](https://savrn.com/models/ppt-pythia-160m-uniform250-previous-mse-seed1024-stage1)

Model · Text generation

### [ppt-pythia-160m-uniform250-previous_mse-seed324-stage1](https://savrn.com/models/ppt-pythia-160m-uniform250-previous-mse-seed324-stage1)

[Qing Yao](https://savrn.com/model-publishers/qing-yao)

This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset. The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 13 - trainingsteps: 250 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Open weights apache-2.0 85M parameters 2,048 tokens transformers

[View model](https://savrn.com/models/ppt-pythia-160m-uniform250-previous-mse-seed324-stage1)

Model · Text generation

### [stark-mcp-scratch](https://savrn.com/models/stark-mcp-scratch)

[Victor cherif](https://savrn.com/model-publishers/vcheriffit)

Modelo de lenguaje entrenado desde cero (pesos aleatorios) por Victor Cherif. Arquitectura tipo Llama de ~110M de parametros, pensado como asistente en espanol orientado a ciencia, programacion, robotica y fisica, con soporte de tool-calling / MCP entrenado contra servidores MCP reales. RMSNorm + SwiGLU - Al ser un modelo pequeno (~110M), el lenguaje puede ser incoherente en respuestas largas, sobre todo pasados los primeros parrafos (tiende a repetir bloques o rellenar con generalidades). - La aritmetica mental sigue sin ser confiable: depende de que el modelo dispare la tool calculate en vez de calcular "de memoria" -- verificado que esto mejoro pero no es perfecto. - El razonamiento ( )…

Open weights apache-2.0 88M parameters 16,384 tokens transformers

[View model](https://savrn.com/models/stark-mcp-scratch)

Model · Text generation

### [cRia-LM-75M](https://savrn.com/models/cria-lm-75m)

[Shreyan Mohanty](https://savrn.com/model-publishers/sz14)

cRia-LM-75M is a 75.7M-parameter base language model built as a Relaxed Recursive Transformer (RRT). It uses a shared 11-layer recurrent block evaluated twice, with pass-specific LoRA parameters on the second traversal. Training was carried out in three stages. Stage 1 established the 2K base model over 10B tokens. Stage 2 continued training with a 2B-token budget and a capability-focused data curriculum; the released Stage 2 checkpoint is step 10,000, corresponding to about 1.31B continuation tokens. Stage 3 extended the context window from 2,048 to 4,096 tokens with a 50M-token run on codelion/sutra-1B. The released checkpoint continues from that long-context stage with 35 learned…

Open weights apache-2.0 76M parameters 4,096 tokens transformers

[View model](https://savrn.com/models/cria-lm-75m)

## Bornil Phukon

[All models and datasets](https://savrn.com/model-publishers/bornil20005)

## Versions

- [3ef84cbc7418](https://savrn.com/models/my-awesome-eli5-clm-model/versions/3ef84cbc7418) · current 2026-10-06

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-06.
- [Hugging Face record](https://huggingface.co/bornil20005/my_awesome_eli5_clm-model)
- [How the hub is built](https://savrn.com/model-hub/methodology)
