This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset. The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 13 - trainingsteps: 250 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1
Open-weight model · Text generation
ppt-pythia-160m-uniform250-previous_mse-seed208-stage1
by Qing Yao qing-yao/ppt-pythia-160m-uniform250-previous_mse-seed208-stage1
ppt-pythia-160m-uniform250-previous_mse-seed208-stage1 is an open-weight model for text generation from Qing Yao, released under Apache License 2.0. It has 85M parameters and a 2,048-token context. At 16-bit it needs about 0.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset.
Runs On
What it takes to serve ppt-pythia-160m-uniform250-previous_mse-seed208-stage1 (85M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.0 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Model Card
By Qing Yao, published under apache-2.0, revision c0c8545d5775.
This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset. The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 13 - trainingsteps: 250 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1
Read Qing Yao's full model card
This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The following hyperparameters were used during training: - learning_rate: 0.001 - train_batch_size: 16 - eval_batch_size: 16 - seed: 208 - gradient_accumulation_steps: 2 - total_train_batch_size: 32 - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments - lr_scheduler_type: cosine_with_min_lr - lr_scheduler_warmup_steps: 13 - training_steps: 250
Training results
Framework versions
- Transformers 5.4.0
- Pytorch 2.8.0+cu128
- Datasets 3.2.0
- Tokenizers 0.22.1
Configuration
- Architecture
- GPTNeoXForCausalLM
- Context length (tokens)
- 2,048
- Layers
- 12
- Hidden size
- 768
- Feed-forward size
- 3,072
- Attention heads
- 12
- Vocabulary size
- 10
- Model type
- gpt_neox
Identity and Version
- Repository
- qing-yao/ppt-pythia-160m-uniform250-previous_mse-seed208-stage1
- Publisher
- Qing Yao
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 85M parameters
- Languages
- Not stated by the source
- Revision
- c0c8545d5775ee97dea958ed5d127f94ea27763b
- First published
- 2026-09-25
- Last updated
- 2026-09-25
Files and Weights
66 files, 2.9 GB in total. The weights are 29 files totalling 2.9 GB in bin, pt, pth, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| checkpoint-100/model.safetensors | Weights | 170.2 MB | 8753a568f8fe |
| checkpoint-100/optimizer.pt | Weights | 340.4 MB | a0c67543d3c1 |
| checkpoint-100/rng_state.pth | Weights | 14.7 KB | 2f51e71fc7b1 |
| checkpoint-100/scheduler.pt | Weights | 1.5 KB | 9d7abad781cf |
| checkpoint-100/training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| checkpoint-150/model.safetensors | Weights | 170.2 MB | a21cfe5cf88a |
| checkpoint-150/optimizer.pt | Weights | 340.4 MB | 675670e032ad |
| checkpoint-150/rng_state.pth | Weights | 14.7 KB | 2f51e71fc7b1 |
| checkpoint-150/scheduler.pt | Weights | 1.5 KB | a3dffde530f0 |
| checkpoint-150/training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| checkpoint-200/model.safetensors | Weights | 170.2 MB | ee262c9e660d |
| checkpoint-200/optimizer.pt | Weights | 340.4 MB | 002ab087afc0 |
| checkpoint-200/rng_state.pth | Weights | 14.7 KB | 2f51e71fc7b1 |
| checkpoint-200/scheduler.pt | Weights | 1.5 KB | aeabaaa65d19 |
| checkpoint-200/training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| checkpoint-250/model.safetensors | Weights | 170.2 MB | a00d0ca20cbd |
| checkpoint-250/optimizer.pt | Weights | 340.4 MB | 6c132c45a375 |
| checkpoint-250/rng_state.pth | Weights | 14.7 KB | 2f51e71fc7b1 |
| checkpoint-250/scheduler.pt | Weights | 1.5 KB | 2395199a83e4 |
| checkpoint-250/training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| checkpoint-50/model.safetensors | Weights | 170.2 MB | 15749148b0f6 |
| checkpoint-50/optimizer.pt | Weights | 340.4 MB | d03f33810ce0 |
| checkpoint-50/rng_state.pth | Weights | 14.7 KB | 2f51e71fc7b1 |
| checkpoint-50/scheduler.pt | Weights | 1.5 KB | 84927109043e |
| checkpoint-50/training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| final/model.safetensors | Weights | 170.2 MB | a00d0ca20cbd |
| final/training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| model.safetensors | Weights | 170.2 MB | a00d0ca20cbd |
| training_args.bin | Weights | 5.3 KB | bfd2747a40ba |
| checkpoint-100/config.json | Configuration | 778 B | — |
| checkpoint-100/generation_config.json | Configuration | 225 B | — |
| checkpoint-100/trainer_state.json | Configuration | 1.3 KB | — |
| checkpoint-150/config.json | Configuration | 778 B | — |
| checkpoint-150/generation_config.json | Configuration | 225 B | — |
| checkpoint-150/trainer_state.json | Configuration | 1.4 KB | — |
| checkpoint-200/config.json | Configuration | 778 B | — |
| checkpoint-200/generation_config.json | Configuration | 225 B | — |
| checkpoint-200/trainer_state.json | Configuration | 1.6 KB | — |
| checkpoint-250/config.json | Configuration | 778 B | — |
| checkpoint-250/generation_config.json | Configuration | 225 B | — |
| checkpoint-250/trainer_state.json | Configuration | 1.7 KB | — |
| checkpoint-50/config.json | Configuration | 778 B | — |
| checkpoint-50/generation_config.json | Configuration | 225 B | — |
| checkpoint-50/trainer_state.json | Configuration | 1.1 KB | — |
| config.json | Configuration | 778 B | — |
| experiment.json | Configuration | 5.1 KB | — |
| final/config.json | Configuration | 778 B | — |
| final/generation_config.json | Configuration | 225 B | — |
| generation_config.json | Configuration | 225 B | — |
| trainer_state.json | Configuration | 2.0 KB | — |
| README.md | Documentation | 1.4 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| checkpoint-100/tokenizer.json | Tokenizer | 3.6 MB | — |
| checkpoint-100/tokenizer_config.json | Tokenizer | 349 B | — |
| checkpoint-150/tokenizer.json | Tokenizer | 3.6 MB | — |
| checkpoint-150/tokenizer_config.json | Tokenizer | 349 B | — |
| checkpoint-200/tokenizer.json | Tokenizer | 3.6 MB | — |
| checkpoint-200/tokenizer_config.json | Tokenizer | 349 B | — |
| checkpoint-250/tokenizer.json | Tokenizer | 3.6 MB | — |
| checkpoint-250/tokenizer_config.json | Tokenizer | 349 B | — |
| checkpoint-50/tokenizer.json | Tokenizer | 3.6 MB | — |
| checkpoint-50/tokenizer_config.json | Tokenizer | 349 B | — |
| final/tokenizer.json | Tokenizer | 3.6 MB | — |
| final/tokenizer_config.json | Tokenizer | 349 B | — |
| tokenizer.json | Tokenizer | 3.6 MB | — |
| tokenizer_config.json | Tokenizer | 349 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 2.9 GB
Released by Qing Yao through its official repository on Hugging Face. Read the license.
Built From
- Derived from EleutherAI/pythia-160m
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 2.9 GB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About ppt-pythia-160m-uniform250-previous_mse-seed208-stage1
How much GPU memory does ppt-pythia-160m-uniform250-previous_mse-seed208-stage1 need?
About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (85M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run ppt-pythia-160m-uniform250-previous_mse-seed208-stage1 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use ppt-pythia-160m-uniform250-previous_mse-seed208-stage1 commercially?
Yes. ppt-pythia-160m-uniform250-previous_mse-seed208-stage1 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is ppt-pythia-160m-uniform250-previous_mse-seed208-stage1's context length?
2,048 tokens, from the maximum position embeddings in its published configuration.
Similar Models
This model is a fine-tuned version of EleutherAI/pythia-160m on an unknown dataset. The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 13 - trainingsteps: 250 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1
Modelo de lenguaje entrenado desde cero (pesos aleatorios) por Victor Cherif. Arquitectura tipo Llama de ~110M de parametros, pensado como asistente en espanol orientado a ciencia, programacion, robotica y fisica, con soporte de tool-calling / MCP entrenado contra servidores MCP reales. RMSNorm + SwiGLU - Al ser un modelo pequeno (~110M), el lenguaje puede ser incoherente en respuestas largas, sobre todo pasados los primeros parrafos (tiende a repetir bloques o rellenar con generalidades). - La aritmetica mental sigue sin ser confiable: depende de que el modelo dispare la tool calculate en vez de calcular "de memoria" -- verificado que esto mejoro pero no es perfecto. - El razonamiento ( )…
DistilGPT2 (short for Distilled-GPT2) is an English-language model pre-trained with the supervision of the smallest version of Generative Pre-trained Transformer 2 (GPT-2). Like GPT-2, DistilGPT2 can be used to generate text. Users of this model card should also consider information about the design, training, and limitations of GPT-2. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. As the developers of GPT-2 (OpenAI) note in their model card, “language models like GPT-2 reflect the biases inherent to the systems they were trained on.” Significant research has explored bias and…
This model is a fine-tuned version of distilbert/distilgpt2 on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 2e-05 - trainbatchsize: 8 - evalbatchsize: 8 - lrschedulertype: linear - numepochs: 3.0 - Transformers 5.18.0 - Pytorch 2.11.0+cu130 - Datasets 4.8.5 - Tokenizers 0.23.2
This model is a fine-tuned version of distilbert/distilgpt2 on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 4 - evalbatchsize: 4 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 16 - lrschedulertype: linear - numepochs: 1 - mixedprecisiontraining: Native AMP - Transformers 5.17.0 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.23.1