SAVRN
Search Contact SAVRN

Open-weight model · Text generation

ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2

by Qing Yao qing-yao/ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2

ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 is an open-weight model for text generation from Qing Yao, released under Apache License 2.0. It has 162M parameters and a 2,048-token context. At 16-bit it needs about 0.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

This model is a fine-tuned version of qing-yao/ppt-pythia-160m-uniform250-previousmse-seed324-stage1 on the None dataset.

Parameters162M
Context2,048
Weights5.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 (162M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 on every accelerator the SAVRN Index prices, at every precision

Model Card

By Qing Yao, published under apache-2.0, revision 2f8019eefe1e.

This model is a fine-tuned version of qing-yao/ppt-pythia-160m-uniform250-previousmse-seed324-stage1 on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 500 - trainingsteps: 10000 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Read Qing Yao's full model card

This model is a fine-tuned version of qing-yao/ppt-pythia-160m-uniform250-previous_mse-seed324-stage1 on the None dataset. It achieves the following results on the evaluation set: - Loss: 3.8428

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training: - learning_rate: 0.001 - train_batch_size: 16 - eval_batch_size: 16 - seed: 324 - gradient_accumulation_steps: 2 - total_train_batch_size: 32 - optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments - lr_scheduler_type: cosine_with_min_lr - lr_scheduler_warmup_steps: 500 - training_steps: 10000

Training results

Training Loss Epoch Step Validation Loss
9.6038 0.005 50 8.1510
7.3770 0.01 100 7.0060
6.7299 0.015 150 6.5891
6.4354 0.02 200 6.3406
6.1888 0.025 250 6.1097
6.0553 0.03 300 6.3539
5.9732 0.035 350 5.8630
5.7494 0.04 400 5.7048
5.6526 0.045 450 5.6101
5.5732 0.05 500 5.5119
5.7924 0.055 550 9.3161
6.7918 0.06 600 6.2639
6.0353 0.065 650 5.8930
5.7229 0.07 700 5.6003
5.5250 0.075 750 5.4818
5.4270 0.08 800 5.3856
5.3561 0.085 850 5.3287
5.3038 0.09 900 5.3082
5.2610 0.095 950 5.2320
5.2645 0.1 1000 5.1920
5.2006 0.105 1050 5.1566
5.1024 0.11 1100 5.0896
5.0347 0.115 1150 5.0381
5.0048 0.12 1200 4.9851
4.9592 0.125 1250 4.9543
4.9163 0.13 1300 4.8828
4.8945 0.135 1350 4.8513
4.8232 0.14 1400 4.7946
4.7942 0.145 1450 4.7792
4.7668 0.15 1500 4.7186
4.7252 0.155 1550 4.6795
4.6654 0.16 1600 4.6638
4.6792 0.165 1650 4.6339
4.6466 0.17 1700 4.6007
4.6177 0.175 1750 4.5868
4.5921 0.18 1800 4.5785
4.5639 0.185 1850 4.5835
4.5461 0.19 1900 4.5197
4.5255 0.195 1950 4.4907
4.4940 0.2 2000 4.4662
4.4792 0.205 2050 4.4410
4.4653 0.21 2100 4.4147
4.4419 0.215 2150 4.4063
4.4335 0.22 2200 4.3988
4.4118 0.225 2250 4.3770
4.4041 0.23 2300 4.3611
4.3997 0.235 2350 4.3500
4.4034 0.24 2400 4.3473
4.3507 0.245 2450 4.3394
4.3947 0.25 2500 4.3459
4.3657 0.255 2550 4.3175
4.3167 0.26 2600 4.2959
4.3145 0.265 2650 4.2815
4.3056 0.27 2700 4.2750
4.2975 0.275 2750 4.2670
4.2922 0.28 2800 4.2664
4.2706 0.285 2850 4.2613
4.2690 0.29 2900 4.2457
4.2295 0.295 2950 4.2309
4.2693 0.3 3000 4.2168
4.2378 0.305 3050 4.2634
4.2645 0.31 3100 4.2508
4.2629 0.315 3150 4.2062
4.2252 0.32 3200 4.1806
4.2108 0.325 3250 4.2118
4.2372 0.33 3300 4.1715
4.1996 0.335 3350 4.1688
4.2138 0.34 3400 4.1458
4.1881 0.345 3450 4.1320
4.1725 0.35 3500 4.1248
4.1536 0.355 3550 4.1313
4.1641 0.36 3600 4.1512
4.1833 0.365 3650 4.1375
4.1498 0.37 3700 4.1139
4.1989 0.375 3750 4.1414
4.2028 0.38 3800 4.1466
4.1501 0.385 3850 4.1343
4.1407 0.39 3900 4.1087
4.1698 0.395 3950 4.1157
4.1116 0.4 4000 4.0921
4.1143 0.405 4050 4.0806
4.1162 0.41 4100 4.0680
4.0757 0.415 4150 4.0625
4.0810 0.42 4200 4.0656
4.0901 0.425 4250 4.0534
4.0616 0.43 4300 4.0465
4.0674 0.435 4350 4.0447
4.0740 0.44 4400 4.0908
4.0844 0.445 4450 4.0450
4.0762 0.45 4500 4.0368
4.0332 0.455 4550 4.0408
4.0397 0.46 4600 4.0265
4.0584 0.465 4650 4.0207
4.0525 0.47 4700 4.0172
4.0353 0.475 4750 4.0051
4.0294 0.48 4800 4.0017
4.0326 0.485 4850 3.9968
4.0292 0.49 4900 3.9927
4.0138 0.495 4950 3.9873
4.0118 0.5 5000 4.0240
4.0562 0.505 5050 3.9909
4.0363 0.51 5100 3.9825
4.0153 0.515 5150 3.9870
4.0100 0.52 5200 3.9756
3.9983 0.525 5250 3.9699
3.9871 0.53 5300 3.9633
3.9860 0.535 5350 3.9620
3.9828 0.54 5400 3.9578
3.9814 0.545 5450 3.9563
3.9922 0.55 5500 3.9539
3.9843 0.555 5550 3.9496
3.9916 0.56 5600 3.9542
3.9968 0.565 5650 3.9479
3.9754 0.57 5700 3.9409
3.9472 0.575 5750 3.9457
3.9836 0.58 5800 3.9378
3.9589 0.585 5850 3.9338
3.9799 0.59 5900 3.9359
3.9705 0.595 5950 3.9324
3.9565 0.6 6000 3.9236
3.9723 0.605 6050 3.9231
3.9667 0.61 6100 3.9217
3.9500 0.615 6150 3.9283
3.9510 0.62 6200 3.9329
3.9505 0.625 6250 3.9222
3.9665 0.63 6300 3.9171
3.9419 0.635 6350 3.9102
3.9283 0.64 6400 3.9121
3.9439 0.645 6450 3.9036
3.9209 0.65 6500 3.8993
3.9299 0.655 6550 3.8971
3.9199 0.66 6600 3.8946
3.9132 0.665 6650 3.9021
3.9114 0.67 6700 3.8953
3.9267 0.675 6750 3.9015
3.9247 0.68 6800 3.8963
3.9100 0.685 6850 3.8910
3.9107 0.69 6900 3.8877
3.9109 0.695 6950 3.8874
3.9050 0.7 7000 3.8893
3.9057 0.705 7050 3.8835
3.9555 0.71 7100 3.8790
3.9186 0.715 7150 3.8777
3.9345 0.72 7200 3.8754
3.9271 0.725 7250 3.8739
3.9292 0.73 7300 3.8720
3.9455 0.735 7350 3.8712
3.8990 0.74 7400 3.8697
3.9048 0.745 7450 3.8679
3.9063 0.75 7500 3.8664
3.8935 0.755 7550 3.8650
3.8870 0.76 7600 3.8643
3.8932 0.765 7650 3.8629
3.8986 0.77 7700 3.8617
3.9045 0.775 7750 3.8609
3.9264 0.78 7800 3.8631
3.8773 0.785 7850 3.8605
3.9106 0.79 7900 3.8586
3.9113 0.795 7950 3.8578
3.8744 0.8 8000 3.8599
3.8833 0.805 8050 3.8567
3.8780 0.81 8100 3.8552
3.8735 0.815 8150 3.8544
3.8733 0.82 8200 3.8543
3.8861 0.825 8250 3.8536
3.8917 0.83 8300 3.8522
3.8809 0.835 8350 3.8521
3.8756 0.84 8400 3.8531
3.8763 0.845 8450 3.8515
3.8847 0.85 8500 3.8508
3.9023 0.855 8550 3.8501
3.8742 0.86 8600 3.8496
3.8826 0.865 8650 3.8491
3.8691 0.87 8700 3.8499
3.8961 0.875 8750 3.8485
3.8449 0.88 8800 3.8514
3.8867 0.885 8850 3.8491
3.8905 0.89 8900 3.8479
3.8992 0.895 8950 3.8473
3.8945 0.9 9000 3.8469
3.8676 0.905 9050 3.8481
3.8669 0.91 9100 3.8468
3.8770 0.915 9150 3.8465
3.8792 0.92 9200 3.8460
3.8674 0.925 9250 3.8455
3.8835 0.93 9300 3.8453
3.8884 0.935 9350 3.8450
3.8840 0.94 9400 3.8446
3.8741 0.945 9450 3.8444
3.8782 0.95 9500 3.8440
3.8695 0.955 9550 3.8440
3.9037 0.96 9600 3.8436
3.8770 0.965 9650 3.8441
3.8731 0.97 9700 3.8437
3.8522 0.975 9750 3.8434
3.8719 0.98 9800 3.8433
3.8693 0.985 9850 3.8428
3.9135 0.99 9900 3.8429
3.8449 0.995 9950 3.8436
3.8879 1.0 10000 3.8428

Framework versions

  • Transformers 5.4.0
  • Pytorch 2.8.0+cu128
  • Datasets 3.2.0
  • Tokenizers 0.22.1

Configuration

Architecture
GPTNeoXForCausalLM
Context length (tokens)
2,048
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
50,304
Model type
gpt_neox

Identity and Version

Repository
qing-yao/ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2
Publisher
Qing Yao
Task
Text generation
Modality
Text
Library
transformers
Parameters
162M parameters
Languages
Not stated by the source
Revision
2f8019eefe1eb071399f3b147d5f527934309fb7
First published
2026-09-26
Last updated
2026-09-26

Files and Weights

66 files, 5.5 GB in total. The weights are 29 files totalling 5.5 GB in bin, pt, pth, safetensors.

Weights29 files · 5.5 GB
Configuration21 files · 347.1 KB
Tokenizer14 files · 25.0 MB
Documentation1 file · 12.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
checkpoint-10000/model.safetensorsWeights324.7 MB e99d9ad91cea
checkpoint-10000/optimizer.ptWeights649.4 MB d41a2ed3e76b
checkpoint-10000/rng_state.pthWeights14.6 KB d686d3f57bd1
checkpoint-10000/scheduler.ptWeights1.5 KB c101a5f2562c
checkpoint-10000/training_args.binWeights5.4 KB 7152f3772afe
checkpoint-2000/model.safetensorsWeights324.7 MB 7c14811a58c1
checkpoint-2000/optimizer.ptWeights649.4 MB 29d5933e77ed
checkpoint-2000/rng_state.pthWeights14.6 KB 4c07f248a4ec
checkpoint-2000/scheduler.ptWeights1.5 KB 70a8cd55d7a2
checkpoint-2000/training_args.binWeights5.4 KB 7152f3772afe
checkpoint-4000/model.safetensorsWeights324.7 MB 385357bc16c1
checkpoint-4000/optimizer.ptWeights649.4 MB 9ab59dcc4cbb
checkpoint-4000/rng_state.pthWeights14.6 KB f7ad907e449a
checkpoint-4000/scheduler.ptWeights1.5 KB 59828c1e7072
checkpoint-4000/training_args.binWeights5.4 KB 7152f3772afe
checkpoint-6000/model.safetensorsWeights324.7 MB 5bf9cdecce3a
checkpoint-6000/optimizer.ptWeights649.4 MB f1006c3dcd6b
checkpoint-6000/rng_state.pthWeights14.6 KB d9ff15a0556e
checkpoint-6000/scheduler.ptWeights1.5 KB 5359576a40af
checkpoint-6000/training_args.binWeights5.4 KB 7152f3772afe
checkpoint-8000/model.safetensorsWeights324.7 MB 66f5d4e2d185
checkpoint-8000/optimizer.ptWeights649.4 MB 24781ae86176
checkpoint-8000/rng_state.pthWeights14.6 KB a28df9a13af3
checkpoint-8000/scheduler.ptWeights1.5 KB 79dc8472cd40
checkpoint-8000/training_args.binWeights5.4 KB 7152f3772afe
final/model.safetensorsWeights324.7 MB e99d9ad91cea
final/training_args.binWeights5.4 KB 7152f3772afe
model.safetensorsWeights324.7 MB e99d9ad91cea
training_args.binWeights5.4 KB 7152f3772afe
checkpoint-10000/config.jsonConfiguration781 B —
checkpoint-10000/generation_config.jsonConfiguration225 B —
checkpoint-10000/trainer_state.jsonConfiguration74.2 KB —
checkpoint-2000/config.jsonConfiguration781 B —
checkpoint-2000/generation_config.jsonConfiguration225 B —
checkpoint-2000/trainer_state.jsonConfiguration15.4 KB —
checkpoint-4000/config.jsonConfiguration781 B —
checkpoint-4000/generation_config.jsonConfiguration225 B —
checkpoint-4000/trainer_state.jsonConfiguration30.1 KB —
checkpoint-6000/config.jsonConfiguration781 B —
checkpoint-6000/generation_config.jsonConfiguration225 B —
checkpoint-6000/trainer_state.jsonConfiguration44.8 KB —
checkpoint-8000/config.jsonConfiguration781 B —
checkpoint-8000/generation_config.jsonConfiguration225 B —
checkpoint-8000/trainer_state.jsonConfiguration59.4 KB —
config.jsonConfiguration781 B —
experiment.jsonConfiguration41.7 KB —
final/config.jsonConfiguration781 B —
final/generation_config.jsonConfiguration225 B —
generation_config.jsonConfiguration225 B —
trainer_state.jsonConfiguration74.4 KB —
README.mdDocumentation12.1 KB —
.gitattributesRepository1.5 KB —
checkpoint-10000/tokenizer.jsonTokenizer3.6 MB —
checkpoint-10000/tokenizer_config.jsonTokenizer349 B —
checkpoint-2000/tokenizer.jsonTokenizer3.6 MB —
checkpoint-2000/tokenizer_config.jsonTokenizer349 B —
checkpoint-4000/tokenizer.jsonTokenizer3.6 MB —
checkpoint-4000/tokenizer_config.jsonTokenizer349 B —
checkpoint-6000/tokenizer.jsonTokenizer3.6 MB —
checkpoint-6000/tokenizer_config.jsonTokenizer349 B —
checkpoint-8000/tokenizer.jsonTokenizer3.6 MB —
checkpoint-8000/tokenizer_config.jsonTokenizer349 B —
final/tokenizer.jsonTokenizer3.6 MB —
final/tokenizer_config.jsonTokenizer349 B —
tokenizer.jsonTokenizer3.6 MB —
tokenizer_config.jsonTokenizer349 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
5.5 GB
Download from Qing Yao

Released by Qing Yao through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published5.5 GB
16-bit0.3 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2

How much GPU memory does ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 need?

About 0.4 GB at 16-bit and 0.1 GB at 4-bit: the weights (162M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 commercially?

Yes. ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is ppt-pythia-160m-uniform250-previous_mse_delta_shuffle1-seed324-stage2's context length?

2,048 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This model is a fine-tuned version of qing-yao/ppt-pythia-160m-uniform250-previousmse-seed324-stage1 on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 500 - trainingsteps: 10000 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Open weights apache-2.0 162M parameters 2,048 tokens transformers

Model · Text generation

astrallm

Pihu Pandey

AstralLM is a compact, efficient causal language model built for Hugging Face Transformers. It features a clean decoder-only transformer architecture with Grouped Query Attention (GQA), SwiGLU feed-forward networks, RoPE positional embeddings, and per-head QK-norm — delivering strong performance at a small parameter count (~138M). AstralLM is a decoder-only transformer with the following design choices: The model uses 8 query heads and 4 key/value heads (2:1 ratio), halving the KV cache memory footprint during inference without measurable quality degradation. Each head operates over a head dimension of 80. Per-head RMSNorm is applied to both query and key projections before the rotary…

Open weights apache-2.0 160M parameters 3,072 tokens transformers

Model · Text generation

Ru-Small-Instruct

LongTime

Ru-Small-Instruct — экспериментальная компактная русскоязычная языковая модель класса SLM (Small Language Model) с объемом параметров ~0.2B (~165M). Разработана с упором на суверенность весов (Zero-Fingerprint): модель обучена с нуля без заимствования базовых чекпоинтов у сторонних корпоративных сетей (Llama 3 от Meta, Qwen от Alibaba, Mistral). Модель предназначена для исследований локального инференса, работы на маломощном оборудовании, CPU и мобильных чипах, где критичны нулевая задержка (Time-To-First-Token) и полная независимость весов. Для компактной модели в 165M параметров, обученной на одном домашнем GPU за 48 часов, способность держать роль, грамотно формулировать сложные термины…

Open weights mit 165M parameters 512 tokens transformers

Model · Text generation

slm-125m-dpo

Tijani Ohiokpehai

tohio/slm-125m-dpo is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 1,024 tokens transformers

Model · Text generation

slm-125m-sft

Tijani Ohiokpehai

tohio/slm-125m-sft is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 1,024 tokens transformers

Model · Text generation

slm-125m-base

Tijani Ohiokpehai

tohio/slm-125m-base is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 2,048 tokens transformers