SAVRN
Search Contact SAVRN

Independent publisher

Tijani Ohiokpehai

tohio

Models in Library3
Datasets in Library0
Models on Hugging Face3
Followers—

Models

Model · Text generation

slm-125m-dpo

Tijani Ohiokpehai

tohio/slm-125m-dpo is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 1,024 tokens transformers

Model · Text generation

slm-125m-sft

Tijani Ohiokpehai

tohio/slm-125m-sft is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 1,024 tokens transformers

Model · Text generation

slm-125m-base

Tijani Ohiokpehai

tohio/slm-125m-base is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 2,048 tokens transformers