SAVRN
Search Contact SAVRN

Open-weight model · Text generation

trackfit-llm-small

by Deepak Samuel deepaksamuel-cuk/trackfit-llm-small

trackfit-llm-small is an open-weight model for text generation from Deepak Samuel, released under MIT License. It has 32M parameters and a 128-token context. At 16-bit it needs about 0.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A 31.9M-parameter Llama-architecture causal language model trained from scratch to reconstruct particle-track parameters from detector hit patterns, treating track fitting as a language translation problem: an array of encoded hit positions is translated into…

Parameters32M
Context128
Weights127.8 MB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve trackfit-llm-small (32M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

trackfit-llm-small on every accelerator the SAVRN Index prices, at every precision

Model Card

By Deepak Samuel, published under mit, revision d02d44e92d3d.

A 31.9M-parameter Llama-architecture causal language model trained from scratch to reconstruct particle-track parameters from detector hit patterns, treating track fitting as a language translation problem: an array of encoded hit positions is translated into a short domain-specific "call" that regenerates it. A detector stack has 12 layers, each with 32 strips. A charged particle (e.g. a cosmic muon) crosses the stack roughly in a straight line, hitting strip round(slp layer + icpt) in every layer it geometrically passes through (slp = slope in strips/layer, icpt = intercept in strips). Each hit is encoded as a single integer layer 32 + strip in [0, 383]. - add — spurious hits not on the…

Read Deepak Samuel's full model card

A 31.9M-parameter Llama-architecture causal language model trained from scratch to reconstruct particle-track parameters from detector hit patterns, treating track fitting as a language translation problem: an array of encoded hit positions is translated into a short domain-specific "call" that regenerates it.

Task

A detector stack has 12 layers, each with 32 strips. A charged particle (e.g. a cosmic muon) crosses the stack roughly in a straight line, hitting strip round(slp * layer + icpt) in every layer it geometrically passes through (slp = slope in strips/layer, icpt = intercept in strips). Each hit is encoded as a single integer layer * 32 + strip in [0, 383].

Real data is noisy: - add — spurious hits not on the track (detector noise). - rem — genuine track hits that failed to register (detector inefficiency).

Given the (possibly noisy) sorted list of encoded input hits, the model outputs:

gen_evt(slp=[<slope>], icpt=[<intercept>], add=[<noise hit codes>], rem=[<missing hit codes>])

i.e. in one generation pass it must jointly recover the true straight-line track parameters and classify which input hits are noise and which true-track hits are missing.

Architecture

LlamaForCausalLM — 8 layers, hidden size 512, 8 attention heads, ~31.95M parameters — trained from scratch with a custom fixed-vocabulary tokenizer (TrackCallTokenizer, 4569 tokens: the 384 hit codes, gen_evt/slp/icpt/ add/rem/syntax tokens, and numeric literals for slp/icpt). custom_tokenizer.py in this repo has the full implementation.

This checkpoint is best_by_full_20k from epoch 777 of continued training (run small_transformer_cosmic_aligned_cont1000). Validation score at checkpoint time: perfect=0.8993 good=0.9091 S=0.9857 D=0.0122.

Results

Synthetic test set (19,270 events, known ground truth)

metric all clean add rem add+rem
parsed % 100.00 100.00 100.00 100.00 100.00
mean similarity S 0.997 1.000 0.992 1.000 0.997
perfect (S=1, D=0) % 97.44 99.98 92.85 100.00 97.50
exact call % 62.57 100.00 80.00 42.43 28.86
model |Δslope| median 0.000 0.000 0.000 0.010 0.030
OLS-fit |Δslope| median 0.035 0.014 0.055 0.029 0.086
model |Δintercept| median 0.000 0.000 0.000 0.060 0.180
OLS-fit |Δintercept| median 0.196 0.081 0.305 0.162 0.478

The model recovers exact track parameters (median error = 0) even on noisy events, beating an independent per-hit ordinary-least-squares fit by roughly an order of magnitude in dispersion (Gaussian-fit σ: 0.049 vs. 0.504 strips/layer for slope, 0.271 vs. 1.871 strips for intercept) and lands closer to the truth than the fit in ~83% of individual events. The gap is driven by noisy events: the fit weights every input hit equally, so a single injected noise hit can pull the fitted line far from the true track, while the model implicitly classifies which hits are noise before "fitting" the rest.

Real cosmic-muon events (100,000 events, no ground truth — compared against two independent fits)

metric value
parsed % 99.98
mean similarity S 0.982
perfect (S=1, D=0) % 88.87
|Δslope| vs. reference (per-layer-mean) fit, median 0.067
|Δintercept| vs. reference fit, median 0.404

Usage

from transformers import LlamaForCausalLM
from custom_tokenizer import TrackCallTokenizer  # included in this repo

tokenizer = TrackCallTokenizer.from_pretrained("deepaksamuel-cuk/trackfit-llm-small")
model = LlamaForCausalLM.from_pretrained("deepaksamuel-cuk/trackfit-llm-small")

hits = [29, 30, 122, 123, 153, 154, 184]        # encoded hit codes (layer*32 + strip), sorted, deduped
prompt = "<s>[" + ",".join(str(h) for h in hits) + "]<CALL>"
input_ids = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).input_ids
out = model.generate(input_ids, max_new_tokens=56, do_sample=False,
                      pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(out[0, input_ids.shape[1]:], skip_special_tokens=True))
# gen_evt (slp=[-1.09], icpt=[12.98], add=[...], rem=[...])

TrackCallTokenizer (custom_tokenizer.py, included here) is a from-scratch fixed-vocabulary tokenizer, not a standard BPE/WordPiece tokenizer — AutoTokenizer will not auto-detect it; import the class directly as above.

Limitations

  • Single-track only. Trained on single-particle events; behavior on multi-track events is untested.
  • Fixed geometry. The 12-layer × 32-strip layout and the add/rem noise model are specific to this detector; the model will not generalize to a different strip/layer count without retraining.
  • No physical units. slp/icpt are in strips/layer and strips, not calibrated to any physical geometry — converting to a physical angle requires the detector's actual strip pitch and layer spacing.

Configuration

Architecture
LlamaForCausalLM
Context length (tokens)
128
Layers
8
Hidden size
512
Feed-forward size
1,536
Attention heads
8
Key/value heads
8
Head dimension
64
Vocabulary size
4,569
Model type
llama

Identity and Version

Repository
deepaksamuel-cuk/trackfit-llm-small
Publisher
Deepak Samuel
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
32M parameters
Languages
Not stated by the source
Revision
d02d44e92d3dff100f71eba616f6ce286f3347d7
First published
2026-09-23
Last updated
2026-09-23

Files and Weights

8 files, 127.9 MB in total. The weights are 1 file totalling 127.8 MB in safetensors.

Weights1 file · 127.8 MB
Configuration2 files · 933 B
Tokenizer3 files · 79.4 KB
Documentation1 file · 5.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights127.8 MB ef1cf2dbe588
config.jsonConfiguration717 B —
generation_config.jsonConfiguration216 B —
README.mdDocumentation5.2 KB —
.gitattributesRepository1.5 KB —
custom_tokenizer.pyTokenizer3.9 KB —
tokenizer_config.jsonTokenizer1.2 KB —
vocab.jsonTokenizer74.3 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
127.8 MB
Download from Deepak Samuel

Released by Deepak Samuel through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published127.8 MB
16-bit0.1 GB
8-bit0.0 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About trackfit-llm-small

How much GPU memory does trackfit-llm-small need?

About 0.1 GB at 16-bit and 0 GB at 4-bit: the weights (32M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run trackfit-llm-small on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use trackfit-llm-small commercially?

Yes. trackfit-llm-small is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is trackfit-llm-small's context length?

128 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

lightning-30m-ft

AobanZ

Lightning is a small, autoregressive transformer which utilizes FlashAttention and MHA. This model is trained on a variety of books from a dataset(300 MB). Lightning utilizes FlashAttention and AdamW for performance and capability. Lightning is designed to provide quick, coherent outputs, with the downside of limited embedding. Lightning is intended to be used for research, analysis and fine-tuning, stories and other. It is not intended to be used for professional advice, real writing or any kind of heavy work as generated outputs may be incorrect. Lightning can be used directly for text generation, experimentation, and conversational interactions. Users can provide text prompts and…

Open weights mit 29M parameters transformers