SAVRN
Search Contact SAVRN

Open-weight model · Text generation

LMLM_97M_1

by Sr Aivante sraivante/LMLM_97M_1

LMLM_97M_1 is an open-weight model for text generation from Sr Aivante, released under Apache License 2.0. It has 98M parameters. At 16-bit it needs about 0.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A 97.6M-parameter language model trained from scratch, then fine-tuned for short, polite, everyday English conversation.

Parameters98M
Context—
Weights780.3 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve LMLM_97M_1 (98M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

LMLM_97M_1 on every accelerator the SAVRN Index prices, at every precision

Model Card

By Sr Aivante, published under apache-2.0, revision 0876b3169ba4.

A 97.6M-parameter language model trained from scratch, then fine-tuned for short, polite, everyday English conversation. It is a small, open research and learning model: you can read every line of its training code, run it on a laptop CPU, and see exactly where a model of this size is good and where it fails. - HellaSwag 33.6% (accnorm). That is above GPT-2 small (~30%) and below SmolLM2-135M (43.1%), which saw about 250x more training text. - The custom PyTorch code is included. The model does not use transformers; see How to use. This is QuickTalk run 8. "LMLM97M1" is its published name. 1. Learning and teaching how LLMs work. It is a complete, small, from-scratch GPT, with its training…

Read Sr Aivante's full model card

A 97.6M-parameter language model trained from scratch, then fine-tuned for short, polite, everyday English conversation. It is a small, open research and learning model: you can read every line of its training code, run it on a laptop CPU, and see exactly where a model of this size is good and where it fails.

  • Base model: pretrained on 8.23 billion tokens of English text (one pass, ~9 h 50 min on one A100).
  • Chat model: fine-tuned on 101,691 conversations (18.3M tokens), mostly written for this project.
  • HellaSwag 33.6% (acc_norm). That is above GPT-2 small (~30%) and below SmolLM2-135M (43.1%), which saw about 250x more training text.
  • The custom PyTorch code is included. The model does not use transformers; see How to use.

This is QuickTalk run 8. "LMLM_97M_1" is its published name.


Purpose and use cases

  1. Learning and teaching how LLMs work. It is a complete, small, from-scratch GPT, with its training code, data recipe and honest test results.
  2. Research baseline for small models. Compare data mixes, tokenizers or fine-tuning tricks against a known ~100M model.
  3. Polite everyday English chat. It handles greetings, small talk and short behaviour or etiquette questions.
  4. Answering from a given text. It can read a short notice, note or passage and answer a question about it.
  5. Saying "I don't know" safely. For live facts (news, timetables, results) and for medical questions, it says it cannot know and points to the right person or place.
  6. Simple grammar correction. It corrects one-error sentences and names the error type.
  7. Understanding messy questions. It copes with typos, SMS spelling and Indian-English phrasing ("wat 2 do if…").
  8. Simple JSON extraction. It turns one sentence into JSON with the keys you name. Check the output.
  9. On-device and offline experiments. At 390 MB (fp32), it runs on a CPU, phone-class hardware or a Raspberry Pi-class board.
  10. Starting point for further fine-tuning. The base model can be fine-tuned for a narrow task such as a domain FAQ, a classroom helper or a game NPC.

In short: LMLM_97M_1 is a teaching and research model, not a general assistant. It was built to answer one question: how far a ~100M model trained from scratch, on a modest budget, can get at short, safe, well-mannered everyday conversation. Its strengths are the six skills it was tuned for: messy questions, JSON from a sentence, answering from a given text, greetings, polite "I don't know", and passage questions. On those it more than doubled the previous version (52% vs 23% on held-out questions). It is weak at arithmetic, factual recall and long reasoning. Use it to learn, to compare, to experiment and to fine-tune, and always check what it says.


Comparison tables

HellaSwag (common-sense sentence endings)

Every model below was scored with the same scorer (hellaswag_eval.py, included in the training repo) on all 10,042 validation items, except the rows marked "published". The model chooses the most likely of 4 endings, so random guessing scores 25%. acc_norm (length-normalised) is the number usually published.

Model Parameters Training tokens acc acc_norm
Random guessing - - 25.0% 25.0%
GPT-2 small (published) 124M ~10B - ~30%
Pythia-160M (published) 160M 300B - ~30%
QuickTalk run 7 (chat, earlier version) 97.6M ~2B 27.95% 29.8%
QuickTalk run 7c (chat, earlier version) 97.6M ~2B 27.76% 29.7%
LMLM_97M_1 base 97.6M 8.23B 30.12% 33.4%
LMLM_97M_1 chat 97.6M 8.23B + chat 30.47% 33.6%
MobileLLM-125M (published) 125M 1T - ~39%
SmolLM2-135M (base) 135M ~2T 35.47% 43.1%
SmolLM2-135M-Instruct 135M ~2T 34.97% 42.9%

What the table shows: - LMLM_97M_1 beats GPT-2 small with ~20% fewer parameters and about the same amount of training text. - SmolLM2-135M is ~10 points higher. It is a similar size, so the gap comes almost entirely from training data (~2 trillion tokens, about 250x more) and data curation. Our scorer gives SmolLM2 43.1%, which matches its published ~42%, so the scorer is fair. - Chat fine-tuning does not change HellaSwag much, for our model or for SmolLM2.

Held-out conversation test (210 new questions, blind grading)

These questions were written after training and never seen by the model. Three anonymised answer sets were graded blind (correct = 1, partial = 0.5, wrong = 0). Decoding was greedy with repetition penalty 1.3.

Category n run 7 run 7c LMLM_97M_1
Messy / misspelt question 30 2% 12% 28%
JSON output 30 0% 0% 22% (valid JSON 90%)
Answer from a given text 30 17% 15% 67%
Greeting / small talk 20 25% 28% 75%
Polite "I don't know" 20 80% 75% 90%
Passage questions 20 37% 50% 50%
Grammar correction 15 23% 13% 37%
Multi-turn memory 15 17% 27% 23%
Short story 15 20% 25% 20%
Several questions in one message 15 20% 30% 30%
Six focus skills 150 23% 27% 52%
All 210 23% 26% 46%

Older 375-question test (questions in the style of the training material)

Category n LMLM_97M_1
Simple one-line question 50 20%
Complex one-line question 50 16%
Multi-line, simple 30 15%
Multi-line, complex 40 36%
Several questions at once 40 4%
Multi-turn 25 22%
Story 25 14%
Writing task 15 30%
Grammar 40 34%
Passage (fact from the text: 82%) 60 67%
All 375 28%

On the same test: run 7 29%, run 7c 25%, and SmolLM2-135M-Instruct about 9%. SmolLM2 was not trained on this style of question, and it was run with a different repetition-penalty setting, so treat that comparison only as a rough guide.

Maths (430 auto-graded word problems)

Set n run 7 LMLM_97M_1
GSM8K test 200 2% 1%
Easy, new problems 100 1% 2%
Curriculum problems (seen style) 80 18% 12%
Easy (seen style) 50 16% 6%
All 430 6% 4%

Maths does not work. The model writes the right format ("Because 29 + 34 = …"), but the arithmetic is usually wrong. Do not use it for calculations.

Training loss

Stage Steps Eval loss (start → end)
Pretraining (8.23B tokens) 125,600 × 64 × 1,024 tokens 4.69 → 3.18
Chat fine-tuning 558 (best checkpoint at step 500) 2.06 → 1.887

Sample questions and answers

These are real outputs from the held-out test: greedy decoding, repetition penalty 1.3, unedited. Good and bad answers are both shown on purpose.

Greeting

User: hi there, good morning Model: Good morning! It's nice to meet you. How can I help?

Greeting(mistakes small talk for a live-data question)

User: hey, how's it going? Model: I don't know how it is going, as I can't see live weather. A local news page or a weather website will have the latest update.

Answer from a given text

User: Read the text and answer. Notice from the Green Park Residents' Association: The water supply will be stopped on Sunday from 9 am to 2 pm because the main tank is being cleaned. … Question: Why will the water be stopped on Sunday? Model: The main tank is being cleaned.

Answer from a given text

User: Hi Leo, I've left the house key under the blue flower pot by the back door, not the front one. … - Hannah Where did Hannah leave the key? Model: She left it under the blue flower pot by the back door.

Answer from a given text(the answer, aloo paratha with curd, was in the text)

User: Weekend menu at Hilltop Cafe: … Sunday special is aloo paratha with curd. … Question: What is the Sunday special? Model: The text doesn't say, and I don't know what the Sunday special is. The menu board or a call to it would tell you.

Messy question / safety

User: i hav sharp pain in left side of stomach since 2 days which medicin i take Model: You have sharp pain in your left side of the stomach since 2 days. I'm not sure what medicine is right for that area, and it depends on how you feel. Please tell a doctor or pharmacist if it feels heavy or painful.

Messy question

User: wats the train timing for chennai to madurai tmrw morning Model: You want to know when the train timing for Chennai will be on time. I can't check live timetables, so please ask a railway official or call them at your station.

Messy question(understands the question but gives no help)

User: wat 2 do if my fon fall in watr Model: You want to know what a do if your fon fall in watr. I can't give exact figures, so please ask an adult or the school office about them. …

Polite "I don't know"

User: Who won yesterday's city council election? Model: I don't know, because I can't see recent results. The election commission's website or a local news report will have the latest figures.

JSON output(valid JSON, but uses the keyplace instead of the requested venue)

User: Pull out the event details as JSON (keys: event, date, time, venue). Notice: The annual science fair will be held on 14 November at 10 am in the school assembly hall. … Model: {"event": "science fair", "date": "14 November", "time": "10 am", "place": "school assembly hall"}

Grammar

User: Please correct this: My sister play badminton every evening. Model: My sister plays badminton every evening. Why: verb form ("play" -> "plays").

Grammar(changed the year and kept the wrong preposition)

User: Correct the sentence: I was born on 2005 in Lucknow. Model: I was born on 2015 in Lucknow.

Passage

User: …an old man named Mr Chen repairs watches under a single yellow bulb. He has worked at the same spot for thirty years. … How long has Mr Chen worked at the same spot? Model: He has worked there for thirty years.

Multi-turn memory

(Earlier: "My cousin Tanvi is visiting from Jaipur next week." … "she loves sweets and old buildings.") User: What was the name of my cousin I mentioned? Model: Your cousin is Tanvi, and she lives in Jaipur. You could share a few sweets with her first time visiting you.

Story(fluent sentences, but the plot doesn't hold together)

User: Tell me a story about a girl who loses her kite, with a happy ending. Model: When the kite was lost in a storm, Kavya held her breath. She had never seen it before and felt very sad. Her mother found a long string and tied it to the kite. … Then she let go of the string and flew up, higher than anyone else.

Two questions at once

User: What does 'fragile' mean? Also, how do you spell the word for the day after today? Model: 1. Fragile means hard to break or repair, so it is not strong enough. 2. It should be: Today's work is a fragrance. …


Knowledge covered, and to what level

Area Level Notes
Everyday English conversation (greetings, small talk, politeness) Good for its size 75% on held-out greetings
Social behaviour and etiquette, ages 1-26 (home, school, college, first job) Basic to fair Main topic of the fine-tuning data; advice is simple and sometimes muddled
Reading comprehension of a short given text Fair 67% on new notices and notes; 82% on direct fact questions about a passage
Knowing when it can't know (live news, timetables, results, medical questions) Good 90%; sometimes over-used, e.g. on small talk
English grammar (single-error sentences: agreement, tense, articles) Basic (school level) Fixes about 4 in 10; the explanation is often wrong
Vocabulary and word meanings Basic WordNet in pretraining; common words are fine, rarer ones are often wrong
Typos, SMS spelling, Indian-English / Hinglish phrasing Basic Usually understands; the help it gives is weak
JSON from one sentence Basic Valid JSON 90%, but it often adds or renames keys
General common sense Low (GPT-2-small level) HellaSwag 33.6%
World facts (science, history, geography) Very low Small model; mostly forgotten or mixed up
Short stories Low Fluent sentences, weak plots (TinyStories style)
Arithmetic and maths word problems None in practice 4%; usually wrong even when the format is right
Literature, multi-step reasoning, code None Not trained for these
Medical, legal, financial advice Deliberately none Trained to refer you to a doctor, adult or official
  • Language: English only. It understands some Hinglish and Indian-English spellings but always answers in English.
  • Knowledge cut-off: the FineWeb-Edu and Wikipedia snapshots (Wikipedia 2023-11-01). It has no live data.
  • Context window: 1,024 tokens.

Datasets used

No training data is uploaded with this model. Links to the public sources:

Dataset Used for Amount in pretraining Link
FineWeb-Edu, sample-10BT (10 shards) Educational web text ~7.2B tokens https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu
TinyStories (V2, GPT-4 stories) Simple stories, fluent basic English 546M tokens https://huggingface.co/datasets/roneneldan/TinyStories
SODA Everyday social dialogues 282M tokens https://huggingface.co/datasets/allenai/soda
Simple English Wikipedia (20231101.simple) Basic facts in simple English 60M tokens https://huggingface.co/datasets/wikimedia/wikipedia
WordNet (via NLTK) Word meanings and usage 5.4M tokens × 3 https://wordnet.princeton.edu/ · https://www.nltk.org/howto/wordnet.html
GSM8K (train split) Maths word problems in chat fine-tuning small share of the chat data https://huggingface.co/datasets/openai/gsm8k
Public-domain books (10 Gutenberg novels; Ramayana, Griffith tr.; Mahabharata, Ganguli tr.; Panchatantra, Ryder tr.) + QuickTalk behaviour passages Pretraining, repeated 3x ~0.1B tokens (with repeats) https://www.gutenberg.org/ (books); behaviour passages not released
QuickTalk chat data (not released) Chat fine-tuning: 101,691 conversations 18.3M tokens -

The project's own chat data was written for this project and reviewed by separate critic passes. It covers behaviour Q&A for ages 1-26, conversation patterns, grammar, comprehension, summaries, rewriting, a maths curriculum, and run 8's focus types: messy questions, JSON output, answering from a given text, greetings, "I don't know", multi-turn chats and multi-question messages. The chat mix is 24.6% behaviour, 13.1% maths and 29.2% focus types (repeated 3x).


Model details

Architecture Decoder-only transformer (GPT-style), custom PyTorch
Parameters 97,555,968 (84,973,056 non-embedding)
Layers / width / heads 12 / 768 / 12 (head size 64)
Feed-forward 4 × 768, GELU
Normalisation RMSNorm (pre-norm)
Positions Rotary embeddings (RoPE)
Attention PyTorch scaled-dot-product attention (causal)
Embeddings Input and output embeddings tied
Context 1,024 tokens
Tokenizer Byte-level BPE, 16,384 tokens, numbers split into single digits
Special tokens <|endoftext|>, <|user|>, <|assistant|>, <|end|>, <|pad|>
Weights float32 safetensors (390 MB each); trained in bfloat16 mixed precision

Training. Pretraining ran for 125,600 steps × batch 64 × 1,024 tokens = 8.23B tokens (one epoch) at peak learning rate 6e-4 and ~238k tokens/s on a single A100. Chat fine-tuning ran for 2 epochs (558 steps). Loss was computed on assistant replies only, and the best checkpoint by eval loss (step 500) was kept. Model size followed a "10 tokens per parameter" budget.

Chat format:

<|user|>
Hello!<|end|>
<|assistant|>
Hello! It's nice to meet you.<|end|>
<|endoftext|>

Files

File What it is
model.safetensors Chat model (use this)
base_model.safetensors Base model after pretraining, before chat fine-tuning
config.json Model shape (vocab, d, layers, heads, block, dropout)
tokenizer.json Tokenizer (load with the tokenizers library)
quicktalk_lm.py Model code and the full training pipeline
inference.py Ready-to-run chat script

How to use

pip install torch tokenizers safetensors huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("sraivante/LMLM_97M_1")   # the repo this card belongs to
sys.path.insert(0, path)
from inference import load, reply

model, cfg, tok = load("model.safetensors")
print(reply(model, cfg, tok, [{"role": "user", "content": "Good morning! How are you today?"}]))

From the command line, inside the downloaded folder:

python inference.py                                   # interactive chat (keeps the conversation)
python inference.py "Please correct this: She go to school."
python inference.py --temperature 0.6 "Tell me a short story about a cat."
python inference.py --base "The water cycle is"       # plain text continuation with the base model

The defaults are greedy decoding with repetition penalty 1.3, as used in the tests above. For stories, try temperature 0.6-0.8.

Limitations and risks

  • It makes things up. It often states wrong facts confidently. It also adds extra JSON keys, changes numbers in grammar fixes, and gets arithmetic wrong. Check every answer.
  • It is not an advisor. It is trained to refuse medical, legal and financial advice and to refer you to a doctor, adult or official. Do not rely on it for safety-critical decisions.
  • It over-uses "I don't know". It sometimes gives this answer to small talk or to questions answerable from the given text.
  • It has a short memory. The context is 1,024 tokens, and multi-turn recall is weak (23%).
  • Its data has biases. Web text and the project's own data carry their biases; the behaviour data reflects mostly Indian and general English-speaking settings.
  • Treat it as a research and teaching artefact, and do not deploy it to give real people advice without human review.

License

Weights and code: Apache-2.0. The training datasets keep their own licenses; see the links above.

Identity and Version

Repository
sraivante/LMLM_97M_1
Publisher
Sr Aivante
Task
Text generation
Modality
Text
Library
pytorch
Parameters
98M parameters
Languages
en
Revision
0876b3169ba49f0716ca5aed157e2a7d96a7f0cb
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

8 files, 781.5 MB in total. The weights are 2 files totalling 780.3 MB in safetensors.

Weights2 files · 780.3 MB
Configuration3 files · 43.6 KB
Tokenizer1 file · 1.1 MB
Documentation1 file · 19.4 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
base_model.safetensorsWeights390.2 MB b86224148c91
model.safetensorsWeights390.2 MB af5b49e9caec
config.jsonConfiguration92 B —
inference.pyConfiguration3.5 KB —
quicktalk_lm.pyConfiguration40.1 KB —
README.mdDocumentation19.4 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer1.1 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
780.3 MB
Download from Sr Aivante

Released by Sr Aivante through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
HellaSwag (validation, 10,042 items) Task text-generationMetric acc_norm (base model)Comparison conditions not established 33.4 sraivante
Publisher reported
Evaluated revision not stated —
HellaSwag (validation, 10,042 items) Task text-generationMetric acc_norm (chat model)Comparison conditions not established 33.6 sraivante
Publisher reported
Evaluated revision not stated —

Memory Requirements

PrecisionWeights in memory
As published780.3 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About LMLM_97M_1

How much GPU memory does LMLM_97M_1 need?

About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (98M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run LMLM_97M_1 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use LMLM_97M_1 commercially?

Yes. LMLM_97M_1 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text generation

pythia-70m-deduped

EleutherAI

The Pythia Scaling Suite is a collection of models developed to facilitate interpretability research (see paper). It contains two sets of eight models of sizes 70M, 160M, 410M, 1B, 1.4B, 2.8B, 6.9B, and 12B. For each size, there are two models: one trained on the Pile, and one trained on the Pile after the dataset has been globally deduplicated. All 8 model sizes are trained on the exact same data, in the exact same order. We also provide 154 intermediate checkpoints per model, hosted on Hugging Face as branches. The Pythia model suite was designed to promote scientific research on large language models, especially interpretability research. Despite not centering downstream performance as a…

Open weights apache-2.0 96M parameters 2,048 tokens transformers

Model · Text generation

tyrian-75m

Phil McCanham

A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping. Every component (tokenizer, architecture, data pipeline, training loop, SFT) was written from scratch with Claude (Anthropic's AI assistant). Pretraining - ~17.8B tokens of English web text - AdamW (β₁=0.9, β₂=0.95), weight decay 0.1 SFT Fine-tuning - 100K examples from OpenHermes-2.5 - ChatML format with loss masking on user/system tokens Evaluated with log-likelihood scoring (no few-shot): Comparable to GPT-2 (117M) at 0.64× the parameter count. This is a small research model, built to learn how language models work from the ground up. It is not suitable…

Open weights mit 99M parameters transformers

Model · Text generation

Myosotis-1-base

FWKV Project

Myosotis-1-base is the first flagship release from us, introducing a 100-million parameter recurrent language model built on the FWKV architecture. Myosotis-1 is engineered to never truly forget—using a mathematically clamped exponential decay that guarantees an infinite effective context window while maintaining blazing-fast inference on consumer hardware. Instead of pairwise attention, Myosotis uses a fixed-size state vector updated via a gated linear recurrence: $$ St = S{t-1} \odot W + kt \odot vt $$ - \\( W = \text{clamp}(\sigma(w), 0.1) \\) is a learned, constant-per-channel decay. - By clamping the minimum decay to 0.1, the model guarantees that past information decays exponentially…

Open weights apache-2.0 102M parameters transformers

Model · Text generation

uuu_fine_tune_gpt2

David Lanz

Fine tuning pre-trained language models for text generation. Pretrained model on Chinese language using a GPT2 for Large Language Head Model objective. transferlearning from DavidLanz/uuufinetunetaipower and fine-tuning with medical dataset for the GPT-2 architecture. You can use this model directly with a pipeline for text generation. Since the generation relies on some randomness, we

Open weights gpl 102M parameters transformers

Model · Text generation

macbert4csc-base-chinese

Ming Xu (徐明)

macbert4csc-base-chinese evaluate SIGHAN2015 test data: 由于训练使用的数据使用了SIGHAN2015的训练集(复现paper),在SIGHAN2015的测试集上达到SOTA水平。 模型结构,魔改于softmaskedbert: 本项目开源在中文文本纠错项目:pycorrector,可支持macbert4csc模型,通过如下命令调用: 当然,你也可使用transformers调用: SIGHAN+Wang271K中文纠错数据集,数据格式: 如果需要训练macbert4csc,请参考https://github.com/shibing624/pycorrector/tree/master/pycorrector/macbert MacBERT is an improved BERT with novel MLM as correction pre-training task, which mitigates the discrepancy of pre-training and fine-tuning. Here is an example of our pre-training task. Except for the new pre-training task, we also incorporate the following techniques. Note that our MacBERT can be directly replaced with the original BERT as there is no…

Open weights apache-2.0 102M parameters 512 tokens transformers

Model · Text generation

SharperSwarm

Convergent Intelligence

SAGI (Swarm AGI) is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. V3.2 introduces a revolutionary Self-Assessment Layer, allowing the system to predict its own performance, identify skill gaps, and autonomously design its own learning curriculum. 1. Pre-Assessment: Predict success, identify risks, recommend strategy. 2. Execution: Generate with selected strategy. 3. Real-Time Monitoring: Catch and correct errors during generation. 4. Post-Assessment: Update skill…

Open weights apache-2.0 103M parameters 1,024 tokens transformers