This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open weights
15M parameters
512 tokens
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open weights
11M parameters
transformers
A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below). It is not a quality release yet; the card states that plainly. Standard LLaMA block, no sliding window, no GQA: The parameter count is the learnable total: the raw safetensors sum is 14,216,640, which double-counts the tied embedding (tok.weight 12288×320 = 3,932,160) that head.weight aliases. Tied, the true count is 10,284,480. - Trained from scratch (no base model). The model learned real context — val loss 4.602 is well below the 7.38 unigram floor…
Open weights
mit
10M parameters
512 tokens
transformers
Model · Text generation
Leecz
Hush-Nano-Chat is an English, single-turn instruction-tuned version of Soulitude/Hush-Nano. It starts from the 22M-parameter pretrained model and uses supervised fine-tuning (SFT) on instruction–response pairs. Due to the model's limited parameters, its response can be inaccurate, incomplete, or inconsistent. Hush-Nano-Chat has the following features: The base model was pretrained on 8.5B tokens (8,554,042,292) drawn from the following subsets: Then it was fine-tuned on a mixture of the following datasets: unsloth/alpaca-cleaned, databricks/databricks-dolly-15k, and HuggingFaceH4/norobots. Only the assistant response and ending EOS token contribute to the training loss. Zero-shot normalized…
Open weights
apache-2.0
23M parameters
1,024 tokens
transformers
Model · Text generation
Ornith
Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…
Open weights
mit
1M parameters
262,144 tokens
transformers
This package uses the standard Transformers Llama causal-language-model architecture with OpenWALDO's schema-1 byte tokenizer. Load the tokenizer with trustremotecode=True. BOM.json inventories every release file and EU-BOM.json contains the EU GPAI training-content disclosure mapping.
Open weights
820,736 parameters
512 tokens
transformers