This model is a custom-code derivative of AxiomicLabs/GPT-X2-125M, adapted for experimental long-context causal language modeling and architecture research. The repository includes a Hugging Face Transformers-compatible GPT-X2 implementation with optional Symplectic Metric-RoPE Governor support and training utilities built around CIxOpt, a heterogeneous optimizer developed for efficient parameter routing across large projection matrices, sensitive normalization parameters, and optional governor modules. The model is intended as a research checkpoint for compact long-context generation, positional encoding experiments, optimizer testing, and continued fine-tuning. This implementation uses a…
Open weights
apache-2.0
126M parameters
32,768 tokens
transformers
PMA-1.2 is Patriot Memory's 127.9M-parameter on-device language model. It speaks English and Traditional Chinese, answers as PMA from Patriot Memory, and fits in a 128 MB parameter budget built for edge hardware. It is a new architecture, a new tokenizer, and an order of magnitude more training, aimed at the same job: a small, fast, honest assistant for Patriot Memory and Viper Gaming questions and general chat. It provides accurate information regarding: If you run into multi-GPU tensor device mismatch errors: Run the script with CUDAVISIBLEDEVICES=0 to isolate execution to GPU 0. trustremotecode=True is required: PMA-1.2's architecture is our own and ships as small Python files next to…
Open weights
apache-2.0
126M parameters
32,768 tokens
transformers
This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-shuff-dyck-100mbseed10. It has been trained using TRL. This model was trained with SFT.
Open weights
125M parameters
transformers
This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-Dp-100mbseed10. It has been trained using TRL. This model was trained with SFT.
Open weights
125M parameters
transformers
This model is a fine-tuned version of fpadovani/arb-arab-100mb-ppt-shuff-dyck-10mbseed10. It has been trained using TRL. This model was trained with SFT.
Open weights
125M parameters
transformers
This model is a fine-tuned version of goldfish-models/swalatn100mb. It has been trained using TRL. This model was trained with SFT.
Open weights
125M parameters
transformers