Independent publisher
Shreyansh singh
shreyansh12183
AI researcher and systems builder exploring domain-adapted Small Language Models (SLMs) for Indian statutory law, semiconductor RTL design, computational biochemistry, and formal mathematical reasoning.
Models
Shreyansh-STEM-AI-2B-v3 is a sovereign compact foundation model for scientific, physical, and mathematical derivation, engineered through SOLAR-style Depth Up-Scaling (DUS) and Continual Pre-Training (CPT) Seam Healing. To surpass standard 2B parameter capacity without requiring training from scratch, intermediate transformer layers were duplicated and spliced, expanding the model depth to 22 layers with a hidden dimension of $d=2048$. Splicing transformer blocks introduces interface discontinuity along the residual stream. To heal these seams, the model underwent Continual Pre-Training (CPT) over the multi-gigabyte shreyansh-1B-SLM-pretrain-stem-english corpus (over 2,400 parquet shards).…
Vidhi-AI is a specialized instruction-tuned legal reasoning prototype focused on Indian Jurisprudence, the Constitution of India, Bharatiya Nyaya Sanhita (BNS), and Supreme Court case-law precedents.
This repository contains official GGUF quantizations of the Vigyan AI STEM Model Series engineered for 100% offline, air-gapped on-device deployment across consumer hardware, laptops, and edge devices. These checkpoints represent intermediate research iterations developed by Vigyan AI to evaluate sovereign neuro-symbolic SLM execution combined with deterministic symbolic solvers (SymPy) and embedded C++ GraphRAG.
Vigyan-7B-PhD-Pure-Math is a domain-specialized LoRA adapter on OLMo-2-1124-7B optimized for rigorous axiomatic mathematics, abstract algebra, and topology.
Vigyan-7B-BioMed-Chem is a domain-specialized LoRA adapter trained on OLMo-2-1124-7B dedicated to organic chemical synthesis, pharmacology, and molecular biology.
Vigyan-7B-Silicon-RTL-EDA is a domain-specialized Low-Rank Adaptation (LoRA) adapter built upon AllenAI's open-weights OLMo-2-1124-7B architecture. It is fine-tuned specifically for Electronic Design Automation (EDA), register-transfer level (RTL) Verilog synthesis, and formal verification assertions.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). - PEFT 0.19.1
Vigyan-7B-Astro-Logic is a domain-specialized LoRA adapter for OLMo-2-1124-7B fine-tuned for orbital mechanics, gravitational physics, and celestial dynamics.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). - PEFT 0.21.0
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
Vigyan-1.5B 4× MoE is an experimental Sparse Mixture of Experts model upcycled from 4 specialized domain LoRA adapters (Science, Technology, Engineering, Mathematics). It was assembled to evaluate whether post-hoc expert stitching on compact language models ($\le 3\text{B}$) can deliver multi-domain specialization at single-expert inference latency with a sub-1GB RAM footprint. A pre-quantized standalone GGUF (vigyan-1.5b-4x-moe.Q4KM.gguf) is available directly within this repository for instantaneous local execution on CPU or mobile. Test this model on Google Colab with an embedded Gradio chat interface: - shreyansh12183/vigyan-1.5b-adapter-science…
Vigyan-32B Titan is a heavyweight sovereign model designed for graduate-level scientific reasoning, semiconductor VLSI timing closure, aerospace astrodynamics, and formal mathematical proofs.
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains an experimental domain-specialized LoRA adapter trained as part of the Vigyan AI Sovereign Mixture of Experts (MoE) Upcycling Initiative. The goal of this experimental line was to train isolated, high-rank domain experts on specialized STEM corpora and examine whether post-hoc upcycling via uncalibrated linear routers (mergekit-moe) could synthesize dense reasoning models into edge-deployable sparse MoE networks. During post-training MoE upcycling experiments, this adapter was used in a 4-expert sparse mixture configuration. The experimental run revealed pivotal architectural insights: 1. Attention-FFN Desynchronization: Because this LoRA was trained across both…
This repository contains experimental Group Relative Policy Optimization (GRPO) LoRA weights trained to evaluate test-time compute and reinforcement learning on mathematical, physical, and scientific reasoning. Test this model instantly on Google Colab with an embedded Gradio chat interface: Released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. Academic research, university coursework, student study, and personal experimentation. Any commercial deployment, paid SaaS API, commercial tutoring platform, or enterprise wrapper requires an explicit commercial license.
Vigyan-3B 4× MoE is an experimental Sparse Mixture of Experts model upcycled from 4 specialized domain LoRA adapters (Science, Technology, Engineering, Mathematics). It was assembled to evaluate whether post-hoc expert stitching on compact language models ($\le 3\text{B}$) can deliver multi-domain specialization at single-expert inference latency. - shreyansh12183/vigyan-3b-adapter-science - shreyansh12183/vigyan-3b-adapter-technology - shreyansh12183/vigyan-3b-adapter-engineering - shreyansh12183/vigyan-3b-adapter-mathematics 1. Delimiter Blindness: Without general conversational anchor replay during adapter training, the model degrades into token echo loops when receiving general chat…
Shreyansh Singh · Founder, ExperimentLab.in · Varanasi, India GitHub · ExperimentLab.in · License: CC BY-NC 4.0 Fine-tuned LoRA adapters on OLMo-2-7B for specialized scientific sub-domains: All models released under CC BY-NC 4.0 — free for research, non-commercial use.
A GGUF-quantized Mixture-of-Experts (MoE) variant of the Vigyan-7B STEM model, designed for efficient on-device inference. - For research and non-commercial use only (CC-BY-NC-4.0) - Experimental MoE architecture — use production Vigyan-7B-STEM-Instruct for stable inference If you use this model, please cite the Vigyan AI project
Developed by Vigyan AI in Varanasi, Uttar Pradesh, India. Vigyan-7B-STEM-DPO-v1 is a specialized 7-billion parameter language model trained on deep engineering and physical sciences. Unlike general-purpose foundation models that guess mathematical calculations through token auto-regression, Vigyan-7B is aligned using Direct Preference Optimization (DPO) to: 1. Eliminate Arithmetic Drift: Active negative preference margins suppress probabilistic token guessing for quantitative operations. 2. Enforce Physical Conservation Laws: Enforces physical invariant guardrails (e.g., $|\Gamma| \le 1.0$, $\eta{\text{Carnot}} 0$, $v{\text{orbit}} > 0$). 3. Structured Tool Invocations: Trained to emit…
Developed by Vigyan AI in Varanasi, Uttar Pradesh, India. Vigyan-2B-STEM-Instruct-v1 is India's sovereign edge Small Language Model (SLM) engineered specifically for low-latency, edge-deployable STEM reasoning. At ~2.4 billion parameters, it delivers: 1. Ultra-Low VRAM Footprint: Requires only ~1.8 GB VRAM in 4-bit NF4 precision, making it deployable on consumer laptops, edge gateways, Raspberry Pi 5, and mobile devices. 2. Sub-Second Token Generation: Yields 45–60 tokens/second on a single GPU, enabling real-time interactive STEM tutoring and on-device circuit debugging. 3. Structured Tool Calling: Trained to invoke deterministic Python execution blocks ( ) for exact arithmetic…