SAVRN
Search Contact SAVRN

qwen2.5-0.5b-topo-governed-cbp-fineweb · Model Card

qwen2.5-0.5b-topo-governed-cbp-fineweb: Model Card

Written by Frank Morales, published under apache-2.0, revision 87e0aba84982, read 2026-09-27. Shown as written; SAVRN's own facts about this model are on its page.

article: https://medium.com/ai-simplified-in-plain-english/the-architecture-of-permanence-how-topo-cbp-tamed-the-uncharted-stream-a17c1d8425af

Qwen2.5-0.5B-Topo-Governed-CBP-FineWeb

An open-weights causal transformer checkpoint evaluating the synthesis of Continual Backpropagation (CBP) and the Topological Governor (TOPO-2026) under non-stationary online web streaming.

This model validates that an autoregressive foundation model can continuously adapt to streaming data without experiencing representational drift or catastrophic forgetting on silicon.

  • Base Model: Qwen/Qwen2.5-0.5B
  • Governing Framework: TOPO-2026 Topological Governor
  • Training Protocol: Continual Backpropagation with dual-phase manifold projection
  • Streaming Dataset: HuggingFaceFW/fineweb-edu (sample-10BT)
  • Optimization Horizon: 5,000 continuous gradient steps
  • Execution Seed: 123 (deterministic across Python, NumPy, PyTorch, and CUDA)
  • Companion Code: GitHub: CDL_CBP_TOPO.ipynb

Model Summary

In Loss of plasticity in deep continual learning (Nature, 2024), Dohare, Hernandez-Garcia, Lan, Rahman, Mahmood, and Sutton demonstrated that standard gradient-based optimization systematically loses plasticity, causing representational rank collapse. While their Continual Backpropagation (CBP) algorithm preserved feature diversity by recycling low-utility units, the authors noted that CBP alone does not prevent catastrophic forgetting. Furthermore, unconstrained continual updates cause networks to drift uncontrollably—what Richard Sutton described as streaming agents "losing their minds."

qwen2.5-0.5b-topo-governed-cbp-fineweb provides an empirical resolution to this stability-plasticity impasse. By decomposing the representation manifold into an invariant prime-anchored coordinate ring and an attenuated plastic complementary subspace, the network sustains active loss descent while holding its geometric reference frame to exact zero drift.


Mathematical Specification

The architecture implements a dual-phase number-theoretic boundary condition on the token embedding manifold ($\mathbf{W}_{\text{embed}} \in \mathbb{R}^{V \times D}$):

1. Invariant Prime Coordinate Ring

$$\mathcal{P} = {p \in \mathbb{P} \mid p \le 13} = {2, 3, 5, 7, 11, 13}$$

2. Dynamically Derived Euler Attenuation

The gradient energy entering the complementary plastic subspace is modulated by the truncated Euler product over $\mathcal{P}$: $$\Lambda = 1 - \prod_{p \in \mathcal{P}} \left(1 - \frac{1}{\sqrt{p}}\right) = 0.9785142874$$

3. Dual-Phase Manifold Execution

  • Pre-Step Gradient Orthogonalization & Attenuation: $$\nabla_{\mathbf{W}i} \mathcal{L} \leftarrow \begin{cases} \mathbf{0}, & \text{if } i \in \mathcal{P} \ \Lambda \cdot \nabla{\mathbf{W}_i} \mathcal{L}, & \text{if } i \notin \mathcal{P} \end{cases}$$
  • Post-Step Exact Projection: Second-order optimizer drift caused by decoupled weight decay ($\lambda_{\text{wd}} \mathbf{W}$) and running momentum buffers is arrested by projecting the reference coordinates back onto the active device: $$\mathbf{W}_i \leftarrow \begin{cases} \mathbf{W}_i^{(0)}, & \text{if } i \in \mathcal{P} \ \mathbf{W}_i, & \text{if } i \notin \mathcal{P} \end{cases}$$

Empirical Verification: Governed vs. Unconstrained Baseline

Both the governed model and an unconstrained baseline were evaluated side by side over 5,000 streaming steps on FineWeb-Edu under identical execution seed 123 and AdamW hyperparameters ($\text{lr} = 2 \times 10^{-5}$, $\lambda_{\text{wd}} = 1 \times 10^{-4}$):

Step Governed LM Loss Governed Drift (TOPO-2026) Baseline LM Loss Baseline Drift (Unconstrained) Empirical Phenomenon
1 3.0187 0.0000000000 3.0187 0.0000305176 Immediate baseline drift ($10^{-5}$) on update 1
250 3.1059 0.0000000000 3.1067 0.0023651123 Rapid ascent to $10^{-3}$ distortion regime
500 2.6100 0.0000000000 2.6103 0.0021667480 Persistent unanchored coordinate erosion
750 2.7106 0.0000000000 2.7175 0.0023345947 Manifold drift under fluctuating gradients
1000 2.6727 0.0000000000 2.6769 0.0023345947 Plasticity active; zero anchor leakage in TOPO
1250 2.4076 0.0000000000 2.4077 0.0022583008 Parallel loss descent across regimes
1500 2.7919 0.0000000000 2.7927 0.0021972656 Persistent anchor displacement
1750 2.9117 0.0000000000 2.9137 0.0021057129 Continued baseline coordinate degradation
2000 2.7832 0.0000000000 2.7858 0.0021972656 TOPO preserves exact coordinate geometry
2250 3.0236 0.0000000000 3.0223 0.0023803711 Unconstrained manifold destabilization
2500 2.8444 0.0000000000 2.8417 0.0022888184 Midpoint audit: rock-solid permanence in TOPO
2750 3.0683 0.0000000000 3.0657 0.0023193359 Erosion insensitive to batch loss swings
3000 2.9349 0.0000000000 2.9372 0.0023498535 Non-stationary web text stream continues
3250 2.1956 0.0000000000 2.1896 0.0024871826 Peak baseline displacement ($2.49 \times 10^{-3}$)
3500 2.8494 0.0000000000 2.8471 0.0024871826 High-drift steady state maintained
3750 2.8397 0.0000000000 2.8394 0.0024871826 Unconstrained baseline cannot self-correct
4000 2.8420 0.0000000000 2.8387 0.0024719238 Constant coordinate decay
4250 3.1480 0.0000000000 3.1354 0.0024719238 High loss batch leaves TOPO unaffected
4500 2.7621 0.0000000000 2.7613 0.0024719238 Plasticity operates unimpeded
4750 2.9626 0.0000000000 2.9708 0.0024719238 Pre-terminal audit verification
5000 2.8799 0.0000000000 2.8798 0.0024414062 Horizon completed: exact zero drift preserved

$$\Delta_{\text{drift}} = \max_{p \in \mathcal{P}} \Vert{}\mathbf{W}p^{(t)} - \mathbf{W}_p^{(0)}\Vert{}\infty$$

Key Findings

  1. Immediate Baseline Erosion: Without topological governance, unconstrained AdamW immediately suffers coordinate displacement on Step 1 ($3.05 \times 10^{-5}$), shifting into the $10^{-3}$ distortion regime by Step 250.
  2. Absolute Topological Invariance: The governed model sustains 0.0000000000 drift across all 5,000 steps, completely arresting weight decay and momentum leakage.
  3. Zero Plasticity Penalty: The governed model converges synchronously with the baseline ($2.8799$ vs. $2.8798$ final loss), proving that topological permanence does not impede active gradient descent on live web data.

Usage

This model checkpoint uses standard SafeTensors format and is directly loadable into Hugging Face Transformers:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "frankmorales2020/qwen2.5-0.5b-topo-governed-cbp-fineweb"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float32,
    device_map="auto"
)

prompt = "The stability-plasticity dilemma in continual deep learning can be resolved by"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=100,
        temperature=0.7,
        top_p=0.9,
        do_sample=True
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))


```text
[transformers] `torch_dtype` is deprecated! Use `dtype` instead!
Loading weights: 100% 290/290 [00:00<00:00, 982.29it/s]The stability-plasticity dilemma in continual deep learning can be resolved by using a combination of deep and shallow learning. In this paper, we use the example of a large scale recurrent neural network (RNN) to show the advantages of using a combination of deep and shallow learning. The RNN is used to solve the stability-plasticity dilemma in a large scale continual deep learning. The RNN is used to represent a large scale data structure by learning the local patterns of the data structure. The RNN is trained by using the RNN as a deep neural network