RL checkpoints of Qwen3.5-4B trained with ACT/WRITE co-evolving memory RL on SWE-Bench-CL 6p6 episode data (miles framework, GRPO, no cpu-offload). Weights exported from Megatron torchdist via language-tower conversion, with model.visual. / mtp. tensors (312) grafted verbatim from the base Qwen/Qwen3.5-4B checkpoint (training never touched them), stored in model-graft.safetensors. Mamba fp32-family tensors (Alog, linearattn norms, 48 keys) are upcast bf16→fp32 to match the base dtype contract — numerically identical to the trained values (Megatron stores the model state in bf16). Each export passed a 4-gate validation: full key/shape/dtype manifest parity vs base, language-tower difference…
Independent publisher
XuQixin
Racktic
NLP, mutimodel
Models in Library1
Datasets in Library0
Models on Hugging Face12
Followers3