Per-expert ARVQ v3 cold experts, NVFP4 hot experts, ARVQ-scored REAP allocation. 75/75 MoE layers replaced by the full-corpus sequential PV campaign. Other layers retain their previously published weights; see pvprogress.json. Each layer's two tensor files and reports are replaced together in one commit. Training draws sequentially from 18,001,846 text tokens at context 1024. Fixed validation and development-audit sets each contain 16,384 tokens. Adam trains FP4-constrained per-expert books and FP8-constrained per-block scales. The effective batch is 262,144 tokens, accumulated in four 65,536-token passes. Layers 4–26 use 69 updates: book/scale LR.048/.032 through update 45, then.012/.008.…
Independent publisher
Jarrel Seah
jarrelscy
Models
Complete uploaded 5% hot checkpoint: 1325 of 26496 routed experts use NVFP4. Remaining experts use per-expert ARVQ FP4 books and FP16 block scales. Global allocation ranks routing-weighted output-error benefit measured on 65536 training-only calibration tokens. All 69 layers have prepared weights. Sequential PV is continuing through the remaining layers. Layers 1-20 reuse the earlier accepted all-cold candidates (layer 9 retained its initial fit). Their cold weight tensor payloads were verified identical to those underlying the cached inputs for layer 21. Layers 21 and 22 have completed PV; layers 23-69 follow sequentially, freezing the selected NVFP4 hot experts. Each accepted layer…
ONNX conversion of nvidia/NV-Reason-CT for ONNX Runtime Web with WebGPU. These files are used by https://jarrelscy.github.io/nv-reason-ct-web/ (source: https://github.com/jarrelscy/nv-reason-ct-web), which runs the model entirely in the browser. Each.json manifest lists the ONNX graph, its external data chunks (split to fit browser buffer limits) and the decoder state names. The decoders use fp32 activations and MatMulNBits weights, and the lmhead is pruned to the last token. Token embeddings are tied to lmhead, so they are not stored separately; the web app dequantises embedding rows from the lmhead weights (described under embed in decoder.json). Image tokens use 3D MRoPE positions as in…