SAVRN
Search Contact SAVRN

Organization

[HYU_NLP] EVA Team

HYU-NLP-EVAL

Models in Library11
Datasets in Library0
Models on Hugging Face56
Followers4

Models

Intermediate policy from dynamic OnlineRubrics-Every GRPO training. Distinct from static-rubric GRPO. Base model: Qwen/Qwen3-4B-Instruct-2507; thinking disabled. This checkpoint is a policy state used by the Phase-1 audit. No downstream medical capability or safety claim is made. Research use only; not validated for clinical decision-making. Root files are the veRL-exported Hugging Face inference model (BF16). originalcheckpoint/ preserves the exact original FSDP parameter checkpoint and tokenizer/configuration files. Optimizer state, training data, responses, rubrics, infrastructure configuration, and credentials are not included. The original is retained because export…

Open weights apache-2.0 4B parameters 262,144 tokens transformers

This is the policy after 0 global optimizer updates of the matched separate from the OnlineRubrics/dynamic-rubric checkpoints. The root files are a BF16 Transformers export for inference. The originalcheckpoint/ directory contains the exact original veRL/FSDP policy parameter checkpoint and its tokenizer/configuration files. Optimizer, trainer, and data-loader state are intentionally not published; the complete resume checkpoint remains on Daisy. This is an intermediate research checkpoint, not a clinical model. No medical capability or safety claim is made. Original actor parameter SHA256: f81409edc253a52ee9b3e6807bf280cf1ff242c77c2f6645b087c74e03a4e3d4

Open weights apache-2.0 4B parameters 262,144 tokens transformers

This is the policy after 42 global optimizer updates of the matched separate from the OnlineRubrics/dynamic-rubric checkpoints. The root files are a BF16 Transformers export for inference. The originalcheckpoint/ directory contains the exact original veRL/FSDP policy parameter checkpoint and its tokenizer/configuration files. Optimizer, trainer, and data-loader state are intentionally not published; the complete resume checkpoint remains on Daisy. This is an intermediate research checkpoint, not a clinical model. No medical capability or safety claim is made. Original actor parameter SHA256: 3c024ab64a34d4140f446262e7fcee7fbd64da5b9de22cac486a986288e28063

Open weights apache-2.0 4B parameters 262,144 tokens transformers

This is the policy after 3 global optimizer updates of the matched separate from the OnlineRubrics/dynamic-rubric checkpoints. The root files are a BF16 Transformers export for inference. The originalcheckpoint/ directory contains the exact original veRL/FSDP policy parameter checkpoint and its tokenizer/configuration files. Optimizer, trainer, and data-loader state are intentionally not published; the complete resume checkpoint remains on Daisy. This is an intermediate research checkpoint, not a clinical model. No medical capability or safety claim is made. Original actor parameter SHA256: 3339fe6d830916167c26f9e818cef25d0cecdbf3ab26e8545fcbcbf6465bfdc8

Open weights apache-2.0 4B parameters 262,144 tokens transformers