Latent-belief RL on the passive BayesClue detective game. Base: Qwen3.5-4B. Belief is read from forced-choice probe logits (never verbalized). - Merge the LoRA with credalverl08/remergesftqwen35.py (the model.layers.→model.languagemodel.layers. remap); the stock verl.modelmerger writes a base copy. - Reward R = α·(−CE(pH,qH)) + (1−α)·means(−CE(pR,qR)), α=0.6. The −1.225 reward plateau IS the optimum −H(p), not a truncation artifact.
Open weights
other
Muhammad Faizan Khan. ChemEmbed positive-mode model and paired reference database from release v1.0.1. The computational-metabolomics maintainers record MIT for the model based on the upstream project licence, not a confirmed checkpoint-specific grant. Upstream CITATION.cff also describes a CC-BY-4.0 models/data deposit, but the exact released model's relationship to that deposit has not been established. The reference database is separately CC-BY-4.0; upstream documents its Parquet conversion from the cited Zenodo data. No conversion is performed by this mirror. Per-file licences below override this default; the model-card licence field does not relicense accompanying data. The software…
Open weights
mit
Action Chunking with Transformers (ACT) is an imitation-learning method that predicts short action chunks instead of single steps. It learns from teleoperated data and often achieves high success rates. This policy has been trained and pushed to the Hub using LeRobot. Learn how to train and run it in the LeRobot act guide, or browse the full documentation. The policy consumes these observation features and produces these action features. Inputs Outputs New to LeRobot? These guides cover the full workflow: - Install LeRobot — set up the lerobot package. - Hardware setup — assemble, wire, and calibrate your robot and cameras. - Record data & train a policy — the end-to-end imitation-learning…
Open weights
apache-2.0
52M parameters
lerobot
This repository contains the CI-Net processing, training, inference, and validation code. The directories under code follow the processing order: 1. datapreparing: read and align satellite and radar inputs. 2. labeling: create cloud labels. 3. finalpreprocess: convert the prepared fields into model input files. 4. training: train the model and run inference. 5. validation: create validation targets and calculate validation metrics. Each code directory contains Python files in src and its configuration and shell entry points in run. rawdata contains the input example included with this repository. result contains prepared data, model files, inference outputs, and validation outputs. See…
Open weights
cc-by-4.0
pytorch
Open weights
2.2B parameters
262,144 tokens
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights
Open weights