Get the code and instructions, or read the write-up and watch the video. Two independently acting chefs share this policy and serve six soups in 512 ticks in PufferLib's Cramped Room. Each sees its public observation expressed in text, game rules, and eight recent action outcomes. The policy scores stay, up, down, left, right, and interact. Each selection executes one native action, without a planner, pathfinder, action macro, or assigned role. This repository contains LoRA adapters and the trained original NLI classifier head, not a standalone base model. The root is cooperative checkpoint 220. single-chef/ contains checkpoint 330, which initialized cooperative training. Both use OpenJev…
Open weights
mit
peft
This is an unchanged mirror of AlexWortega's pretrained OpenJev v5 0.8B model, published so adapters can identify and load this specific base without confusing it with the 2B and 4B models in the upstream repository. All model, tokenizer, and configuration files are byte-for-byte copies. I did not train this base model. The source is AlexWortega/openjev, qwen3.5-0.8b-nli-v5, pinned to revision 552759daad712f1af6c4c13dabcb1e047886fc9c. OpenJev turns the Qwen3.5-0.8B backbone into a three-class natural language inference classifier. Credit for the base model and its training belongs to AlexWortega and the Qwen team. See the upstream model card for the method and reported evaluations. The…
Open weights
mit
853M parameters
262,144 tokens
transformers