SAVRN
Search Contact SAVRN

Open-weight model · Text classification

decision-model-rl-overcooked

by Rodney Lafuente-Mercado rodneyslafuente/decision-model-rl-overcooked

decision-model-rl-overcooked is an open-weight model for text classification from Rodney Lafuente-Mercado, released under MIT License. Its published files total 132.2 MB.

Get the code and instructions, or read the write-up and watch the video. Two independently acting chefs share this policy and serve six soups in 512 ticks in PufferLib's Cramped Room.

Parameters—
Context—
Weights86.7 MB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By Rodney Lafuente-Mercado, published under mit, revision 6e3017c4b74f.

Get the code and instructions, or read the write-up and watch the video. Two independently acting chefs share this policy and serve six soups in 512 ticks in PufferLib's Cramped Room. Each sees its public observation expressed in text, game rules, and eight recent action outcomes. The policy scores stay, up, down, left, right, and interact. Each selection executes one native action, without a planner, pathfinder, action macro, or assigned role. This repository contains LoRA adapters and the trained original NLI classifier head, not a standalone base model. The root is cooperative checkpoint 220. single-chef/ contains checkpoint 330, which initialized cooperative training. Both use OpenJev…

Read Rodney Lafuente-Mercado's full model card

A decision model that learned to cook and cooperate

Get the code and instructions, or read the write-up and watch the video.

Two independently acting chefs share this policy and serve six soups in 512 ticks in PufferLib's Cramped Room. Each sees its public observation expressed in text, game rules, and eight recent action outcomes. The policy scores stay, up, down, left, right, and interact. Each selection executes one native action, without a planner, pathfinder, action macro, or assigned role.

This repository contains LoRA adapters and the trained original NLI classifier head, not a standalone base model. The root is cooperative checkpoint 220. single-chef/ contains checkpoint 330, which initialized cooperative training. Both use OpenJev v5's 0.8B checkpoint. The ancestry is Qwen3.5-0.8B, then OpenJev v5 0.8B, then this cooking adapter. The adapter points to an unchanged, attributed mirror of the exact upstream subfolder. A separate base repository avoids inheriting the unrelated 4B ancestry from OpenJev's shared repository.

Use

The code repository downloads the exact base model, builds the pinned native environment, evaluates the policy, and renders video with the original prompts and scores.

python download.py
python evaluate.py
python render.py runs/evaluation.json 220 --folder videos

For standard Transformers and PEFT loading:

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from peft import PeftModel

base_id = "rodneyslafuente/openjev-v5-0.8b"
revision = "61af02b317b1e462a7594721823c234974ec1a5e"
tokenizer = AutoTokenizer.from_pretrained(base_id, revision=revision)
base = AutoModelForSequenceClassification.from_pretrained(
    base_id, revision=revision, dtype=torch.bfloat16
).to("cuda")
base.config.get_text_config().pad_token_id = tokenizer.pad_token_id
model = PeftModel.from_pretrained(base, "rodneyslafuente/decision-model-rl-overcooked").eval()

The adapter metadata also supports automatic PEFT base resolution through the pinned 0.8B mirror. Use the repository's pinned dependencies. policy.yaml contains the prompt rules and training settings. The underlying NLI template comes from the base model config. For each proposed button, score the hypothesis Your next action should be {action}., take the entailment probability (class 1), and normalize across the six candidates. These are action-selection scores, not calibrated probabilities of success. The code uses optimized shared-prefix scoring for the reported evaluation.

Training and result

I trained text-backbone LoRA adapters with rank 16 and alpha 32 and fully trained the existing three-class NLI head. The adapter contains 10,825,728 trained parameters. No action-specific head was added. The algorithm is a clipped group-relative policy gradient with discounted reward-to-go, adaptive entropy, observation novelty, and a repeated-no-op penalty. The game provides rewards, including shaping for adding ingredients, starting cooking, and plating. There are no demonstrations or teacher model.

Training used one 16 GB RTX 5080. The successful single-chef stage took about 19.5 hours, followed by about 32.6 hours of cooperative training to checkpoint 220, excluding an earlier unsuccessful run. Cooperative updates collect 8,192 agent decisions across eight kitchens. Learning rates are 5e-6 for LoRA and 1e-5 for the head. The recorded checkpoint-220 update peaked at 7.45 GiB allocated VRAM.

The six-soup result is one greedy evaluation, seed 0, selected as the best checkpoint using that same seed. It is not a multi-seed average or held-out test. The two chefs share weights, but each chooses independently from its own observation. Cooperation appears in the trace as one chef plating or serving while the other adds onions. A trial on another layout did not transfer successfully. Retention on unrelated NLI tasks and visual control were not evaluated for these trained weights.

evaluation.json records all 512 ticks, both chefs' exact states, action scores, and outcomes. provenance.json records the base/environment revisions and the original adapter checksum. Training changed settings during development. The final configs support evaluation and continued training, but do not capture every historical intervention.

Attribution

Base: OpenJev, by AlexWortega, with a Qwen3.5-0.8B backbone. Environment: PufferLib. This release distributes the trained adapter under MIT. The base weights and dependencies retain their upstream licenses. The write-up includes the full method and references.

Identity and Version

Repository
rodneyslafuente/decision-model-rl-overcooked
Publisher
Rodney Lafuente-Mercado
Task
Text classification
Modality
Text
Library
peft
Parameters
Not stated by the source
Languages
en
Revision
6e3017c4b74feb9e57a794db477ffa56bb1ac081
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

17 files, 132.2 MB in total. The weights are 2 files totalling 86.7 MB in safetensors.

Weights2 files · 86.7 MB
Configuration6 files · 5.4 MB
Tokenizer4 files · 40.0 MB
Documentation2 files · 6.5 KB
Other2 files · 15.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter_model.safetensorsWeights43.4 MB 6fc8c69621bf
single-chef/adapter_model.safetensorsWeights43.4 MB 0730e43ad011
adapter_config.jsonConfiguration11.6 KB —
evaluation.jsonConfiguration5.4 MB —
policy.yamlConfiguration2.1 KB —
provenance.jsonConfiguration1.1 KB —
single-chef/adapter_config.jsonConfiguration11.6 KB —
single-chef/policy.yamlConfiguration1.8 KB —
LICENSEDocumentation1.1 KB —
README.mdDocumentation5.4 KB —
chat_template.jinjaOther7.8 KB —
single-chef/chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
single-chef/tokenizer.jsonTokenizer20.0 MB 06b9509352d2
single-chef/tokenizer_config.jsonTokenizer1.2 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
86.7 MB
Download from Rodney Lafuente-Mercado

Released by Rodney Lafuente-Mercado through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published86.7 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About decision-model-rl-overcooked

Can I use decision-model-rl-overcooked commercially?

Yes. decision-model-rl-overcooked is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Text classification

finbert

Prosus AI

FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…

Open weights 512 tokens transformers

Model · Text classification

twitter-roberta-base-sentiment-latest

Cardiff NLP

This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.

Open weights cc-by-4.0 514 tokens transformers

Model · Text classification

finbert-tone

Yi

FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…

Open weights 512 tokens transformers

Model · Text classification

ms-marco-MiniLM-L-6-v2

Joshua

https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights 512 tokens transformers.js

Model · Text classification

twitter-xlm-roberta-base-sentiment

Cardiff NLP

This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.

Open weights 514 tokens transformers

Model · Text classification

emotion-english-distilroberta-base

Hartmann

With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…

Open weights 514 tokens transformers