SAVRN
Search Contact SAVRN

Open-weight model · Text generation

DietRecommendation-Qwen2.5-0.5B

by Yubraj Sigdel syubraj/DietRecommendation-Qwen2.5-0.5B

DietRecommendation-Qwen2.5-0.5B is an open-weight model for text generation from Yubraj Sigdel, released under Apache License 2.0. Its published files total 24.7 MB.

A LoRA adapter for Qwen/Qwen2.5-0.5B-Instruct that turns a short user profile (age, gender, height, weight, activity level, dietary preference, daily calorie target) into a one-day meal plan with breakfast, lunch, snack and dinner.

Parameters—
Context—
Weights8.8 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By Yubraj Sigdel, published under apache-2.0, revision 8dce44a03fb9.

A LoRA adapter for Qwen/Qwen2.5-0.5B-Instruct that turns a short user profile (age, gender, height, weight, activity level, dietary preference, daily calorie target) into a one-day meal plan with breakfast, lunch, snack and dinner. It is a small, fast, single-purpose model: about 2.2M trainable parameters on top of a 0.5B base, usable on CPU. The adapter was trained on one fixed prompt layout. Use the same system prompt and the same field names and order, otherwise quality drops. Two details matter: - Pass the system prompt explicitly. Without it, the chat template inserts Qwen's default system prompt, which the adapter never saw. - The first field is spelled Ages: in the training data, so…

Read Yubraj Sigdel's full model card

A LoRA adapter for Qwen/Qwen2.5-0.5B-Instruct that turns a short user profile (age, gender, height, weight, activity level, dietary preference, daily calorie target) into a one-day meal plan with breakfast, lunch, snack and dinner.

It is a small, fast, single-purpose model: about 2.2M trainable parameters on top of a 0.5B base, usable on CPU.

Not medical advice. This is an experimental model trained on a small dataset. Its output has not been reviewed by a dietitian and must not be used to manage a health condition. See Limitations.

Model details

Developed by syubraj
Model type LoRA adapter (PEFT) for a causal language model
Base model Qwen/Qwen2.5-0.5B-Instruct
Language English
License Apache-2.0
Training data syubraj/DietRecommendation-dataset-Qwen-2.5-0.5b
Related model syubraj/DietRecommender_4bit_Qwen2.5-0.5B (trained on the same dataset)

How to use

The adapter was trained on one fixed prompt layout. Use the same system prompt and the same field names and order, otherwise quality drops. Two details matter:

  • Pass the system prompt explicitly. Without it, the chat template inserts Qwen's default system prompt, which the adapter never saw.
  • The first field is spelled Ages: in the training data, so keep it that way.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

BASE = "Qwen/Qwen2.5-0.5B-Instruct"
ADAPTER = "syubraj/DietRecommendation-Qwen2.5-0.5B"

tokenizer = AutoTokenizer.from_pretrained(ADAPTER)
model = AutoModelForCausalLM.from_pretrained(BASE, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval()

# Exact system prompt used in training (including the line break).
SYSTEM_PROMPT = (
    "Act as a nutrition expert. Based on the user’s age, gender, height, weight, "
    "activity level, diet preference, and calorie target, \n"
    "suggest a balanced meal plan with breakfast, lunch, snacks, and dinner."
)


def build_profile(age, gender, height_cm, weight_kg, activity, diet, kcal):
    return (
        f"Ages: {age}\n"
        f"Gender: {gender}\n"
        f"Height: {height_cm} cm\n"
        f"Weight: {weight_kg} kg\n"
        f"Activity Level: {activity}\n"
        f"Dietary Preference: {diet}\n"
        f"Daily Calorie Target: {kcal} kcal"
    )


messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": build_profile(32, "Female", 165, 65, "Lightly Active", "Vegetarian", 1600)},
]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)

with torch.no_grad():
    output = model.generate(**inputs, max_new_tokens=96, do_sample=False)

print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

device_map="auto" needs accelerate; drop that argument to load on CPU.

Input values seen in training

Field Values
Gender Male, Female
Activity Level Sedentary, Lightly Active, Moderately Active, Very Active
Dietary Preference Omnivore, Vegetarian, Vegan
Height / Weight centimetres / kilograms
Daily Calorie Target kcal

Output format

Four lines, one per meal. This example is a row from the training set and shows the format the model was taught to produce:

Breakfast: Tofu scramble with veggies
Lunch: Lentil soup with whole wheat bread
Snack: Apple with almond butter
Dinner: Vegetable stir-fry with brown rice

Merging the adapter

To ship a standalone model with no PEFT dependency:

merged = model.merge_and_unload()
merged.save_pretrained("DietRecommendation-Qwen2.5-0.5B-merged")
tokenizer.save_pretrained("DietRecommendation-Qwen2.5-0.5B-merged")

Intended uses

  • Prototypes and demos of profile-to-meal-plan generation.
  • A starting point for meal-idea features where a person reviews the result.
  • A worked example of LoRA fine-tuning a very small instruct model on a structured task.

Out of scope

  • Medical nutrition therapy, or diet planning for any health condition (diabetes, kidney disease, food allergies, pregnancy, eating disorders and so on). The prompt has no field for conditions, allergies or medication.
  • Children and teenagers. The training examples sampled for this card were all adults.
  • Any setting where the output is delivered to people as professional advice without qualified review.
  • General chat. The adapter is tuned for one prompt format.

Training data

syubraj/DietRecommendation-dataset-Qwen-2.5-0.5b: 1,698 examples (1,358 train / 340 validation), derived from the Kaggle dataset Nutrition Daily Meals in Diseases Cases.

Each example is a single text string already rendered in ChatML (system prompt, user profile, assistant meal plan), 525 to 678 characters long.

Training procedure

LoRA configuration

Setting Value
Rank (r) 4
lora_alpha 16
lora_dropout 0.1
Bias none
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Task type CAUSAL_LM
Trainable parameters about 2.2M (roughly 0.44% of the base model)

Hyperparameters

Setting Value
Learning rate 2e-4
Train / eval batch size 4 / 4
Gradient accumulation steps 4 (effective batch size 16)
Optimizer AdamW (adamw_torch), betas (0.9, 0.999), epsilon 1e-8
LR scheduler cosine, 100 warmup steps
Epochs 15
Mixed precision native AMP
Seed 42

Training results

Training loss Epoch Step Validation loss
3.0693 1.18 100 0.2517
0.9641 2.35 200 0.2243
0.8782 3.53 300 0.2218
0.8378 4.71 400 0.2253
0.8114 5.88 500 0.2179
0.7791 7.06 600 0.2193
0.7539 8.24 700 0.2178
0.7247 9.41 800 0.2185
0.6962 10.59 900 0.2234
0.6731 11.76 1000 0.2265
0.6363 12.94 1100 0.2317
0.6184 14.12 1200 0.2343

Final reported validation loss: 0.2343 (step 1200).

How to read these numbers:

  • Validation loss bottoms out at step 700 (0.2178) and then climbs slowly while training loss keeps falling. That is mild overfitting; around 8 epochs would have been enough.
  • Training and validation loss are on different scales in this log (training is several times higher throughout), so compare each column only with itself.
  • The dataset stores each example as one full ChatML string, and every example shares the same system prompt. A loss computed over the whole string is pulled down by that repeated text, so the absolute value says little about meal-plan quality.

Framework versions

  • PEFT 0.14.0
  • Transformers 4.47.0
  • PyTorch 2.5.1+cu121
  • Datasets 3.3.1
  • Tokenizers 0.21.0

Evaluation

Only validation loss was measured. There is no task-level evaluation yet: no check of nutritional adequacy, calorie accuracy, adherence to the dietary preference, or expert review. Treat the model as unvalidated for all of these.

Limitations

  • No quantities. Outputs name dishes only, with no portion sizes, calories or macronutrients. Nothing ensures the plan meets the requested calorie target.
  • Weak personalisation. In the training data, the same meal plan appears for different profiles and calorie targets, and identical profiles appear with different plans. Expect the dietary preference to shape the output much more than age, height, weight or calorie target.
  • Narrow menu. The training plans draw on a limited, largely Western set of dishes (oatmeal, tofu scramble, grilled chicken salad, lentil soup, salmon with vegetables). Outputs will be repetitive and may not suit other cuisines or budgets.
  • Three diet types only. Omnivore, vegetarian and vegan. No support for allergies, intolerances, religious diets, keto, gluten-free and so on.
  • Small base model, small dataset. A 0.5B model trained on about 1.4k examples can still produce a non-vegan dish in a vegan plan or other inconsistencies. Validate outputs in code if that matters for your use.
  • Possibly optimistic validation loss. The training split contains repeated and near-identical examples; if the validation split overlaps with them, the validation loss understates the real error.
  • Fixed format, English only. Behaviour on free-form requests, other languages or other field layouts is untested.

Recommendations

Show a clear "not medical advice" notice wherever outputs reach end users, keep a person in the loop, and check dietary-preference compliance separately rather than trusting the model.

Citation

@misc{syubraj2025dietrecommendation,
  title  = {DietRecommendation-Qwen2.5-0.5B: a LoRA adapter for meal-plan recommendation},
  author = {syubraj},
  year   = {2025},
  url    = {https://huggingface.co/syubraj/DietRecommendation-Qwen2.5-0.5B}
}

Base model:

@misc{qwen2.5,
  title  = {Qwen2.5: A Party of Foundation Models},
  url    = {https://qwenlm.github.io/blog/qwen2.5/},
  author = {Qwen Team},
  month  = {September},
  year   = {2024}
}

Identity and Version

Repository
syubraj/DietRecommendation-Qwen2.5-0.5B
Publisher
Yubraj Sigdel
Task
Text generation
Modality
Text
Library
peft
Parameters
Not stated by the source
Languages
en
Revision
8dce44a03fb96a790bb223fbe42e1c9739b561d7
First published
2025-03-09
Last updated
2026-10-05

Files and Weights

12 files, 24.7 MB in total. The weights are 2 files totalling 8.8 MB in bin, safetensors.

Weights2 files · 8.8 MB
Configuration4 files · 2.2 KB
Tokenizer4 files · 15.9 MB
Documentation1 file · 10.1 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter_model.safetensorsWeights8.8 MB 18d2cb182d85
training_args.binWeights5.4 KB d7ac75510ce4
adapter_config.jsonConfiguration798 B —
added_tokens.jsonConfiguration605 B —
generation_config.jsonConfiguration156 B —
special_tokens_map.jsonConfiguration613 B —
README.mdDocumentation10.1 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB 540b7fbf60b8
tokenizer_config.jsonTokenizer7.3 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
8.8 MB
Download from Yubraj Sigdel

Released by Yubraj Sigdel through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published8.8 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DietRecommendation-Qwen2.5-0.5B

Can I use DietRecommendation-Qwen2.5-0.5B commercially?

Yes. DietRecommendation-Qwen2.5-0.5B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers