SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm

by Namho Koh namhokaist/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm

qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm is an open-weight model for image and text to text from Namho Koh, released under Apache License 2.0. It has 770,288 parameters and a 262,144-token context. At 16-bit it needs about 0 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the…

Parameters770,288
Context262,144
Weights17.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm (770,288 parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm on every accelerator the SAVRN Index prices, at every precision

Model Card

By Namho Koh, published under apache-2.0, revision f10209c4660c.

Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its junction with the main stem. The model does not choose the target. Compared with the box+point models in this account (qwen3-vl-8b-tomato-cutpoint-bp-, which find the petiole themselves), this one measures localization given identity. Answer: {"cutpointuv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2…

Read Namho Koh's full model card

Qwen3-VL-8B tomato petiole cut point, BOX-CUED (target box given), RGB + depth, 9 mm labels — full fine-tune cued-full-8b-B-rgbd-lr1e-5-ep5

Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its junction with the main stem. The model does not choose the target. Compared with the box+point models in this account (qwen3-vl-8b-tomato-cutpoint-bp-*, which find the petiole themselves), this one measures localization given identity.

Answer: {"cut_point_uv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2 decimals; decode x*848/1000, y*408/1000, top-left edge origin, x right, y down). Greedy decoding (generation_config.json says do_sample; that is NOT how it was evaluated or served).

How to run

run_cutpoint_cued_standalone.py in this repo is self-contained (transformers 4.57.6, torch, pillow, numpy; any CUDA GPU with ~20 GB free):

python run_cutpoint_cued_standalone.py --model namhokaist/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm --rgb frame.png --depth frame_depth_m.npy --cue-box 681,276,848,338 [--valid valid.png] [--overlay out.png]
python run_cutpoint_cued_standalone.py --model namhokaist/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm --frames frames.jsonl --out predictions.jsonl   # keys rgb, depth, cue_box_xyxy, optional valid, id

It reproduces the served pipeline exactly: verified bit-for-bit against six archived endpoint requests (raw model text identical) and the cue-box text against the training encoder on 2,000 random boxes. The four weight shards are sha256-identical to the checkpoint that served our box-cued endpoint (provenance/endpoint_checkpoint_pin.json).

Inputs: RGB exactly 848x408 (do not crop or resize); depth float32 (408, 848) optical-axis Z in metres, pixel-aligned (optional valid mask, 255 = valid); the depth is rendered as turbo(inverse depth) over 0.1-1.5 m with invalid pixels black (code in the script). Cue box: [xmin, ymin, xmax, ymax] in pixel edges of the 848x408 frame, 0 <= xmin < xmax <= 848, 0 <= ymin < ymax <= 408; rounded to 0.1 px, then normalized to [0,1000) at 2 decimals in the prompt. Training cues were integer boxes of the target petiole's stalk render mask (xmin/ymin inclusive, xmax/ymax exclusive), reaching its junction with the main stem and excluding the leaflet blades; the prompt's phrase "together with its leaflets" does not describe the trained boxes. A box that includes leaflets or several structures, or cuts off the junction end, is outside the training distribution.

Prompt (exact): system = the cued system prompt in the script; user = RGB image, depth image, then The target petiole lies inside the box [x1, y1, x2, y2]. Trace it to its junction with the main stem and return the nominal cut point 9 mm along the petiole from that junction. Use the original full image. followed by the depth note.

Training data and recipe

Synthetic tomato-greenhouse release unified_release_v3_cued (manifest sha256 c7e6b2e565b086021be9a618a7a4eeacec58996f731e22b49a7303cf6335700e): 4,988 head-camera frames 848x408 with aligned noise-free rendered depth, labels from the renderer geometry (9 mm along the petiole centreline from the stem attachment; label epochs greenhouse.native848_all_petiole_9mm.v2 and continuations), one labelled target per frame (multi-candidate frames held out), splits by plant family: train 3,815 rows / 10 families / 1,405 targets, validation 609 / 5 / 240, test 564 / 4 / 280 (test unused here). Sources thor1 / thor3 / a794 (three simulator builds). The cue fed in training is the ground-truth stalk box of the labelled petiole.

Full-parameter fine-tune (vision tower at 0.1x LR), lr 1e-05, 5 epochs, per-device batch 4 x accumulation 2 x 4 GPUs = 32 examples/step, cosine schedule with 3 % warm-up, seed 41, +/-96 px shift augmentation (the cue box moves and clips with the content), bf16, DeepSpeed ZeRO-3 on 4x H200, 2402 s, final train loss 0.309. Depth input native_optical_z_inverse_turbo_0.1_1.5m.v1. W&B: https://wandb.ai/namhokoh-korea-advanced-institute-of-science-and-technology/tomato-pi/runs/d18lhtjj. Code: examples/greenhouse_sim/sim_data on branch koh-dev/sim-vlm of the project repository (train h200_train.py, eval h200_evaluate.py, scorer tomatopi-eval-v1; this upload at commit 9d4ff287).

Results (validation split, 609 rows from 5 unseen plant families, ground-truth box as the cue; invalid answers count as failures)

cut-point median px median mm p90 px <=5 px <=10 px <=20 px point inside the cue box valid answers
all 2.2 1.3 8.5 0.846 0.920 0.938 0.990 1.000
source a794 1.9 — 4.3 0.943 0.995 0.995 — —
source thor1 2.0 — 3.9 0.977 0.977 0.986 — —
source thor3 3.2 — 36.7 0.614 0.787 0.832 — —

Per-family cut median (px): seed17_full {'n': 158, 'median_px': 4.8, 'p75_px': 10.7, 'p90_px': 44.6, 'mean_px': 26.7, 'within_5px': 0.513, 'within_10px': 0.728, 'within_20px': 0.785, 'rows': 158}, seed29_full {'n': 29, 'median_px': 3.8, 'p75_px': 5.2, 'p90_px': 5.9, 'mean_px': 4.2, 'within_5px': 0.69, 'within_10px': 1.0, 'within_20px': 1.0, 'rows': 29}, seed53_full {'n': 155, 'median_px': 1.4, 'p75_px': 2.2, 'p90_px': 3.0, 'mean_px': 1.9, 'within_5px': 0.981, 'within_10px': 0.994, 'within_20px': 0.994, 'rows': 155}, seed7_full {'n': 217, 'median_px': 2.0, 'p75_px': 3.0, 'p90_px': 4.0, 'mean_px': 2.6, 'within_5px': 0.977, 'within_10px': 0.977, 'within_20px': 0.986, 'rows': 217}, seed97_full {'n': 50, 'median_px': 1.4, 'p75_px': 2.0, 'p90_px': 2.4, 'mean_px': 1.5, 'within_5px': 1.0, 'within_10px': 1.0, 'within_20px': 1.0, 'rows': 50}. Image-free position prior: median 328.5 px. Context: the untrained Qwen3-VL-8B base given the same cue returns a valid answer on 99.5 % of rows but lands within 10 px on 1.6 %; the box+point models that must find the petiole themselves reach 0.59-0.69 within 10 px on the same frames. Millimetres are depth-scaled image-plane distances at the label depth, not 3D errors.

Caveats

Simulated frames and noise-free rendered depth; labels are automatic (9 mm rule from the renderer geometry), flagged upstream training_approved: false. This checkpoint predates the project's move to the 3 mm anatomical cut convention and to metric 3D (camera-XYZ) answers; it answers in 2D pixels with the 9 mm rule. Perception only: never a claim that a blade motion is safe or executable.

Provenance

provenance/: run_contract.json (release validation, processor checks, hyperparameters, all train ids), completed.json, trainer_state.json, eval_validation_report.json (tomatopi-eval-v1), endpoint_checkpoint_pin.json (sha256 of the 18 served files). Weight shards: model-00001-of-00004.safetensors 29838f6c606a…, model-00002-of-00004.safetensors 47288937fdfd…, model-00003-of-00004.safetensors ac75fb3e2048…, model-00004-of-00004.safetensors bbfab9b99260….

Configuration

Architecture
Qwen3VLForConditionalGeneration
Context length (tokens)
262,144
Layers
36
Hidden size
4,096
Feed-forward size
12,288
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
151,936
RoPE base
5,000,000
Model type
qwen3_vl

Identity and Version

Repository
namhokaist/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm
Publisher
Namho Koh
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
770,288 parameters
Languages
Not stated by the source
Revision
f10209c4660cd69a7efc379a3cda2ecd2798ecf9
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

26 files, 17.6 GB in total. The weights are 5 files totalling 17.5 GB in bin, safetensors.

Weights5 files · 17.5 GB
Configuration14 files · 387.7 KB
Tokenizer4 files · 15.9 MB
Documentation1 file · 7.6 KB
Other1 file · 5.3 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00004.safetensorsWeights5.0 GB 29838f6c606a
model-00002-of-00004.safetensorsWeights4.9 GB 47288937fdfd
model-00003-of-00004.safetensorsWeights4.9 GB ac75fb3e2048
model-00004-of-00004.safetensorsWeights2.7 GB bbfab9b99260
training_args.binWeights7.2 KB 431b90508c1f
added_tokens.jsonConfiguration707 B —
config.jsonConfiguration1.5 KB —
generation_config.jsonConfiguration213 B —
grounding_adapter.jsonConfiguration450 B —
model.safetensors.index.jsonConfiguration67.8 KB —
preprocessor_config.jsonConfiguration782 B —
provenance/completed.jsonConfiguration680 B —
provenance/endpoint_checkpoint_pin.jsonConfiguration2.1 KB —
provenance/eval_validation_report.jsonConfiguration91.6 KB —
provenance/run_contract.jsonConfiguration195.2 KB —
provenance/trainer_state.jsonConfiguration12.2 KB —
run_cutpoint_cued_standalone.pyConfiguration13.0 KB —
special_tokens_map.jsonConfiguration613 B —
video_preprocessor_config.jsonConfiguration817 B —
README.mdDocumentation7.6 KB —
chat_template.jinjaOther5.3 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB aeb13307a71a
tokenizer_config.jsonTokenizer5.4 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
17.5 GB
Download from Namho Koh

Released by Namho Koh through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published17.5 GB
16-bit0.0 GB
8-bit0.0 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm

How much GPU memory does qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm need?

About 0 GB at 16-bit and 0 GB at 4-bit: the weights (770,288 parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm commercially?

Yes. qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

qwen3.5-35b-a3b-instruct-merged8083-sft

TR AKR

This model is a fine-tuned version of /mnt/shared-storage-user/mineru2-shared/niujunbo/ldy/models/Qwen3.5-35B-A3B-Instruct on the merged8083ulogocohqwen35 dataset. The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1.0 - Transformers 5.6.0 - Pytorch 2.10.0+cu128 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 664,944 parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…

Open weights apache-2.0 gguf

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0