Qwen3-VL-8B tomato petiole cut point, BOX-CUED (target box given), RGB + depth, 9 mm labels — full fine-tune cued-full-8b-B-rgbd-lr1e-5-ep5
Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned
variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a
rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its
junction with the main stem. The model does not choose the target. Compared with the box+point models in this account
(qwen3-vl-8b-tomato-cutpoint-bp-*, which find the petiole themselves), this one measures localization given identity.
Answer: {"cut_point_uv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2 decimals; decode x*848/1000,
y*408/1000, top-left edge origin, x right, y down). Greedy decoding (generation_config.json says do_sample; that is NOT how it was
evaluated or served).
How to run
run_cutpoint_cued_standalone.py in this repo is self-contained (transformers 4.57.6, torch, pillow, numpy; any CUDA GPU with ~20 GB free):
python run_cutpoint_cued_standalone.py --model namhokaist/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm --rgb frame.png --depth frame_depth_m.npy --cue-box 681,276,848,338 [--valid valid.png] [--overlay out.png]
python run_cutpoint_cued_standalone.py --model namhokaist/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm --frames frames.jsonl --out predictions.jsonl # keys rgb, depth, cue_box_xyxy, optional valid, id
It reproduces the served pipeline exactly: verified bit-for-bit against six archived endpoint requests (raw model text identical) and the
cue-box text against the training encoder on 2,000 random boxes. The four weight shards are sha256-identical to the checkpoint that served
our box-cued endpoint (provenance/endpoint_checkpoint_pin.json).
Inputs: RGB exactly 848x408 (do not crop or resize); depth float32 (408, 848) optical-axis Z in metres, pixel-aligned (optional valid mask,
255 = valid); the depth is rendered as turbo(inverse depth) over 0.1-1.5 m with invalid pixels black (code in the script). Cue box:
[xmin, ymin, xmax, ymax] in pixel edges of the 848x408 frame, 0 <= xmin < xmax <= 848, 0 <= ymin < ymax <= 408; rounded to 0.1 px, then
normalized to [0,1000) at 2 decimals in the prompt. Training cues were integer boxes of the target petiole's stalk render mask
(xmin/ymin inclusive, xmax/ymax exclusive), reaching its junction with the main stem and excluding the leaflet blades; the prompt's phrase
"together with its leaflets" does not describe the trained boxes. A box that includes leaflets or several structures, or cuts off the junction
end, is outside the training distribution.
Prompt (exact): system = the cued system prompt in the script; user = RGB image, depth image, then
The target petiole lies inside the box [x1, y1, x2, y2]. Trace it to its junction with the main stem and return the nominal cut point 9 mm along the petiole from that junction. Use the original full image.
followed by the depth note.
Training data and recipe
Synthetic tomato-greenhouse release unified_release_v3_cued (manifest sha256 c7e6b2e565b086021be9a618a7a4eeacec58996f731e22b49a7303cf6335700e): 4,988 head-camera
frames 848x408 with aligned noise-free rendered depth, labels from the renderer geometry (9 mm along the petiole centreline from the stem
attachment; label epochs greenhouse.native848_all_petiole_9mm.v2 and continuations), one labelled target per frame (multi-candidate frames
held out), splits by plant family: train 3,815 rows / 10 families / 1,405 targets,
validation 609 / 5 / 240, test 564 / 4 / 280 (test unused here).
Sources thor1 / thor3 / a794 (three simulator builds). The cue fed in training is the ground-truth stalk box of the labelled petiole.
Full-parameter fine-tune (vision tower at 0.1x LR), lr 1e-05, 5 epochs, per-device batch 4 x accumulation
2 x 4 GPUs = 32 examples/step, cosine schedule with 3 % warm-up, seed 41, +/-96 px shift
augmentation (the cue box moves and clips with the content), bf16, DeepSpeed ZeRO-3 on 4x H200, 2402 s, final train loss
0.309. Depth input native_optical_z_inverse_turbo_0.1_1.5m.v1. W&B: https://wandb.ai/namhokoh-korea-advanced-institute-of-science-and-technology/tomato-pi/runs/d18lhtjj.
Code: examples/greenhouse_sim/sim_data on branch koh-dev/sim-vlm of the project repository (train h200_train.py, eval h200_evaluate.py,
scorer tomatopi-eval-v1; this upload at commit 9d4ff287).
Results (validation split, 609 rows from 5 unseen plant families, ground-truth box as the cue; invalid answers count as failures)
|
cut-point median px |
median mm |
p90 px |
<=5 px |
<=10 px |
<=20 px |
point inside the cue box |
valid answers |
| all |
2.2 |
1.3 |
8.5 |
0.846 |
0.920 |
0.938 |
0.990 |
1.000 |
| source a794 |
1.9 |
— |
4.3 |
0.943 |
0.995 |
0.995 |
— |
— |
| source thor1 |
2.0 |
— |
3.9 |
0.977 |
0.977 |
0.986 |
— |
— |
| source thor3 |
3.2 |
— |
36.7 |
0.614 |
0.787 |
0.832 |
— |
— |
Per-family cut median (px): seed17_full {'n': 158, 'median_px': 4.8, 'p75_px': 10.7, 'p90_px': 44.6, 'mean_px': 26.7, 'within_5px': 0.513, 'within_10px': 0.728, 'within_20px': 0.785, 'rows': 158}, seed29_full {'n': 29, 'median_px': 3.8, 'p75_px': 5.2, 'p90_px': 5.9, 'mean_px': 4.2, 'within_5px': 0.69, 'within_10px': 1.0, 'within_20px': 1.0, 'rows': 29}, seed53_full {'n': 155, 'median_px': 1.4, 'p75_px': 2.2, 'p90_px': 3.0, 'mean_px': 1.9, 'within_5px': 0.981, 'within_10px': 0.994, 'within_20px': 0.994, 'rows': 155}, seed7_full {'n': 217, 'median_px': 2.0, 'p75_px': 3.0, 'p90_px': 4.0, 'mean_px': 2.6, 'within_5px': 0.977, 'within_10px': 0.977, 'within_20px': 0.986, 'rows': 217}, seed97_full {'n': 50, 'median_px': 1.4, 'p75_px': 2.0, 'p90_px': 2.4, 'mean_px': 1.5, 'within_5px': 1.0, 'within_10px': 1.0, 'within_20px': 1.0, 'rows': 50}. Image-free position prior: median 328.5 px.
Context: the untrained Qwen3-VL-8B base given the same cue returns a valid answer on 99.5 % of rows but lands within 10 px on 1.6 %; the
box+point models that must find the petiole themselves reach 0.59-0.69 within 10 px on the same frames. Millimetres are depth-scaled image-plane
distances at the label depth, not 3D errors.
Caveats
Simulated frames and noise-free rendered depth; labels are automatic (9 mm rule from the renderer geometry), flagged upstream
training_approved: false. This checkpoint predates the project's move to the 3 mm anatomical cut convention and to metric 3D (camera-XYZ)
answers; it answers in 2D pixels with the 9 mm rule. Perception only: never a claim that a blade motion is safe or executable.
Provenance
provenance/: run_contract.json (release validation, processor checks, hyperparameters, all train ids), completed.json,
trainer_state.json, eval_validation_report.json (tomatopi-eval-v1), endpoint_checkpoint_pin.json (sha256 of the 18 served files).
Weight shards: model-00001-of-00004.safetensors 29838f6c606a…, model-00002-of-00004.safetensors 47288937fdfd…, model-00003-of-00004.safetensors ac75fb3e2048…, model-00004-of-00004.safetensors bbfab9b99260….