Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its junction with the main stem. The model does not choose the target. Compared with the box+point models in this account (qwen3-vl-8b-tomato-cutpoint-bp-, which find the petiole themselves), this one measures localization given identity. Answer: {"cutpointuv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2…
Independent publisher
Namho Koh
namhokaist
Models in Library1
Datasets in Library0
Models on Hugging Face204
Followers—