Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its junction with the main stem. The model does not choose the target. Compared with the box+point models in this account (qwen3-vl-8b-tomato-cutpoint-bp-, which find the petiole themselves), this one measures localization given identity. Answer: {"cutpointuv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2…
Open-weight model · Image and text to text
qwen3.5-35b-a3b-instruct-merged8083-sft
by TR AKR AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft
qwen3.5-35b-a3b-instruct-merged8083-sft is an open-weight model for image and text to text from TR AKR, released under other. It has 664,944 parameters and a 262,144-token context. At 16-bit it needs about 0 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
This model is a fine-tuned version of /mnt/shared-storage-user/mineru2-shared/niujunbo/ldy/models/Qwen3.5-35B-A3B-Instruct on the merged8083ulogocohqwen35 dataset.
Runs On
What it takes to serve qwen3.5-35b-a3b-instruct-merged8083-sft (664,944 parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.
Model Card
This model is a fine-tuned version of /mnt/shared-storage-user/mineru2-shared/niujunbo/ldy/models/Qwen3.5-35B-A3B-Instruct on the merged8083ulogocohqwen35 dataset. The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1.0 - Transformers 5.6.0 - Pytorch 2.10.0+cu128 - Datasets 4.0.0 - Tokenizers 0.22.2
Excerpt from the card by TR AKR, licensed other.
Configuration
- Architecture
- Qwen3_5MoeForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 40
- Hidden size
- 2,048
- Attention heads
- 16
- Key/value heads
- 2
- Head dimension
- 256
- Vocabulary size
- 248,320
- Experts
- 256
- Experts active per token
- 8
- Model type
- qwen3_5_moe
Identity and Version
- Repository
- AKRTR/qwen3.5-35b-a3b-instruct-merged8083-sft
- Publisher
- TR AKR
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 664,944 parameters
- Languages
- Not stated by the source
- Revision
- 5614bd26b324bfe51368efc19943f459bad23485
- First published
- 2026-10-06
- Last updated
- 2026-10-07
Files and Weights
31 files, 70.2 GB in total. The weights are 17 files totalling 70.2 GB in bin, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00016.safetensors | Weights | 4.3 GB | c75eca69e9e1 |
| model-00002-of-00016.safetensors | Weights | 4.5 GB | ca401db401f3 |
| model-00003-of-00016.safetensors | Weights | 5.0 GB | 422cffc6786e |
| model-00004-of-00016.safetensors | Weights | 4.0 GB | cb2881cdaaa9 |
| model-00005-of-00016.safetensors | Weights | 4.5 GB | 30001aff22b1 |
| model-00006-of-00016.safetensors | Weights | 5.0 GB | 1013ca4037cf |
| model-00007-of-00016.safetensors | Weights | 4.0 GB | 68b99f6a7cff |
| model-00008-of-00016.safetensors | Weights | 4.5 GB | e3cf60d948bf |
| model-00009-of-00016.safetensors | Weights | 5.0 GB | 8283573ef196 |
| model-00010-of-00016.safetensors | Weights | 4.0 GB | 653312085bcf |
| model-00011-of-00016.safetensors | Weights | 4.5 GB | 24d757010507 |
| model-00012-of-00016.safetensors | Weights | 5.0 GB | 2692e8fa6396 |
| model-00013-of-00016.safetensors | Weights | 4.0 GB | c9f637801064 |
| model-00014-of-00016.safetensors | Weights | 4.5 GB | 4451b8fbee1c |
| model-00015-of-00016.safetensors | Weights | 5.0 GB | 7ab5e36930b2 |
| model-00016-of-00016.safetensors | Weights | 2.6 GB | 97c7edaac726 |
| training_args.bin | Weights | 7.7 KB | f947eb568e90 |
| all_results.json | Configuration | 208 B | — |
| config.json | Configuration | 3.3 KB | — |
| generation_config.json | Configuration | 199 B | — |
| model.safetensors.index.json | Configuration | 96.9 KB | — |
| processor_config.json | Configuration | 1.2 KB | — |
| train_results.json | Configuration | 208 B | — |
| trainer_state.json | Configuration | 24.5 KB | — |
| README.md | Documentation | 1.6 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| trainer_log.jsonl | Other | 26.3 KB | — |
| training_loss.png | Other | 62.1 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 70.2 GB
Released by TR AKR through its official repository on Hugging Face.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 70.2 GB |
| 16-bit | 0.0 GB |
| 8-bit | 0.0 GB |
| 4-bit | 0.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About qwen3.5-35b-a3b-instruct-merged8083-sft
How much GPU memory does qwen3.5-35b-a3b-instruct-merged8083-sft need?
About 0 GB at 16-bit and 0 GB at 4-bit: the weights (664,944 parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run qwen3.5-35b-a3b-instruct-merged8083-sft on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What license is qwen3.5-35b-a3b-instruct-merged8083-sft released under?
other, as its publisher declares it. Read the license text before commercial use.
What is qwen3.5-35b-a3b-instruct-merged8083-sft's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…
Model · Image and text to text
Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…
This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…
Model · Image and text to text
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF
in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…
Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…