# ProactiveInquirer-Qwen3-8B-GGUF
[Ido Levy](https://scholar.google.com/citations?user=Ok_7M80AAAAJ)1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM 2Weizmann Institute of Science
[](https://dolev31.github.io/ProactiveInquirer/)
[](https://arxiv.org/abs/2609.37236)
[](https://github.com/dolev31/ProactiveInquirer)
[](https://www.apache.org/licenses/LICENSE-2.0)
GGUF quantizations of the trained questioner from Asking for What Was Never Requested: Horizontal
and Vertical Proactivity in Agents, for llama.cpp, Ollama, LM Studio and Jan. They were made from the
merged model with llama.cpp
(commit 9adc7f4).
| File |
Quantization |
Size |
Notes |
ProactiveInquirer-Qwen3-8B-Q4_K_M.gguf |
Q4_K_M |
5.0 GB |
the usual choice, runs on a laptop |
ProactiveInquirer-Qwen3-8B-Q5_K_M.gguf |
Q5_K_M |
5.9 GB |
a step closer to the full model |
ProactiveInquirer-Qwen3-8B-Q8_0.gguf |
Q8_0 |
8.7 GB |
closest to the full model |
Before upload, each file ran the adapter card's
two-turn example with greedy decoding. Every file asked the same two questions as the full-precision model: "Who directed the film The Great Flamarion?" and, once the evidence named the director, "Who was the spouse of film director Anthony Mann?".
Results
The results are the trained questioner's, as the paper reports them: see the
adapter card's Results. The paper evaluated the unquantized model, not these files.
Run it
Ollama
ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
llama.cpp
llama-server -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M --jinja
LM Studio: search for ProactiveInquirer in the model browser.
The questioner reads the prompt template it was trained on, in the adapter repository's
prompts/, and replies
with one JSON action per step: {"action": "ask", "question": ...} or {"action": "stop", ...}. It
was trained with Qwen3's thinking off, so keep it off: in Ollama run it with --think=false (or send
"think": false to its API), and with llama.cpp's server send "chat_template_kwargs": {"enable_thinking": false}.
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"messages": [{"role": "user", "content": "<the filled template>"}],
"chat_template_kwargs": {"enable_thinking": false},
"temperature": 0
}'
Limitations
- The questioner's own limitations, from the paper: it has learned what to ask more readily than when to
stop, the extra evidence it finds does not yet translate into better final answers, and its user-facing
results come from a simulated customer, not from real people.
- It is a component inside an agent, meant to be called with its prompt template. It is not a chat
assistant, and it was trained and evaluated in English.
- Quantization can change the model's choices. Each file was checked on one example, as above: a
check, not an evaluation.
Citation
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint arXiv:2609.37236},
url = {https://arxiv.org/abs/2609.37236},
year = {2026}
}
License
Apache-2.0, like the base model Qwen3-8B.