# ProactiveInquirer-Qwen3-8B
[Ido Levy](https://scholar.google.com/citations?user=Ok_7M80AAAAJ)1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM 2Weizmann Institute of Science
[](https://dolev31.github.io/ProactiveInquirer/)
[](https://arxiv.org/abs/2609.37236)
[](https://github.com/dolev31/ProactiveInquirer)
[](https://www.apache.org/licenses/LICENSE-2.0)
This is the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical
Proactivity in Agents. It is a LoRA adapter on Qwen3-8B, trained
with Q&D (questioner and drafter).
A tool-using agent usually does what it is asked, yet a task often needs information the user never
mentions. This model decides, one step at a time, which question to send to the agent's retriever next,
or that it is time to stop. It goes after two kinds of unstated need:
- Horizontal proactivity: a need the current state already names. A customer gives a name and a
ZIP code, so the agent looks up the account.
- Vertical proactivity: a need that only newly found evidence names. The account lists the order,
the order names the product, and the product lists the size-8 variant.
Results
These numbers are from the paper, on held-out test splits. The MuSiQue reading compares policies after
the same number of questions, so asking more cannot pass for asking better. The τ²-bench readings are
the benchmark's own task success under the same per-dialogue caps.
| Setting |
Measure |
Qwen3-8B, prompted |
This model |
| MuSiQue, equal retrieval spend |
Required evidence recovered |
78% |
90% |
| τ²-bench retail (base prompt), no further training |
Task success |
13% |
34% |
| τ²-bench retail (stop prompt), no further training |
Task success |
12% |
32% |
- At equal retrieval spend it improves both forms of proactivity over the same model, prompted, on
held-out splits of three multi-hop QA benchmarks, and outperforms GPT-OSS-120B, a prompted model 15×
larger in the same role, on two of the three.
- The gain comes from what it asks, not from asking more or longer questions: it holds against a
question-volume control and a length control.
- Placed in a customer-service agent with a simulated customer, with no further training, it raises
retail task success from 13% to 34%, and it asks less and finds more: fewer questions, more of which
reach the records the task needs. Against GPT-OSS-120B in retail, it completes more tasks with fewer
follow-up turns from the customer.
How to use it
The questioner reads one prompt, the template it was trained on, and replies with one JSON action. The
two template files are in prompts/.
import re
import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "dolev31/ProactiveInquirer-Qwen3-8B"
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, REPO) # seed 1; add subfolder="seed2" for the second seed
template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()
def next_action(**state):
fields = dict(state, user_channel=placebo.strip())
prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
ids = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
).to(model.device)
out = model.generate(**ids, max_new_tokens=200, do_sample=False)
return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)
question = "Who was the spouse of the director of the film The Great Flamarion?"
instructions = (
"Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
"queries against that pool before answering; several paragraphs are distractors, and "
"the answer usually requires composing facts from more than one of them."
)
print(next_action(
question=question, instructions=instructions, evidence="(nothing retrieved yet)",
draft="(no draft yet)", history="(nothing asked yet)",
))
evidence = (
"[3f2a9c1b7d4e] The Great Flamarion\n"
"The Great Flamarion is a 1945 American film noir directed by Anthony Mann and starring "
"Erich von Stroheim, Mary Beth Hughes and Dan Duryea."
)
print(next_action(
question=question, instructions=instructions, evidence=evidence,
draft="The film was directed by Anthony Mann; his spouse is not yet known.",
history="Q1: Who directed the film The Great Flamarion?\n"
"A1: The Great Flamarion (1945) was directed by Anthony Mann.",
))
Output, with greedy decoding (run on one A100 with transformers 5.17, peft 0.20 and torch 2.14):
{"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}
{"action": "ASK", "question": "Who was the spouse of film director Anthony Mann?", "rationale": "Need the spouse of the director to answer the task"}
The first question is horizontal: the task names the film, so the questioner goes after its director.
The second is vertical: it uses a value only the retrieved evidence named, Anthony Mann. Compare the
action field case-insensitively, since the model may write ASK or ask.
Inside an agent, feed each answer back: retrieved paragraphs go into evidence as [uid] title plus
text, questions and answers into history as Q1:/A1: lines, and the drafter's text into draft.
The ProactiveInquirer library runs the whole loop
(questioner, retriever, drafter, answerer) and the paper's evaluation. The
project page summarizes the paper and its results.
To serve it with vLLM, download the adapter and pass it as a LoRA module (rank 32):
huggingface-cli download dolev31/ProactiveInquirer-Qwen3-8B --local-dir proactive-inquirer
vllm serve Qwen/Qwen3-8B --enable-lora --max-lora-rank 32 \
--lora-modules proactive-inquirer=./proactive-inquirer
Other formats
Training
- Method. Q&D trains the questioner from the consequences of its own questions. A run is forked
at one state and continued after several candidate questions (and after stopping). The candidate
whose continuation retrieves more of the required evidence is preferred, and asking is preferred
over stopping while evidence is still missing. The drafter is frozen, so every change in what the
agent holds is caused by a question. No reward model or model judge is involved.
- Stages. Three. First imitation: where the required evidence was already in hand the target is to
stop, and otherwise the best sampled question, if its consequence score clears a fixed floor. Then
direct preference optimization on question pairs, and last on question pairs and stop contrasts
together, which rank asking above stopping at unfinished states. This adapter is the last stage.
- Data. Training splits of MuSiQue, StrategyQA and 2WikiMultiHopQA. The final stage uses 31,473
preference pairs over 19,124 states (11,305 question-vs-question, 1,249 question-vs-stop and 18,919
synthetic question-vs-stop pairs). Nothing from τ²-bench is used in training.
- Hyperparameters. LoRA r = 32, α = 64, dropout 0.05 on all attention and MLP projections. DPO with
the sigmoid loss, β = 0.1, learning rate 5e-6, one epoch, gradient accumulation 16, bf16, sequences up to
5,120 tokens, Qwen3 chat template with thinking disabled.
- Seeds. The paper reports two training seeds: seed 1 is at the root of this repository, seed 2 in
seed2/.
Limitations
- It has learned what to ask more readily than when to stop.
- The extra evidence it finds does not yet translate into better final answers.
- User-facing results come from a simulated customer, not from real people.
- It is a component inside an agent, meant to be called with the template above. It is not a chat
assistant, and it was trained and evaluated in English.
Citation
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint arXiv:2609.37236},
url = {https://arxiv.org/abs/2609.37236},
year = {2026}
}
License
Apache-2.0, like the base model Qwen3-8B.