SAVRN
Search Contact SAVRN

Open-weight model · Text generation

ProactiveInquirer-Qwen3-8B

by Ido Levy dolev31/ProactiveInquirer-Qwen3-8B

ProactiveInquirer-Qwen3-8B is an open-weight model for text generation from Ido Levy, released under Apache License 2.0. Its published files total 698.9 MB. It draws 31 downloads a month.

Ido Levy 1,2 · 1,2 This is the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents. It is a LoRA adapter on Qwen3-8B, trained with Q&D (questioner and drafter).

Parameters—
Context—
Weights698.5 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads31

Model Card

By Ido Levy, published under apache-2.0, revision 195506a9be74.

Ido Levy 1,2 · 1,2 This is the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents. It is a LoRA adapter on Qwen3-8B, trained with Q&D (questioner and drafter). A tool-using agent usually does what it is asked, yet a task often needs information the user never mentions. This model decides, one step at a time, which question to send to the agent's retriever next, or that it is time to stop. It goes after two kinds of unstated need: ZIP code, so the agent looks up the account. the order names the product, and the product lists the size-8 variant. These numbers are from the paper, on held-out test splits. The MuSiQue reading compares…

Read Ido Levy's full model card
# ProactiveInquirer-Qwen3-8B [Ido Levy](https://scholar.google.com/citations?user=Ok_7M80AAAAJ)1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM   2Weizmann Institute of Science [![Project page](https://img.shields.io/badge/Project-page-1B5EA8)](https://dolev31.github.io/ProactiveInquirer/) [![Paper](https://img.shields.io/badge/arXiv-2609.37236-b31b1b?logo=arxiv&logoColor=white)](https://arxiv.org/abs/2609.37236) [![Code](https://img.shields.io/badge/GitHub-ProactiveInquirer-181717?logo=github)](https://github.com/dolev31/ProactiveInquirer) [![License](https://img.shields.io/badge/license-Apache--2.0-blue)](https://www.apache.org/licenses/LICENSE-2.0)

This is the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents. It is a LoRA adapter on Qwen3-8B, trained with Q&D (questioner and drafter).

A tool-using agent usually does what it is asked, yet a task often needs information the user never mentions. This model decides, one step at a time, which question to send to the agent's retriever next, or that it is time to stop. It goes after two kinds of unstated need:

  • Horizontal proactivity: a need the current state already names. A customer gives a name and a ZIP code, so the agent looks up the account.
  • Vertical proactivity: a need that only newly found evidence names. The account lists the order, the order names the product, and the product lists the size-8 variant.

Results

These numbers are from the paper, on held-out test splits. The MuSiQue reading compares policies after the same number of questions, so asking more cannot pass for asking better. The τ²-bench readings are the benchmark's own task success under the same per-dialogue caps.

Setting Measure Qwen3-8B, prompted This model
MuSiQue, equal retrieval spend Required evidence recovered 78% 90%
τ²-bench retail (base prompt), no further training Task success 13% 34%
τ²-bench retail (stop prompt), no further training Task success 12% 32%
  • At equal retrieval spend it improves both forms of proactivity over the same model, prompted, on held-out splits of three multi-hop QA benchmarks, and outperforms GPT-OSS-120B, a prompted model 15× larger in the same role, on two of the three.
  • The gain comes from what it asks, not from asking more or longer questions: it holds against a question-volume control and a length control.
  • Placed in a customer-service agent with a simulated customer, with no further training, it raises retail task success from 13% to 34%, and it asks less and finds more: fewer questions, more of which reach the records the task needs. Against GPT-OSS-120B in retail, it completes more tasks with fewer follow-up turns from the customer.

How to use it

The questioner reads one prompt, the template it was trained on, and replies with one JSON action. The two template files are in prompts/.

import re

import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "dolev31/ProactiveInquirer-Qwen3-8B"
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, REPO)  # seed 1; add subfolder="seed2" for the second seed

template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()


def next_action(**state):
    fields = dict(state, user_channel=placebo.strip())
    prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
    ids = tok.apply_chat_template(
        [{"role": "user", "content": prompt}],
        add_generation_prompt=True,
        enable_thinking=False,
        return_tensors="pt",
        return_dict=True,
    ).to(model.device)
    out = model.generate(**ids, max_new_tokens=200, do_sample=False)
    return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)


question = "Who was the spouse of the director of the film The Great Flamarion?"
instructions = (
    "Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
    "queries against that pool before answering; several paragraphs are distractors, and "
    "the answer usually requires composing facts from more than one of them."
)
print(next_action(
    question=question, instructions=instructions, evidence="(nothing retrieved yet)",
    draft="(no draft yet)", history="(nothing asked yet)",
))
evidence = (
    "[3f2a9c1b7d4e] The Great Flamarion\n"
    "The Great Flamarion is a 1945 American film noir directed by Anthony Mann and starring "
    "Erich von Stroheim, Mary Beth Hughes and Dan Duryea."
)
print(next_action(
    question=question, instructions=instructions, evidence=evidence,
    draft="The film was directed by Anthony Mann; his spouse is not yet known.",
    history="Q1: Who directed the film The Great Flamarion?\n"
            "A1: The Great Flamarion (1945) was directed by Anthony Mann.",
))

Output, with greedy decoding (run on one A100 with transformers 5.17, peft 0.20 and torch 2.14):

{"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}
{"action": "ASK", "question": "Who was the spouse of film director Anthony Mann?", "rationale": "Need the spouse of the director to answer the task"}

The first question is horizontal: the task names the film, so the questioner goes after its director. The second is vertical: it uses a value only the retrieved evidence named, Anthony Mann. Compare the action field case-insensitively, since the model may write ASK or ask.

Inside an agent, feed each answer back: retrieved paragraphs go into evidence as [uid] title plus text, questions and answers into history as Q1:/A1: lines, and the drafter's text into draft. The ProactiveInquirer library runs the whole loop (questioner, retriever, drafter, answerer) and the paper's evaluation. The project page summarizes the paper and its results.

To serve it with vLLM, download the adapter and pass it as a LoRA module (rank 32):

huggingface-cli download dolev31/ProactiveInquirer-Qwen3-8B --local-dir proactive-inquirer
vllm serve Qwen/Qwen3-8B --enable-lora --max-lora-rank 32 \
    --lora-modules proactive-inquirer=./proactive-inquirer

Other formats

Training

  • Method. Q&D trains the questioner from the consequences of its own questions. A run is forked at one state and continued after several candidate questions (and after stopping). The candidate whose continuation retrieves more of the required evidence is preferred, and asking is preferred over stopping while evidence is still missing. The drafter is frozen, so every change in what the agent holds is caused by a question. No reward model or model judge is involved.
  • Stages. Three. First imitation: where the required evidence was already in hand the target is to stop, and otherwise the best sampled question, if its consequence score clears a fixed floor. Then direct preference optimization on question pairs, and last on question pairs and stop contrasts together, which rank asking above stopping at unfinished states. This adapter is the last stage.
  • Data. Training splits of MuSiQue, StrategyQA and 2WikiMultiHopQA. The final stage uses 31,473 preference pairs over 19,124 states (11,305 question-vs-question, 1,249 question-vs-stop and 18,919 synthetic question-vs-stop pairs). Nothing from τ²-bench is used in training.
  • Hyperparameters. LoRA r = 32, α = 64, dropout 0.05 on all attention and MLP projections. DPO with the sigmoid loss, β = 0.1, learning rate 5e-6, one epoch, gradient accumulation 16, bf16, sequences up to 5,120 tokens, Qwen3 chat template with thinking disabled.
  • Seeds. The paper reports two training seeds: seed 1 is at the root of this repository, seed 2 in seed2/.

Limitations

  • It has learned what to ask more readily than when to stop.
  • The extra evidence it finds does not yet translate into better final answers.
  • User-facing results come from a simulated customer, not from real people.
  • It is a component inside an agent, meant to be called with the template above. It is not a chat assistant, and it was trained and evaluated in English.

Citation

@article{levy2026asking,
  title   = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
  author  = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
  journal = {arXiv preprint arXiv:2609.37236},
  url     = {https://arxiv.org/abs/2609.37236},
  year    = {2026}
}

License

Apache-2.0, like the base model Qwen3-8B.

Identity and Version

Repository
dolev31/ProactiveInquirer-Qwen3-8B
Publisher
Ido Levy
Task
Text generation
Modality
Text
Library
peft
Parameters
Not stated by the source
Languages
en
Revision
195506a9be743498193435674dd1f78b37bcf135
First published
2026-09-27
Last updated
2026-09-30

Files and Weights

12 files, 698.9 MB in total. The weights are 2 files totalling 698.5 MB in safetensors.

Weights2 files · 698.5 MB
Configuration3 files · 4.6 KB
Documentation1 file · 10.5 KB
Other5 files · 444.3 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter_model.safetensorsWeights349.2 MB c33d3e6e0b7f
seed2/adapter_model.safetensorsWeights349.2 MB 8e8603830b15
adapter_config.jsonConfiguration1.1 KB —
example.pyConfiguration2.3 KB —
seed2/adapter_config.jsonConfiguration1.1 KB —
README.mdDocumentation10.5 KB —
assets/figure1.pngOther340.6 KB b276396d01ac
assets/results_frontier.pngOther52.4 KB —
assets/title-card.pngOther48.9 KB —
prompts/fragment_user_channel_placebo.txtOther257 B —
prompts/inquirer_prompted.txtOther2.1 KB —
.gitattributesRepository1.6 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
698.5 MB
Download from Ido Levy

Released by Ido Levy through its official repository on Hugging Face. Read the license.

Built From

  • Adapter of Qwen/Qwen3-8B
  • Derived from Qwen/Qwen3-8B
  • Described by arXiv:2609.37236
  • Trained on (disclosed) ChilleD/StrategyQA
  • Trained on (disclosed) dgslibisey/MuSiQue
  • Trained on (disclosed) xanhho/2WikiMultihopQA

Memory Requirements

PrecisionWeights in memory
As published698.5 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ProactiveInquirer-Qwen3-8B

Can I use ProactiveInquirer-Qwen3-8B commercially?

Yes. ProactiveInquirer-Qwen3-8B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers