SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Qwen3-1.7B-GPTQ-Int4

by AXERA AXERA-TECH/Qwen3-1.7B-GPTQ-Int4

Qwen3-1.7B-GPTQ-Int4 is an open-weight model for text generation from AXERA, released under Apache License 2.0. Its published files total 2.0 GB. It draws 11 downloads a month.

This version of Qwen3-1.7B-GPTQ-Int4 has been converted to run on the Axera NPU using w4a16 quantization.

Parameters—
Context—
Weights622.3 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads11

Model Card

By AXERA, published under apache-2.0, revision 6bbc2ec7f7cf.

This version of Qwen3-1.7B-GPTQ-Int4 has been converted to run on the Axera NPU using w4a16 quantization. This model has been optimized with the following LoRA: For those who are interested in model conversion, you can try to export axmodel through the original repo: https://huggingface.co/Qwen/Qwen3-1.7B Convert the original Huggingface Qwen3-1.7B-GPTQ-Int4 to axmodel, and then apply the w4a16 quantization to get the final axmodel for axllm runtime. 方式二:一行命令安装(默认分支 axllm): 方式三:下载Github Actions CI 导出的可执行程序(适合没有编译环境的用户): https://github.com/AXERA-TECH/ax-llm/actions?query=branch%3Aaxllm 下载 最新 CI 导出的可执行程序(axllm),然后:

Read AXERA's full model card

This version of Qwen3-1.7B-GPTQ-Int4 has been converted to run on the Axera NPU using w4a16 quantization.

This model has been optimized with the following LoRA:

Compatible with Pulsar2 version: 5.2

Convert tools links:

For those who are interested in model conversion, you can try to export axmodel through the original repo : https://huggingface.co/Qwen/Qwen3-1.7B

Pulsar2 Link, How to Convert LLM from Huggingface to axmodel

AXera NPU LLM Runtime

Convert the original Huggingface Qwen3-1.7B-GPTQ-Int4 to axmodel, and then apply the w4a16 quantization to get the final axmodel for axllm runtime.

export FLOAT_MATMUL_USE_CONV_EU=1 # only support AX650, for better performance, please set this env var before running the conversion command.

# context window size 2048, prefill length 1024
pulsar2 llm_build --input_path Qwen3-1.7B-GPTQ-Int4 --output_path <your path> \
--hidden_state_type bf16 --kv_cache_len 2048 --prefill_len 128 --chip AX650 -c 1 --parallel 32 \
--last_kv_cache_len 128 --last_kv_cache_len 256 --last_kv_cache_len 384 --last_kv_cache_len 512 \
--last_kv_cache_len 640 --last_kv_cache_len 768 --last_kv_cache_len 896 --last_kv_cache_len 1024 -w s4

Support Platform

Chips w4a16 CMM Flash
AX650 12.72 tokens/sec 1.7 GiB 1.9GiB

How to use

安装 axllm

方式一:克隆仓库后执行安装脚本:

git clone -b axllm https://github.com/AXERA-TECH/ax-llm.git
cd ax-llm
./install.sh

方式二:一行命令安装(默认分支 axllm):

curl -fsSL https://raw.githubusercontent.com/AXERA-TECH/ax-llm/axllm/install.sh | bash

方式三:下载Github Actions CI 导出的可执行程序(适合没有编译环境的用户):

如果没有编译环境,请到: https://github.com/AXERA-TECH/ax-llm/actions?query=branch%3Aaxllm 下载 最新 CI 导出的可执行程序(axllm),然后:

chmod +x axllm
sudo mv axllm /usr/bin/axllm

模型下载(Hugging Face)

先创建模型目录并进入,然后下载到该目录:

mkdir -p AXERA-TECH/Qwen3-1.7B-GPTQ-Int4
cd AXERA-TECH/Qwen3-1.7B-GPTQ-Int4
hf download AXERA-TECH/Qwen3-1.7B-GPTQ-Int4 --local-dir .

# structure of the downloaded files
tree -L 3
.
└── AXERA-TECH
    └── Qwen3-1.7B-GPTQ-Int4
        ├── README.md
        ├── config.json
        ├── model.embed_tokens.weight.bfloat16.bin
        ├── post_config.json
        ├── qwen3_p128_l0_together.axmodel
...
        ├── qwen3_p128_l9_together.axmodel
        ├── qwen3_post.axmodel
        └── qwen3_tokenizer.txt

2 directories, 34 files

Inference with AX650 Host, such as M4N-Dock(爱芯派Pro) or AX650N DEMO Board

运行(CLI)

(base) root@ax650:~# axllm run AXERA-TECH/Qwen3-1.7B-GPTQ-Int4/
14:39:39.955 INF Init:890 | LLM init start
tokenizer_type = 1
96% | ############################## | 30 / 31 [3.85s<3.98s, 7.79 count/s] init post axmodel ok,remain_cmm(8230 MB)
14:39:43.809 INF Init:1045 | max_token_len: 2048
14:39:43.809 INF Init:1048 | kv_cache_size: 1024, kv_cache_num: 2048
14:39:43.809 INF init_groups_from_model:606 | prefill_token_num: 128
14:39:43.809 INF init_groups_from_model:820 | decode grp: 0, gid: 0, max_token_len: 2048
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 0, gid: 1, history_cap: 0, total_cap: 128, symbolic_cap: 1
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 1, gid: 2, history_cap: 128, total_cap: 256, symbolic_cap: 128
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 2, gid: 3, history_cap: 256, total_cap: 384, symbolic_cap: 256
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 3, gid: 4, history_cap: 384, total_cap: 512, symbolic_cap: 384
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 4, gid: 5, history_cap: 512, total_cap: 640, symbolic_cap: 512
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 5, gid: 6, history_cap: 640, total_cap: 768, symbolic_cap: 640
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 6, gid: 7, history_cap: 768, total_cap: 896, symbolic_cap: 768
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 7, gid: 8, history_cap: 896, total_cap: 1024, symbolic_cap: 896
14:39:43.809 INF init_groups_from_model:824 | prefill grp: 8, gid: 9, history_cap: 1024, total_cap: 1152, symbolic_cap: 1024
14:39:43.809 INF init_groups_from_model:831 | prefill_max_token_num: 1152
14:39:43.809 INF Init:27 | LLaMaEmbedSelector use mmap
100% | ################################ | 31 / 31 [3.85s<3.85s, 8.04 count/s] embed_selector init ok
14:39:43.810 INF load_config:282 | load config:
14:39:43.810 INF load_config:282 | {
14:39:43.810 INF load_config:282 | "enable_repetition_penalty": false,
14:39:43.810 INF load_config:282 | "enable_temperature": false,
14:39:43.810 INF load_config:282 | "enable_top_k_sampling": false,
14:39:43.810 INF load_config:282 | "enable_top_p_sampling": false,
14:39:43.810 INF load_config:282 | "penalty_window": 20,
14:39:43.810 INF load_config:282 | "repetition_penalty": 1.2,
14:39:43.810 INF load_config:282 | "temperature": 0.9,
14:39:43.810 INF load_config:282 | "top_k": 10,
14:39:43.810 INF load_config:282 | "top_p": 0.8
14:39:43.810 INF load_config:282 | }
14:39:43.810 INF Init:1139 | LLM init ok
Commands:
/q, /exit 退出
/reset 重置 kvcache
/dd 删除一轮对话
/pp 打印历史对话
Ctrl+C: 停止当前生成
----------------------------------------
prompt >> who are you
14:39:51.617 INF SetKVCache:1437 | decode_grpid:0 prefill_grpid:1 history_cap:0 total_cap:128 symbolic_cap:1 precompute_len:0 input_num_token:23 prefer_symbolic_group:0
14:39:51.617 INF SetKVCache:1458 | current prefill_max_token_num:1152
14:39:51.713 INF SetKVCache:1462 | first run
14:39:51.715 INF Run:1553 | input token num: 23, prefill_split_num: 1
14:39:51.715 INF Run:1640 | prefill chunk p=0 history_len=0 grpid=1 kv_cache_num=0 input_tokens=23
14:39:51.715 INF Run:1665 | prefill indices shape: p=0 idx_elems=128 idx_rows=1 pos_rows=0
14:39:51.866 INF Run:1837 | ttft: 151.12 ms
<think>
Okay, the user asked, "Who are you?" I need to respond appropriately. Let me think.

First, I should acknowledge their question and clarify my role. I'm an AI assistant, so I should mention that. But I need to keep it friendly and not too technical. Also, I should mention that I can help with various tasks, like answering questions, writing, or providing information. But I need to make sure not to mention any specific functions that might be too detailed. Also, I should avoid using markdown and keep the response natural.

Wait, the user might be testing if I'm a real person or an AI. I should clarify that I'm an AI assistant, not a human. But I should also mention that I can assist with various tasks. Need to keep it positive and helpful. Also, avoid any mention of specific functions that might be too detailed. Let me structure the response: greet them, state I'm an AI assistant, mention I can help with tasks, and offer to assist. Keep it concise and friendly.
</think>

I'm an AI assistant here to help you with questions, tasks, and more! I can answer your questions, write content, and assist with various tasks. Just let me know what you need!

14:40:12.236 NTC Run:2102 | hit eos,decode avg 12.57 token/s
14:40:12.236 INF GetKVCache:1408 | precompute_len:280, remaining:872
prompt >> /q

启动服务(OpenAI 兼容)

(base) root@ax650:~# axllm serve AXERA-TECH/Qwen3-1.7B-GPTQ-Int4/
14:41:17.619 INF Init:890 | LLM init start
tokenizer_type = 1
 96% | ##############################   |  30 /  31 [2.60s<2.69s, 11.54 count/s] init post axmodel ok,remain_cmm(8230 MB)
14:41:20.219 INF Init:1045 | max_token_len : 2048
14:41:20.219 INF Init:1048 | kv_cache_size : 1024, kv_cache_num: 2048
14:41:20.219 INF init_groups_from_model:606 | prefill_token_num : 128
14:41:20.219 INF init_groups_from_model:820 | decode grp: 0, gid: 0, max_token_len : 2048
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 0, gid: 1, history_cap: 0, total_cap: 128, symbolic_cap: 1
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 1, gid: 2, history_cap: 128, total_cap: 256, symbolic_cap: 128
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 2, gid: 3, history_cap: 256, total_cap: 384, symbolic_cap: 256
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 3, gid: 4, history_cap: 384, total_cap: 512, symbolic_cap: 384
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 4, gid: 5, history_cap: 512, total_cap: 640, symbolic_cap: 512
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 5, gid: 6, history_cap: 640, total_cap: 768, symbolic_cap: 640
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 6, gid: 7, history_cap: 768, total_cap: 896, symbolic_cap: 768
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 7, gid: 8, history_cap: 896, total_cap: 1024, symbolic_cap: 896
14:41:20.219 INF init_groups_from_model:824 | prefill grp: 8, gid: 9, history_cap: 1024, total_cap: 1152, symbolic_cap: 1024
14:41:20.219 INF init_groups_from_model:831 | prefill_max_token_num : 1152
14:41:20.219 INF Init:27 | LLaMaEmbedSelector use mmap
100% | ################################ |  31 /  31 [2.60s<2.60s, 11.92 count/s] embed_selector init ok
14:41:20.220 INF load_config:282 | load config:
14:41:20.220 INF load_config:282 | {
14:41:20.220 INF load_config:282 |     "enable_repetition_penalty": false,
14:41:20.220 INF load_config:282 |     "enable_temperature": false,
14:41:20.220 INF load_config:282 |     "enable_top_k_sampling": false,
14:41:20.220 INF load_config:282 |     "enable_top_p_sampling": false,
14:41:20.220 INF load_config:282 |     "penalty_window": 20,
14:41:20.220 INF load_config:282 |     "repetition_penalty": 1.2,
14:41:20.220 INF load_config:282 |     "temperature": 0.9,
14:41:20.220 INF load_config:282 |     "top_k": 10,
14:41:20.220 INF load_config:282 |     "top_p": 0.8
14:41:20.220 INF load_config:282 | }
14:41:20.220 INF Init:1139 | LLM init ok
Starting server on port 8000 with model 'AXERA-TECH/Qwen3-1.7B-GPTQ-Int4'...
API URLs:
  GET  http://127.0.0.1:8000/health
  GET  http://127.0.0.1:8000/v1/models
  POST http://127.0.0.1:8000/v1/chat/completions
  GET  http://10.126.29.54:8000/health
  GET  http://10.126.29.54:8000/v1/models
  POST http://10.126.29.54:8000/v1/chat/completions
  GET  http://172.17.0.1:8000/health
  GET  http://172.17.0.1:8000/v1/models
  POST http://172.17.0.1:8000/v1/chat/completions
Aliases:
  GET  http://127.0.0.1:8000/models
  POST http://127.0.0.1:8000/chat/completions
  GET  http://10.126.29.54:8000/models
  POST http://10.126.29.54:8000/chat/completions
  GET  http://172.17.0.1:8000/models
  POST http://172.17.0.1:8000/chat/completions
OpenAI API Server starting on http://0.0.0.0:8000
Max concurrency: 1
Models: AXERA-TECH/Qwen3-1.7B-GPTQ-Int4

OpenAI 调用示例

from openai import OpenAI

API_URL = "http://127.0.0.1:8000/v1"
MODEL = "AXERA-TECH/Qwen3-1.7B-GPTQ-Int4"

messages = [
    {"role": "system", "content": [{"type": "text", "text": "you are a helpful assistant."}]},
    {"role": "user", "content": "hello"},
]

client = OpenAI(api_key="not-needed", base_url=API_URL)
completion = client.chat.completions.create(
    model=MODEL,
    messages=messages,
)

print(completion.choices[0].message.content)

OpenAI 流式调用示例

from openai import OpenAI

API_URL = "http://127.0.0.1:8000/v1"
MODEL = "AXERA-TECH/Qwen3-1.7B-GPTQ-Int4"

messages = [
    {"role": "system", "content": [{"type": "text", "text": "you are a helpful assistant."}]},
    {"role": "user", "content": "hello"},
]

client = OpenAI(api_key="not-needed", base_url=API_URL)
stream = client.chat.completions.create(
    model=MODEL,
    messages=messages,
    stream=True,
)

print("assistant:")
for ev in stream:
    delta = getattr(ev.choices[0], "delta", None)
    if delta and getattr(delta, "content", None):
        print(delta.content, end="", flush=True)
print("")

Identity and Version

Repository
AXERA-TECH/Qwen3-1.7B-GPTQ-Int4
Publisher
AXERA
Task
Text generation
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
en
Revision
6bbc2ec7f7cff174d06058de4aed85112df6261b
First published
2025-11-21
Last updated
2026-09-19

Files and Weights

35 files, 2.0 GB in total. The weights are 1 file totalling 622.3 MB in bin.

Weights1 file · 622.3 MB
Configuration2 files · 874 B
Tokenizer1 file · 1.6 MB
Documentation1 file · 12.7 KB
Other29 files · 1.4 GB
Repository1 file · 6.5 KB
Every file
FileTypeSizeSHA-256
model.embed_tokens.weight.bfloat16.binWeights622.3 MB f56b71b6939a
config.jsonConfiguration595 B —
post_config.jsonConfiguration279 B —
README.mdDocumentation12.7 KB —
qwen3_p128_l0_together.axmodelOther36.8 MB 7518c50def8a
qwen3_p128_l10_together.axmodelOther36.8 MB 15b0e72451f1
qwen3_p128_l11_together.axmodelOther36.8 MB fb9dbf15a0b7
qwen3_p128_l12_together.axmodelOther36.8 MB 1d968c49674f
qwen3_p128_l13_together.axmodelOther36.8 MB 1e8f789d1494
qwen3_p128_l14_together.axmodelOther36.8 MB 16363f8f50f6
qwen3_p128_l15_together.axmodelOther36.8 MB 579d07897183
qwen3_p128_l16_together.axmodelOther36.8 MB 27d7ebe1a179
qwen3_p128_l17_together.axmodelOther36.8 MB c1c43670f89b
qwen3_p128_l18_together.axmodelOther36.8 MB b3c983524b19
qwen3_p128_l19_together.axmodelOther36.8 MB 6087f4aea6d6
qwen3_p128_l1_together.axmodelOther36.8 MB 513708376f95
qwen3_p128_l20_together.axmodelOther36.8 MB 424040275685
qwen3_p128_l21_together.axmodelOther36.8 MB d9cbf8ee53f0
qwen3_p128_l22_together.axmodelOther36.8 MB 07a7895a65fc
qwen3_p128_l23_together.axmodelOther36.8 MB 1115f2c826cc
qwen3_p128_l24_together.axmodelOther36.8 MB 7d0dcc1ca9ed
qwen3_p128_l25_together.axmodelOther36.8 MB f7fa17ac56cf
qwen3_p128_l26_together.axmodelOther36.8 MB f3fbcd1c0e55
qwen3_p128_l27_together.axmodelOther36.8 MB 11de33836040
qwen3_p128_l2_together.axmodelOther36.8 MB b361f16598c6
qwen3_p128_l3_together.axmodelOther36.8 MB e526ee37350a
qwen3_p128_l4_together.axmodelOther36.8 MB 186b2509f812
qwen3_p128_l5_together.axmodelOther36.8 MB 44291acf0d90
qwen3_p128_l6_together.axmodelOther36.8 MB 66c005a2f32b
qwen3_p128_l7_together.axmodelOther36.8 MB b285905280e0
qwen3_p128_l8_together.axmodelOther36.8 MB d0b5b0305221
qwen3_p128_l9_together.axmodelOther36.8 MB 932b4663504b
qwen3_post.axmodelOther339.8 MB 40ff9c646cb1
.gitattributesRepository6.5 KB —
qwen3_tokenizer.txtTokenizer1.6 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
622.3 MB
Download from AXERA

Released by AXERA through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published622.3 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Qwen3-1.7B-GPTQ-Int4

Can I use Qwen3-1.7B-GPTQ-Int4 commercially?

Yes. Qwen3-1.7B-GPTQ-Int4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers