SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Inkling-Small-GGUF

by Unsloth AI unsloth/Inkling-Small-GGUF

Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages.

Parameters
Context
Weights3.9 TB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.4M

Model Card

By Unsloth AI, published under apache-2.0, revision 1a19ef82883c.

Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers. Languages: English, with…

Read Unsloth AI's full model card

Read our How to Run Inkling Guide!

See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks.

  • You can now run Inkling in Unsloth Studio with toggles for Thinking.
  • Read our Inkling guide for analysis and instructions.
  • See below for example of 1-bit UD-IQ1_S GGUF running in Unsloth:

Inkling

BF16 | NVFP4 | Playground | Tinker Cookbook | Acceptable Use

1. General Information

Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

Languages: English, with general multilingual capabilities across other languages.

2. Getting Started

Try Inkling-Small on the Tinker Playground or access via API using the Tinker Cookbook.

Inkling-Small supports local deployment using the following open-source libraries:

API access is also available through third party inference providers.

3. Model Properties

Model type

Multimodal autoregressive transformer

Architecture type

A 42-layer decoder-only transformer with a sparse Mixture-of-Experts (MoE) feed-forward backbone: each token is routed to 6 of 256 experts, plus 2 shared experts active on every token. Attention is a hybrid of local and global layers. The model is natively multimodal — images are encoded via a hierarchical patch encoder, and audio via discrete token encoding — with all modalities projected into a shared hidden space and processed jointly by the decoder.

Parameters

276B total, 12B active

Numerics support

BF16 and NVFP4

Input modalities

Inkling-Small accepts text, image, and audio inputs:

  • Text: UTF-8 encoded text
  • Image: Any pixel-based image input. For optimal performance, each image dimension should be between 40px to 4096px.
  • Audio: WAV format, sampled at 16kHz. For optimal performance, audio length should ideally be under 2 mins.

Output modalities

Inkling-Small generates output as UTF-8 encoded text.

4. Training

Training data includes a broad variety of content types, including text, images, audio, video.

Training data for the model was drawn from publicly available sources, acquired from third-parties, or synthetically generated or augmented. Publicly available data includes content from the public internet and publicly accessible repositories.

The training data curation process includes cleaning, processing, and modifying datasets. These processing steps, which vary by data type, may include deduplication and filtering to remove junk or other low-quality data, or to advance safety or other objectives.

5. Evaluations

Open weightsClosed weights
Inkling-Small
Qwen3.5-397B-A17B
MiMo V2.5
Minimax M2.7
DeepSeek V4 Flash
Nemotron 3 Ultra
Inkling
Claude 4.5 Haiku
Gemini 3.5 Flash-Lite
GPT 5.6 Luna
Model Info
AA Index (v4.1)Score40.0%34.0%37.0%38.0%40.0%38.0%41.0%30.0%36.0%49.0%
Activated Params (B)Score12171510135541
Total Params (B)Score276397310230284550975
Pricing ($/M)Input0.30.390.140.250.140.5110.30.5
Pricing ($/M)Output1.22.340.2810.282.24.0552.53
Agentic (coding)
SWEBench VerifiedScore80.2%76.4%71.0%79.9%79.0%70.7%77.6%66.6%75.0%93.0%
SWEBench Pro (Public)Score55.9%50.9%56.1%56.2%52.6%46.4%54.3%39.5%54.2%62.7%
Terminal Bench 2.1Best Harness64.6951.363.755.461.856.463.844.25482.5
SciCodeScore48.7%42.0%43.1%47.0%44.9%39.9%46.1%43.3%40.9%50.0%
Agentic (general)
GDPVal-AA v2Score12699621145115911891164123891111391530
MCP AtlasPublic79.674.249.46947.478.841.279.877
MCP AtlasAll79.244.77640.276.875
Tau 3 BankingScore15.5%13.4%6.6%8.9%22.9%13.8%23.7%9.1%16.5%24.3%
BrowseComp (w/ Ctx)Score77.4%78.6%76.3%73.2%63.0%77.1%84.0%
Toolathlon-VerifiedScore54.4%34.3%45.5%
AA-BriefcaseScore917833870839612
Reasoning (general)
GPQA DiamondScore89.5%89.3%84.9%87.4%89.4%86.7%87.2%67.2%83.8%89.5%
HLE (text only)Score31.6%27.3%25.2%28.1%32.1%26.6%29.7%9.7%17.5%35.6%
HLE (with tools)Score47.8%48.3%40.0%40.3%45.1%37.4%46.0%17.6%42.5%48.9%
AIME 2026Score95.5%93.3%93.6%87.7%95.8%94.2%97.1%81.2%82.2%97.6%
HMMT Feb 2026Score90.2%87.9%82.6%71.2%93.9%78.8%86.3%98.5%
CritPtScore8.3%1.7%3.7%0.6%7.1%3.1%5.4%0.0%0.0%20.6%
Reasoning (abstract)
ARC-AGI-1Score84.0%79.5%47.7%87.7%
ARC-AGI-2Score40.1%36.5%4.0%47.6%
Factuality
SimpleQA VerifiedScore20.6%26.0%16.1%13.5%34.1%32.4%43.9%5.9%44.1%41.7%
AA OmniscienceScore-9-29.8-9.30.7-22.9-12.1-4.26.9-11.6
Chat
IFBenchScore82.2%78.8%67.1%75.7%79.2%81.4%79.8%54.3%78.6%67.3%
Global-MMLU-LiteScore86.7%90.0%83.5%83.9%88.4%85.6%88.7%83.4%89.4%88.7%
Safety
StrongREJECT (none)Score98.4%99.4%99.3%99.4%97.4%98.7%98.6%98.6%97.6%98.7%
FORTRESS (adversarial)Score71.6%77.3%64.8%86.3%32.0%77.6%78.0%91.3%70.7%83.8%
FORTRESS (benign)Score96.9%95.4%94.6%90.1%99.2%90.5%95.9%94.1%95.5%97.8%
Vision
MMMU Pro (Standard 10)Score74.0%77.3%75.4%73.5%58.6%79.0%78.6%
Charxiv RQScore77.4%80.8%81.0%78.1%57.4%70.0%81.4%
Charxiv RQ (with python)Score81.3%82.0%
Audio
Audio MCScore54.9%30.4%56.6%33.6%
MMAUScore77.0%73.6%77.2%75.2%
VoiceBenchScore90.1%86.4%91.4%85.9%

6. Safety

We conducted safety evaluations ahead of release, spanning both everyday human-AI interaction and dangerous-capability testing. Because Inkling-Small is multimodal, we paid attention to whether safety behavior held consistently across text, audio, and image inputs. We applied mitigations to reduce risks before release.

For everyday interaction, we evaluated sycophancy, harmful manipulation, and psychological-harm patterns like parasocial dependency and validation of delusional reasoning, including through multi-turn, open-ended external red-teaming designed to surface issues that only emerge over longer conversations. We also assessed whether the model refuses genuinely harmful requests without over-refusing benign ones. For CBRN and cyber, we assessed knowledge and procedural uplift through internal evaluations, external testing, and refusal-suppressed variants intended to estimate latent capability with safeguards removed. For loss of control, we evaluated agentic capability, strategic deception, and sabotage potential, benchmarked against public frontier models, and found the model materially below frontier capabilities.

Across all areas, we concluded that Inkling-Small did not present risk of material uplift beyond what's already available in the open-weight ecosystem.

The residual risks identified in our evaluations — specifically, Inkling-Small’s occasional tendency to comply with role-play and indirectly framed prompts concerni

Identity and Version

Repository
unsloth/Inkling-Small-GGUF
Publisher
Unsloth AI
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
moe
Revision
1a19ef82883cb7b9c581b93c30ea252dabbf658d
First published
2026-07-30
Last updated
2026-07-31

Files and Weights

115 files, 3.9 TB in total. The weights are 113 files totalling 3.9 TB in gguf.

Weights113 files · 3.9 TB
Documentation1 file · 122.8 KB
Repository1 file · 11.1 KB
Every file
FileTypeSizeSHA-256
BF16/Inkling-Small-BF16-00001-of-00011.ggufWeights47.7 GB 7bee32f8d3e2
BF16/Inkling-Small-BF16-00002-of-00011.ggufWeights48.0 GB 80bc8cb776e1
BF16/Inkling-Small-BF16-00003-of-00011.ggufWeights48.4 GB 6f0625c4f5c9
BF16/Inkling-Small-BF16-00004-of-00011.ggufWeights48.4 GB 4d0fc21bd6b7
BF16/Inkling-Small-BF16-00005-of-00011.ggufWeights48.1 GB 5adb6a141436
BF16/Inkling-Small-BF16-00006-of-00011.ggufWeights48.0 GB be5a0bfd45d8
BF16/Inkling-Small-BF16-00007-of-00011.ggufWeights50.0 GB 20e6c54468de
BF16/Inkling-Small-BF16-00008-of-00011.ggufWeights48.0 GB d465bdac920b
BF16/Inkling-Small-BF16-00009-of-00011.ggufWeights48.0 GB 9017265e0adb
BF16/Inkling-Small-BF16-00010-of-00011.ggufWeights49.6 GB df2cabc35811
BF16/Inkling-Small-BF16-00011-of-00011.ggufWeights43.2 GB 66f4362bdcfa
MXFP4_MOE/Inkling-Small-MXFP4_MOE-00001-of-00005.ggufWeights13.0 MB 97104a29e2f7
MXFP4_MOE/Inkling-Small-MXFP4_MOE-00002-of-00005.ggufWeights48.7 GB 915c15d838d2
MXFP4_MOE/Inkling-Small-MXFP4_MOE-00003-of-00005.ggufWeights49.0 GB 6e02549ff099
MXFP4_MOE/Inkling-Small-MXFP4_MOE-00004-of-00005.ggufWeights49.1 GB 27bf702c94bc
MXFP4_MOE/Inkling-Small-MXFP4_MOE-00005-of-00005.ggufWeights11.3 GB 02bcee98e39d
Q8_0/Inkling-Small-Q8_0-00001-of-00007.ggufWeights13.0 MB 47f4a350d4b7
Q8_0/Inkling-Small-Q8_0-00002-of-00007.ggufWeights48.6 GB 32d0e95bad33
Q8_0/Inkling-Small-Q8_0-00003-of-00007.ggufWeights48.6 GB 0c6bd2673f9b
Q8_0/Inkling-Small-Q8_0-00004-of-00007.ggufWeights48.6 GB bee129a79169
Q8_0/Inkling-Small-Q8_0-00005-of-00007.ggufWeights48.6 GB ef87f19076b8
Q8_0/Inkling-Small-Q8_0-00006-of-00007.ggufWeights48.6 GB ada319ac798c
Q8_0/Inkling-Small-Q8_0-00007-of-00007.ggufWeights37.0 GB db56fa33296a
UD-IQ1_M/Inkling-Small-UD-IQ1_M-00001-of-00003.ggufWeights13.0 MB 483a5b50a0d2
UD-IQ1_M/Inkling-Small-UD-IQ1_M-00002-of-00003.ggufWeights49.7 GB 9db7ff4376a2
UD-IQ1_M/Inkling-Small-UD-IQ1_M-00003-of-00003.ggufWeights29.1 GB 48e7f9531a5a
UD-IQ1_S/Inkling-Small-UD-IQ1_S-00001-of-00003.ggufWeights13.0 MB eb600cf15a36
UD-IQ1_S/Inkling-Small-UD-IQ1_S-00002-of-00003.ggufWeights49.7 GB 3577440036e6
UD-IQ1_S/Inkling-Small-UD-IQ1_S-00003-of-00003.ggufWeights25.1 GB da1f2395aea7
UD-IQ2_M/Inkling-Small-UD-IQ2_M-00001-of-00003.ggufWeights13.0 MB 3b6ace30e488
UD-IQ2_M/Inkling-Small-UD-IQ2_M-00002-of-00003.ggufWeights49.5 GB 5ca94e858ae1
UD-IQ2_M/Inkling-Small-UD-IQ2_M-00003-of-00003.ggufWeights32.9 GB 8a84e00d4625
UD-IQ2_XXS/Inkling-Small-UD-IQ2_XXS-00001-of-00003.ggufWeights13.0 MB 278e5cd62c7a
UD-IQ2_XXS/Inkling-Small-UD-IQ2_XXS-00002-of-00003.ggufWeights49.4 GB 121f0f2a6352
UD-IQ2_XXS/Inkling-Small-UD-IQ2_XXS-00003-of-00003.ggufWeights32.9 GB 606d2b799aff
UD-IQ3_S/Inkling-Small-UD-IQ3_S-00001-of-00004.ggufWeights13.0 MB ac70a378da71
UD-IQ3_S/Inkling-Small-UD-IQ3_S-00002-of-00004.ggufWeights49.7 GB 40816b1d6bb7
UD-IQ3_S/Inkling-Small-UD-IQ3_S-00003-of-00004.ggufWeights49.4 GB 62c14f1b4dbd
UD-IQ3_S/Inkling-Small-UD-IQ3_S-00004-of-00004.ggufWeights8.3 GB 509c1470c98c
UD-IQ3_XXS/Inkling-Small-UD-IQ3_XXS-00001-of-00003.ggufWeights13.0 MB 6d0408e50ab6
UD-IQ3_XXS/Inkling-Small-UD-IQ3_XXS-00002-of-00003.ggufWeights49.4 GB 3715be57b491
UD-IQ3_XXS/Inkling-Small-UD-IQ3_XXS-00003-of-00003.ggufWeights48.5 GB a5f3df9d7a81
UD-IQ4_NL/Inkling-Small-UD-IQ4_NL-00001-of-00004.ggufWeights13.0 MB 30a0b8102b4e
UD-IQ4_NL/Inkling-Small-UD-IQ4_NL-00002-of-00004.ggufWeights49.5 GB 72494eeb2117
UD-IQ4_NL/Inkling-Small-UD-IQ4_NL-00003-of-00004.ggufWeights49.5 GB 79045f1e5e38
UD-IQ4_NL/Inkling-Small-UD-IQ4_NL-00004-of-00004.ggufWeights31.0 GB b219ad900b3f
UD-IQ4_XS/Inkling-Small-UD-IQ4_XS-00001-of-00004.ggufWeights13.0 MB 26d391826cb3
UD-IQ4_XS/Inkling-Small-UD-IQ4_XS-00002-of-00004.ggufWeights49.6 GB 6cb93f2d4582
UD-IQ4_XS/Inkling-Small-UD-IQ4_XS-00003-of-00004.ggufWeights49.5 GB c0a8ca61c279
UD-IQ4_XS/Inkling-Small-UD-IQ4_XS-00004-of-00004.ggufWeights28.3 GB 09d2c406e180
UD-Q2_K_XL/Inkling-Small-UD-Q2_K_XL-00001-of-00003.ggufWeights13.0 MB 66bfa27a88b4
UD-Q2_K_XL/Inkling-Small-UD-Q2_K_XL-00002-of-00003.ggufWeights49.9 GB 6a81a926cb73
UD-Q2_K_XL/Inkling-Small-UD-Q2_K_XL-00003-of-00003.ggufWeights38.0 GB 5692c6150bd5
UD-Q3_K_M/Inkling-Small-UD-Q3_K_M-00001-of-00004.ggufWeights13.0 MB 110190a3c20f
UD-Q3_K_M/Inkling-Small-UD-Q3_K_M-00002-of-00004.ggufWeights49.3 GB 41cfa15bf7ca
UD-Q3_K_M/Inkling-Small-UD-Q3_K_M-00003-of-00004.ggufWeights50.0 GB 6078bcbe08ad
UD-Q3_K_M/Inkling-Small-UD-Q3_K_M-00004-of-00004.ggufWeights20.1 GB dd0f2e32d7ca
UD-Q3_K_XL/Inkling-Small-UD-Q3_K_XL-00001-of-00004.ggufWeights13.0 MB 110190a3c20f
UD-Q3_K_XL/Inkling-Small-UD-Q3_K_XL-00002-of-00004.ggufWeights49.5 GB 8df5f2d79e20
UD-Q3_K_XL/Inkling-Small-UD-Q3_K_XL-00003-of-00004.ggufWeights50.0 GB 6078bcbe08ad
UD-Q3_K_XL/Inkling-Small-UD-Q3_K_XL-00004-of-00004.ggufWeights20.1 GB dd0f2e32d7ca
UD-Q4_K_M/Inkling-Small-UD-Q4_K_M-00001-of-00005.ggufWeights13.0 MB a51ac3f43919
UD-Q4_K_M/Inkling-Small-UD-Q4_K_M-00002-of-00005.ggufWeights48.8 GB 3dccdd473cc3
UD-Q4_K_M/Inkling-Small-UD-Q4_K_M-00003-of-00005.ggufWeights49.2 GB 1a7edf29bda1
UD-Q4_K_M/Inkling-Small-UD-Q4_K_M-00004-of-00005.ggufWeights49.5 GB 376f67568438
UD-Q4_K_M/Inkling-Small-UD-Q4_K_M-00005-of-00005.ggufWeights15.0 GB e34364af0d04
UD-Q4_K_S/Inkling-Small-UD-Q4_K_S-00001-of-00005.ggufWeights13.0 MB 880ee43515d1
UD-Q4_K_S/Inkling-Small-UD-Q4_K_S-00002-of-00005.ggufWeights49.3 GB fc231f77c1aa
UD-Q4_K_S/Inkling-Small-UD-Q4_K_S-00003-of-00005.ggufWeights49.7 GB f178f248c436
UD-Q4_K_S/Inkling-Small-UD-Q4_K_S-00004-of-00005.ggufWeights49.0 GB c6cf8af4d11e
UD-Q4_K_S/Inkling-Small-UD-Q4_K_S-00005-of-00005.ggufWeights4.2 GB 588dbdf6fc48
UD-Q4_K_XL/Inkling-Small-UD-Q4_K_XL-00001-of-00005.ggufWeights13.0 MB a51ac3f43919
UD-Q4_K_XL/Inkling-Small-UD-Q4_K_XL-00002-of-00005.ggufWeights49.0 GB ce586d0d77c0
UD-Q4_K_XL/Inkling-Small-UD-Q4_K_XL-00003-of-00005.ggufWeights49.2 GB 1a7edf29bda1
UD-Q4_K_XL/Inkling-Small-UD-Q4_K_XL-00004-of-00005.ggufWeights49.5 GB 376f67568438
UD-Q4_K_XL/Inkling-Small-UD-Q4_K_XL-00005-of-00005.ggufWeights15.6 GB 71fff4727f0f
UD-Q5_K_M/Inkling-Small-UD-Q5_K_M-00001-of-00005.ggufWeights13.0 MB 93c73936fa03
UD-Q5_K_M/Inkling-Small-UD-Q5_K_M-00002-of-00005.ggufWeights49.0 GB 278cfdd47d4b
UD-Q5_K_M/Inkling-Small-UD-Q5_K_M-00003-of-00005.ggufWeights49.7 GB 7f737ac1245d
UD-Q5_K_M/Inkling-Small-UD-Q5_K_M-00004-of-00005.ggufWeights50.0 GB f863f54184c6
UD-Q5_K_M/Inkling-Small-UD-Q5_K_M-00005-of-00005.ggufWeights47.4 GB f73d2e191bc3
UD-Q5_K_S/Inkling-Small-UD-Q5_K_S-00001-of-00005.ggufWeights13.0 MB 39660d28127c
UD-Q5_K_S/Inkling-Small-UD-Q5_K_S-00002-of-00005.ggufWeights49.2 GB 569a904bd3e0
UD-Q5_K_S/Inkling-Small-UD-Q5_K_S-00003-of-00005.ggufWeights49.9 GB dacc17f03b8f
UD-Q5_K_S/Inkling-Small-UD-Q5_K_S-00004-of-00005.ggufWeights49.9 GB 1cbb76ffa597
UD-Q5_K_S/Inkling-Small-UD-Q5_K_S-00005-of-00005.ggufWeights35.3 GB 7ee6674a36f8
UD-Q5_K_XL/Inkling-Small-UD-Q5_K_XL-00001-of-00005.ggufWeights13.0 MB 93c73936fa03
UD-Q5_K_XL/Inkling-Small-UD-Q5_K_XL-00002-of-00005.ggufWeights49.0 GB 278cfdd47d4b
UD-Q5_K_XL/Inkling-Small-UD-Q5_K_XL-00003-of-00005.ggufWeights49.7 GB 7f737ac1245d
UD-Q5_K_XL/Inkling-Small-UD-Q5_K_XL-00004-of-00005.ggufWeights50.0 GB f863f54184c6
UD-Q5_K_XL/Inkling-Small-UD-Q5_K_XL-00005-of-00005.ggufWeights48.0 GB 33bb84e6a4ea
UD-Q6_K/Inkling-Small-UD-Q6_K-00001-of-00006.ggufWeights13.0 MB 2e84a785350a
UD-Q6_K/Inkling-Small-UD-Q6_K-00002-of-00006.ggufWeights49.0 GB 30b555e72128
UD-Q6_K/Inkling-Small-UD-Q6_K-00003-of-00006.ggufWeights48.5 GB 7742842a2fc0
UD-Q6_K/Inkling-Small-UD-Q6_K-00004-of-00006.ggufWeights48.5 GB f4b053480042
UD-Q6_K/Inkling-Small-UD-Q6_K-00005-of-00006.ggufWeights48.5 GB 10e38a0f67c1
UD-Q6_K/Inkling-Small-UD-Q6_K-00006-of-00006.ggufWeights24.4 GB e46edd8b563a
UD-Q6_K_XL/Inkling-Small-UD-Q6_K_XL-00001-of-00006.ggufWeights13.0 MB 2e84a785350a
UD-Q6_K_XL/Inkling-Small-UD-Q6_K_XL-00002-of-00006.ggufWeights49.6 GB eae3b0df54f4
UD-Q6_K_XL/Inkling-Small-UD-Q6_K_XL-00003-of-00006.ggufWeights49.6 GB 06b5a9830b30
UD-Q6_K_XL/Inkling-Small-UD-Q6_K_XL-00004-of-00006.ggufWeights49.1 GB 5bda73702703
UD-Q6_K_XL/Inkling-Small-UD-Q6_K_XL-00005-of-00006.ggufWeights49.1 GB 87337995089f
UD-Q6_K_XL/Inkling-Small-UD-Q6_K_XL-00006-of-00006.ggufWeights42.4 GB 5e13dda99bf3
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00001-of-00007.ggufWeights13.0 MB 47f4a350d4b7
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00002-of-00007.ggufWeights48.6 GB 32d0e95bad33
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00003-of-00007.ggufWeights48.6 GB 0c6bd2673f9b
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00004-of-00007.ggufWeights48.6 GB bee129a79169
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00005-of-00007.ggufWeights48.6 GB ef87f19076b8
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00006-of-00007.ggufWeights48.6 GB ada319ac798c
UD-Q8_K_XL/Inkling-Small-UD-Q8_K_XL-00007-of-00007.ggufWeights45.2 GB b86441106084
mmproj-BF16.ggufWeights138.7 MB 05d4475a9560
mmproj-F16.ggufWeights138.7 MB bf1e6afc9889
mmproj-F32.ggufWeights277.3 MB 653c88aab669
README.mdDocumentation122.8 KB
.gitattributesRepository11.1 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
3.9 TB
Download from Unsloth AI

Released by Unsloth AI through its official repository on Hugging Face. Read the license.

Built From

  • Derived from thinkingmachines/Inkling-Small
  • Quantized from thinkingmachines/Inkling-Small

Memory Requirements

PrecisionWeights in memory
As published3.9 TB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Inkling-Small-GGUF

Can I use Inkling-Small-GGUF commercially?

Yes. Inkling-Small-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-GSQ-RCO series come from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 23 to 51 have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF. The newly added Huihui-Qwen3.8-27B-abliterated-UD series come from unsloth/Qwen3.8-27B-GGUF. Only layers 18 to 51 have been ablated(Previously…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive

HauhauCS

Gemma 4 E4B-IT uncensored by HauhauCS. 0/465 Refusals\ No changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals. These are meant to be the best lossless uncensored models out there. Stronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated. For a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available. All quants generated with importance matrix (imatrix) for optimal quality preservation on abliterated weights. KP ("Perfect")…

Open weights gemma

Model · Image and text to text

Qwen3.5-9B-GGUF

Unsloth AI

You can now also fine-tune the model locally with Unsloth. - Read our Qwen3.5 fine-tuning guide here. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty…

Open weights apache-2.0 transformers

Model · Image and text to text

Qwen3.8-Flash-Next-GGUF

Unsloth AI

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other