SAVRN
Search Contact SAVRN

Research paper · 2026-09-27

SMAT: Simple and Efficient Merge-Aware Training

Yanggan Gu, Yuanyi Wang, Zhen Li, Shuo Cai, Yuhang Liu, Junzhuo Li, Zihao Wang, Hongxia Yang

40 open models in the SAVRN Model Hub cite SMAT: Simple and Efficient Merge-Aware Training (2026). The most downloaded is CLIP-ViT-large-patch14-SMAT-RESISC45 by Yanggangu (image classification). They are used for image classification, text generation.

Published2026-09-27
Authors8
Citing Models40
arXiv2609.33437

Abstract

Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.

Full paper on arXiv · Code

Details

arXiv identifier
2609.33437
Published
2026-09-27
Authors
Yanggan Gu, Yuanyi Wang, Zhen Li, Shuo Cai, Yuhang Liu, Junzhuo Li, Zihao Wang, Hongxia Yang

Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

ModelTaskSizeLicenseMonthly downloadsCheapest setup at 16-bit
CLIP-ViT-large-patch14-SMAT-RESISC45
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-SMAT-SVHN
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-SMAT-SUN397
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-SMAT-MNIST
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-SMAT-GTSRB
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-SMAT-EuroSAT
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-SMAT-Cars
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-SVHN
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-SUN397
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-MNIST
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-GTSRB
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-EuroSAT
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-DTD
Yanggangu
Image classification — mit — —
CLIP-ViT-large-patch14-FT-Cars
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-SVHN
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-SUN397
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-MNIST
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-RESISC45
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-GTSRB
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-DTD
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-SMAT-Cars
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-SVHN
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-RESISC45
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-SUN397
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-MNIST
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-EuroSAT
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-Cars
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-DTD
Yanggangu
Image classification — mit — —
CLIP-ViT-base-patch32-FT-GTSRB
Yanggangu
Image classification — mit — —
Llama-3.1-8B-Instruct-FT-20Minuten
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-FT-NumGLUE-ds
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-SMAT-20Minuten
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-SMAT-NumGLUE-ds
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-SMAT-NumGLUE-cm
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-FT-NumGLUE-cm
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-FT-ScienceQA
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-SMAT-MeetingBank
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-SMAT-ScienceQA
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-FT-FOMC
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr
Llama-3.1-8B-Instruct-FT-MeetingBank
Yanggangu
Text generation 8B llama3.1 — 1x MI300X $1.85/hr

By task: Image classification (29) · Text generation (11)