SAVRN
Search Contact SAVRN

Research paper · 2024-03-28

Model Stock: All we need is just a few fine-tuned models

Dong-Hwan Jang, Sangdoo Yun, Dongyoon Han

3 open models in the SAVRN Model Hub cite Model Stock: All we need is just a few fine-tuned models (2024). Together they draw 44 downloads a month. The most downloaded is RPBizkit-v7-12B by Ricardo_estep (text generation, 12.2B parameters).

Published2024-03-28
Authors3
Citing Models3
arXiv2403.19522

Abstract

This paper introduces an efficient fine-tuning method for large pre-trained models, offering strong in-distribution (ID) and out-of-distribution (OOD) performance. Breaking away from traditional practices that need a multitude of fine-tuned models for averaging, our approach employs significantly fewer models to achieve final weights yet yield superior accuracy. Drawing from key insights in the weight space of fine-tuned weights, we uncover a strong link between the performance and proximity to the center of weight space. Based on this, we introduce a method that approximates a center-close weight using only two fine-tuned models, applicable during or after training. Our innovative layer-wise weight averaging technique surpasses state-of-the-art model methods such as Model Soup, utilizing only two fine-tuned models. This strategy can be aptly coined Model Stock, highlighting its reliance on selecting a minimal number of models to draw a more optimized-averaged model. We demonstrate the efficacy of Model Stock with fine-tuned models based upon pre-trained CLIP architectures, achieving remarkable performance on both ID and OOD tasks on the standard benchmarks, all while barely bringing extra computational demands. Our code and pre-trained models are available at https://github.com/naver-ai/model-stock.

Full paper on arXiv · Code

Details

arXiv identifier
2403.19522
Published
2024-03-28
Authors
Dong-Hwan Jang, Sangdoo Yun, Dongyoon Han

Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

ModelTaskSizeLicenseMonthly downloadsCheapest setup at 16-bit
RPBizkit-v7-12B
Ricardo_estep
Text generation 12.2B — 44 1x MI300X $1.85/hr
RPBizkit-v8-12B
Ricardo_estep
Text generation — — — —
RPBizkit-v10-12B
Ricardo_estep
Text generation 12.2B — — 1x MI300X $1.85/hr