SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

StockLLM

by The Fin AI TheFinAI/StockLLM

StockLLM is a model for image and text to text from The Fin AI, released under llama3.2 (access requested at publisher). It has 1.2B parameters. At 16-bit it needs about 3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 178 downloads a month.

StockLLM is an open-source fine-tuned 1B large language model as the backbone of our first retrieval-augmented generation (RAG) framework specifically designed for financial time-series forecasting.

Parameters1.2B
Context—
Weights2.5 GB
Licensellama3.2
AccessAccess requested at publisher
Monthly Downloads178

Runs On

What it takes to serve StockLLM (1.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 2.5 GB 3.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.2 GB 1.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

StockLLM on every accelerator the SAVRN Index prices, at every precision

Model Card

StockLLM is an open-source fine-tuned 1B large language model as the backbone of our first retrieval-augmented generation (RAG) framework specifically designed for financial time-series forecasting. Paper or resources for more information: https://arxiv.org/pdf/2502.05878 This repository and its contents are provided for academic and educational purposes only. None of the material constitutes financial, legal, or investment advice. No warranties, express or implied, are offered regarding the accuracy, completeness, or utility of the content. The authors and contributors are not responsible for any errors, omissions, or any consequences arising from the use of the information herein. Users…

Excerpt from the card by The Fin AI, licensed llama3.2.

Identity and Version

Repository
TheFinAI/StockLLM
Publisher
The Fin AI
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
1.2B parameters
Languages
en
Revision
ca465007246ac2a3b0a80f2fe15b74912e0385fe
First published
2025-03-15
Last updated
2026-10-08

Files and Weights

11 files, 2.5 GB in total. The weights are 2 files totalling 2.5 GB in safetensors.

Weights2 files · 2.5 GB
Configuration4 files · 13.4 KB
Tokenizer2 files · 17.3 MB
Documentation1 file · 2.6 KB
Other1 file · 241.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights2.0 GB —
model-00002-of-00002.safetensorsWeights474.0 MB —
config.jsonConfiguration938 B —
generation_config.jsonConfiguration184 B —
model.safetensors.index.jsonConfiguration12.0 KB —
special_tokens_map.jsonConfiguration325 B —
README.mdDocumentation2.6 KB —
Overview.pngOther241.4 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer17.2 MB —
tokenizer_config.jsonTokenizer51.3 KB —

License and Download

License
llama3.2
Access
Access requested at publisher
Download size
2.5 GB
Request access from The Fin AI

The Fin AI grants access through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published2.5 GB
16-bit2.5 GB
8-bit1.2 GB
4-bit0.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About StockLLM

How much GPU memory does StockLLM need?

About 3 GB at 16-bit and 0.7 GB at 4-bit: the weights (1.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run StockLLM on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is StockLLM released under?

llama3.2, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Image and text to text

PaddleOCR-VL-1.6

PaddlePaddle

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. PaddleOCR-VL-1.6 introduces a region-aware data optimization framework that identifies weak regions from the previous model, applies targeted enhancement to those regions, and improves the reliability of supervision signals. It further adopts a progressive post-training recipe based on curated data selection and reinforcement learning, pushing model performance to a higher level through staged optimization. PaddleOCR-VL-1.6 achieves a new state-of-the-art score…

Open weights apache-2.0 959M parameters 131,072 tokens PaddleOCR

Model · Image and text to text

PaddleOCR-VL-1.5

PaddlePaddle

PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing PaddleOCR-VL-1.5 is an advanced next-generation model of PaddleOCR-VL, achieving a new state-of-the-art accuracy of 94.5% on OmniDocBench v1.5. To rigorously evaluate robustness against real-world physical distortions—including scanning artifacts, skew, warping, screen photography, and illumination—we propose the Real5-OmniDocBench benchmark. Experimental results demonstrate that this enhanced model attains SOTA performance on the newly curated benchmark. Furthermore, we extend the model’s capabilities by incorporating seal recognition and text spotting tasks, while remaining a 0.9B ultra-compact VLM…

Open weights apache-2.0 959M parameters 131,072 tokens PaddleOCR

Model · Image and text to text

Qwen3.5-0.8B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Scores of Qwen3.5 models are reported…

Open weights apache-2.0 873M parameters 262,144 tokens transformers

Model · Image and text to text

surya-ocr-2

Datalab

Surya is a 650M param OCR model with these features: - Accuracy - scores 83.3% on olmOCR-bench (top under 3B params) - Multilingual - scores 87.2% on an internal benchmark set of 91 languages (more here) - Layout analysis (table, image, header, etc.) with reading order - Table recognition (rows + columns) It works on a range of documents (see usage and benchmarks). Our managed platform runs both Surya, and variants of our highest accuracy model, Chandra. Get started with $5 in free credits — sign up (takes under 30 seconds) or try our free public playground. Surya is named for the Hindu sun god, who has universal vision. The Surya code is licensed under Apache 2.0. The model weights use a…

Open weights openrail 686M parameters 262,144 tokens transformers

Model · Image and text to text

moondream2

Vik Korrapati

This repository contains the latest version of Moondream 2, our previous generation model. The latest version of Moondream is Moondream 3 (Preview). Moondream is a small vision language model designed to run efficiently everywhere. This repository contains the latest (2025-06-21) release of Moondream 2, as well as historical releases. The model is updated frequently, so we recommend specifying a revision as shown below if you're using it in a production application. Grounded Reasoning Introduces a new step-by-step reasoning mode that explicitly grounds reasoning in spatial positions within the image before answering, leading to more precise visual interpretation (e.g., chart median…

Open weights apache-2.0 1.9B parameters transformers

SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its compact size, requiring only 1.8GB of GPU RAM for video inference, it delivers robust performance on complex multimodal tasks. This efficiency makes it particularly well-suited for on-device applications where computational resources may be limited. SmolVLM2 can be used for inference on multimodal (video / image / text) tasks where the input consists of text queries along with video or one or more images.…

Open weights apache-2.0 507M parameters 8,192 tokens transformers