# FinSeer by The Fin AI: Open-Weight Model
Source: https://savrn.com/models/finseer
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve FinSeer (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 0.2 GB | 0.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 8, 2026.

[FinSeer on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/finseer/gpus)

## Model Card

This is our first dedicated retriever for financial time-series forecasting, Financial TimeSeries Retriever (FinSeer). Paper or resources for more information: https://arxiv.org/pdf/2502.05878 The primary use of FinSeer is research on financial time-series forecasting using retrieval-augmented generation (RAG) framework. Install Package pip install InstructorEmbedding pip install -U FlagEmbedding pip install sentence-transformers==2.2.2 pip install protobuf==3.20.0 pip install yahoo-finance python -m pip install -U angle-emb pip install transformers==4.33.2 # UAE This repository and its contents are provided for academic and educational purposes only. None of the material constitutes…

Excerpt from the card by The Fin AI.

## Identity and Version

Repository

TheFinAI/FinSeer

Publisher

The Fin AI

Task

Image and text to text

Modality

Image and text

Library

transformers

Parameters

109M parameters

Languages

en

Revision

20ab7e723f515124b3557fe4284c730d163b8fd4

First published

2025-03-15

Last updated

2026-10-08

## Files and Weights

8 files, 436.5 MB in total. The weights are 1 file totalling 435.6 MB in safetensors.

Weights1 file · 435.6 MB

Configuration2 files · 1.5 KB

Tokenizer3 files · 944.6 KB

Documentation1 file · 3.0 KB

Repository1 file · 1.5 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 435.6 MB | — |
| config.json | Configuration | 763 B | — |
| special_tokens_map.json | Configuration | 695 B | — |
| README.md | Documentation | 3.0 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 711.6 KB | — |
| tokenizer_config.json | Tokenizer | 1.4 KB | — |
| vocab.txt | Tokenizer | 231.5 KB | — |

## License and Download

License

Not stated by the source

Access

Access requested at publisher

Download size

435.6 MB

[Request access from The Fin AI](https://huggingface.co/TheFinAI/FinSeer)

The Fin AI grants access through its official repository on Hugging Face.

## Built From

- Derived from BAAI/llm-embedder
- Described by arXiv:2502.05878

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 435.6 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About FinSeer

### How much GPU memory does FinSeer need?

About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run FinSeer on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

## Similar Models

Model · Image and text to text

### [Omni-Edu-27B](https://savrn.com/models/omni-edu-27b)

[Hao Liang](https://savrn.com/model-publishers/lhpku20010120)

This model is a fine-tuned version of Qwen/Qwen3.8-27B on the on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 16 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 128 - totalevalbatchsize: 128 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 3M parameters 262,144 tokens transformers

[View model](https://savrn.com/models/omni-edu-27b)

Model · Image and text to text

### [qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm](https://savrn.com/models/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm)

[Namho Koh](https://savrn.com/model-publishers/namhokaist)

Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its junction with the main stem. The model does not choose the target. Compared with the box+point models in this account (qwen3-vl-8b-tomato-cutpoint-bp-, which find the petiole themselves), this one measures localization given identity. Answer: {"cutpointuv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2…

Open weights apache-2.0 770,288 parameters 262,144 tokens

[View model](https://savrn.com/models/qwen3-vl-8b-tomato-cutpoint-cued-rgbd-9mm)

Model · Image and text to text

### [qwen3.5-35b-a3b-instruct-merged8083-sft](https://savrn.com/models/qwen3-5-35b-a3b-instruct-merged8083-sft)

[TR AKR](https://savrn.com/model-publishers/akrtr)

This model is a fine-tuned version of /mnt/shared-storage-user/mineru2-shared/niujunbo/ldy/models/Qwen3.5-35B-A3B-Instruct on the merged8083ulogocohqwen35 dataset. The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1.0 - Transformers 5.6.0 - Pytorch 2.10.0+cu128 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 664,944 parameters 262,144 tokens transformers

[View model](https://savrn.com/models/qwen3-5-35b-a3b-instruct-merged8083-sft)

Model · Image and text to text

### [Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF](https://savrn.com/models/qwen3-8-27b-imatrix-nvfp4-mtp-gguf)

[Michał Piszczek](https://savrn.com/model-publishers/cdiamond)

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

[View model](https://savrn.com/models/qwen3-8-27b-imatrix-nvfp4-mtp-gguf)

Model · Image and text to text

### [Qwen3.8-Flash-Next-GSQ-RCO-GGUF](https://savrn.com/models/qwen3-8-flash-next-gsq-rco-gguf)

[IST Austria Distributed Algorithms and Systems Lab](https://savrn.com/model-publishers/ista-daslab)

Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…

Open weights apache-2.0 gguf

[View model](https://savrn.com/models/qwen3-8-flash-next-gsq-rco-gguf)

Model · Image and text to text

### [Huihui-Qwen3.8-27B-abliterated-GGUF](https://savrn.com/models/huihui-qwen3-8-27b-abliterated-gguf)

[Huihui.ai](https://savrn.com/model-publishers/huihui-ai)

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

[View model](https://savrn.com/models/huihui-qwen3-8-27b-abliterated-gguf)

## The Fin AI

[All models and datasets](https://savrn.com/model-publishers/thefinai)

## Versions

- [20ab7e723f51](https://savrn.com/models/finseer/versions/20ab7e723f51) · current 2026-10-08

## Explore More

- [All image and text to text models](https://savrn.com/models/tasks/image-and-text-to-text)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-08.
- [Hugging Face record](https://huggingface.co/TheFinAI/FinSeer)
- [How the hub is built](https://savrn.com/model-hub/methodology)
