This model is a fine-tuned version of Qwen/Qwen3.8-27B on the on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 16 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 128 - totalevalbatchsize: 128 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2
FinSeer is a model for image and text to text from The Fin AI (access requested at publisher). It has 109M parameters. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 68 downloads a month.
This is our first dedicated retriever for financial time-series forecasting, Financial TimeSeries Retriever (FinSeer).
Runs On
What it takes to serve FinSeer (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.
FinSeer on every accelerator the SAVRN Index prices, at every precision
Model Card
This is our first dedicated retriever for financial time-series forecasting, Financial TimeSeries Retriever (FinSeer). Paper or resources for more information: https://arxiv.org/pdf/2502.05878 The primary use of FinSeer is research on financial time-series forecasting using retrieval-augmented generation (RAG) framework. Install Package pip install InstructorEmbedding pip install -U FlagEmbedding pip install sentence-transformers==2.2.2 pip install protobuf==3.20.0 pip install yahoo-finance python -m pip install -U angle-emb pip install transformers==4.33.2 # UAE This repository and its contents are provided for academic and educational purposes only. None of the material constitutes…
Excerpt from the card by The Fin AI.
Identity and Version
- Repository
- TheFinAI/FinSeer
- Publisher
- The Fin AI
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 109M parameters
- Languages
- en
- Revision
- 20ab7e723f515124b3557fe4284c730d163b8fd4
- First published
- 2025-03-15
- Last updated
- 2026-10-08
Files and Weights
8 files, 436.5 MB in total. The weights are 1 file totalling 435.6 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 435.6 MB | — |
| config.json | Configuration | 763 B | — |
| special_tokens_map.json | Configuration | 695 B | — |
| README.md | Documentation | 3.0 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 711.6 KB | — |
| tokenizer_config.json | Tokenizer | 1.4 KB | — |
| vocab.txt | Tokenizer | 231.5 KB | — |
License and Download
- License
- Not stated by the source
- Access
- Access requested at publisher
- Download size
- 435.6 MB
The Fin AI grants access through its official repository on Hugging Face.
Built From
- Derived from BAAI/llm-embedder
- Described by arXiv:2502.05878
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 435.6 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About FinSeer
How much GPU memory does FinSeer need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run FinSeer on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Similar Models
Full fine-tune of Qwen/Qwen3-VL-8B-Instruct (revision 0c351dd) for the box-conditioned variant of the deleafing cut-point task: the input is one robot head-camera frame (848x408 RGB) with its aligned depth map (metres, fed as a rendered second image) and the bounding box of the target petiole; the output is the nominal cut point 9 mm along that petiole from its junction with the main stem. The model does not choose the target. Compared with the box+point models in this account (qwen3-vl-8b-tomato-cutpoint-bp-, which find the petiole themselves), this one measures localization given identity. Answer: {"cutpointuv":[x,y]} in normalized [0,1000) coordinates of the original 848x408 frame (2…
This model is a fine-tuned version of /mnt/shared-storage-user/mineru2-shared/niujunbo/ldy/models/Qwen3.5-35B-A3B-Instruct on the merged8083ulogocohqwen35 dataset. The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1.0 - Transformers 5.6.0 - Pytorch 2.10.0+cu128 - Datasets 4.0.0 - Tokenizers 0.22.2
I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…
Model · Image and text to text
Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Non-uniform GGUF quantizations of a 512-expert MoE, produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-Flash-Next at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. A…
This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…