SAVRN
Search Contact SAVRN

Open-weight model · Text classification

cryptobert

by Mikolaj Kulakowski ElKulako/cryptobert

For academic reference, cite the following paper: https://ieeexplore.ieee.org/document/10223689 CryptoBERT is a pre-trained NLP model to analyse the language and sentiments of cryptocurrency-related social media posts and messages.

Parameters125M
Context514
Weights997.3 MB
Licensemit
AccessOpen weights
Monthly Downloads490.6k

Runs On

What it takes to serve cryptobert (125M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on cryptobert

Three labels, Bearish, Neutral and Bullish, applied to short social media posts about cryptocurrency. That is the scope of this 125-million-parameter RoBERTa classifier, built by continuing to train vinai's bertweet-base on more than 3.2 million crypto-related posts. Half-precision weights are 0.2 GB and it needs 0.3 GB to run, against 192 GB on the single MI300X we list cheapest at $1.85 per hour. Score posts in batches on a shared GPU rather than renting a card for this alone.

MIT is a short permissive license: commercial use, modification and redistribution, with the notices carried along. Our check would be on the data. The file lists ElKulako/stocktwits-crypto as the training set, and the 514-token context fits a post, not a thread. Confirm your own feed resembles a corpus assembled for a June 2022 release; the weights were last touched in May 2025 and the labels are fixed at three.

Model Card

By Mikolaj Kulakowski, published under mit, revision 9e37c910fe87.

For academic reference, cite the following paper: https://ieeexplore.ieee.org/document/10223689

CryptoBERT

CryptoBERT is a pre-trained NLP model to analyse the language and sentiments of cryptocurrency-related social media posts and messages. It was built by further training the vinai's bertweet-base language model on the cryptocurrency domain, using a corpus of over 3.2M unique cryptocurrency-related social media posts. (A research paper with more details will follow soon.)

Classification Training

The model was trained on the following labels: "Bearish" : 0, "Neutral": 1, "Bullish": 2

CryptoBERT's sentiment classification head was fine-tuned on a balanced dataset of 2M labelled StockTwits posts, sampled from ElKulako/stocktwits-crypto.

CryptoBERT was trained with a max sequence length of 128. Technically, it can handle sequences of up to 514 tokens, however, going beyond 128 is not recommended.

Classification Example

Read the full model card (416 words)

Configuration

Architecture
RobertaForSequenceClassification
Context length (tokens)
514
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
50,265
Stored precision
float32
Model type
roberta

Identity and Version

Repository
ElKulako/cryptobert
Publisher
Mikolaj Kulakowski
Task
Text classification
Modality
Text
Library
transformers
Parameters
125M parameters
Languages
en
Revision
9e37c910fe87727cb842a9ac55c6388256fe0f15
First published
2022-06-20
Last updated
2025-05-26

Files and Weights

10 files, 1.0 GB in total. The weights are 2 files totalling 997.3 MB in bin, safetensors.

Weights2 files · 997.3 MB
Configuration2 files · 1.9 KB
Tokenizer4 files · 3.4 MB
Documentation1 file · 3.7 KB
Repository1 file · 1.2 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights498.6 MB fe34ba09b3d7
pytorch_model.binWeights498.7 MB 80aaa3c6d754
config.jsonConfiguration932 B
special_tokens_map.jsonConfiguration957 B
README.mdDocumentation3.7 KB
.gitattributesRepository1.2 KB
merges.txtTokenizer456.4 KB
tokenizer.jsonTokenizer2.1 MB
tokenizer_config.jsonTokenizer1.3 KB
vocab.jsonTokenizer798.3 KB

License and Download

License
mit
Access
Open weights, no gate
Download size
997.3 MB
Download from Mikolaj Kulakowski

Released by Mikolaj Kulakowski through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) ElKulako/stocktwits-crypto

Memory Requirements

PrecisionWeights in memory
As published997.3 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About cryptobert

How much GPU memory does cryptobert need?

About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (125M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run cryptobert on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use cryptobert commercially?

Yes. cryptobert is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is cryptobert's context length?

514 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

roberta-base-go_emotions

Sam Lowe

Model trained from roberta-base on the goemotions dataset for multi-label classification. A version of this model in ONNX format (including an INT8 quantized ONNX version) is now available at https://huggingface.co/SamLowe/roberta-base-goemotions-onnx. These are faster for inference, esp for smaller batch sizes, massively reduce the size of the dependencies required for inference, make inference of the model more multi-platform, and in the case of the quantized version reduce the model file/download size by 75% whilst retaining almost all the accuracy if you only need inference. goemotions is based on Reddit data and has 28 labels. It is a multi-label dataset where one or multiple labels…

Open weights mit 125M parameters 514 tokens transformers

Model · Text classification

turn-detector

LiveKit

An open-weights language model for contextually-aware end-of-utterance (EOU) detection in voice AI applications. The model predicts whether a user has finished speaking based on the semantic content of their transcribed speech, providing a critical complement to voice activity detection (VAD) systems. Traditional voice agents rely on voice activity detection (VAD) to determine when a user has finished speaking. VAD works by detecting the presence or absence of speech in an audio signal and applying a silence timer. While effective for detecting pauses, VAD lacks language understanding and frequently causes false positives. For example, a user who says "I need to think about that for a…

Open weights other 135M parameters 8,192 tokens transformers

This model is distilled from the zero-shot classification pipeline on the Multilingual Sentiment dataset using this script. In reality the multilingual-sentiment dataset is annotated of course, but we'll pretend and ignore the annotations for the sake of example. Result can be reproduce using the following commands: If you are training this model on Colab, make the following code changes to avoid Out-of-memory error message: - Transformers 4.28.1 - Pytorch 2.0.0+cu118 - Datasets 2.11.0 - Tokenizers 0.13.3

Open weights apache-2.0 135M parameters 512 tokens transformers

Model · Text classification

inclusively-classification

E-MIMIC

This model is an Italian classification model fine-tuned from the Italian BERT model for the classification of inclusive language in Italian. It has been trained to detect three classes: - inclusive: the sentence is inclusive (e.g. "Il personale docente e non docente") - notinclusive: the sentence is not inclusive (e.g. "I professori") - notpertinent: the sentence is not pertinent to the task (e.g. "La scuola è chiusa") The model has been trained on a dataset containing: - 8580 training sentences - 1073 validation sentences - 1072 test sentences The data collection has been manually annotated by experts in the field of inclusive language (dataset is not publicly available yet). The model…

Open weights cc-by-nc-sa-4.0 111M parameters 512 tokens transformers

Model · Text classification

multi-domain-sentiment-bert

ADITYA GUPTA

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 109M parameters 512 tokens transformers

Model · Text classification

assign5autotrain

Harsha B Setty

libraryname: transformers - autotrain - text-classification basemodel: google-bert/bert-base-uncased f1macro: 0.7533020080884588 f1micro: 0.7533333333333333 f1weighted: 0.7533020080884587 precisionmacro: 0.7551310982162045 precisionmicro: 0.7533333333333333 precisionweighted: 0.7551310982162046 recallmacro: 0.7533333333333333 recallmicro: 0.7533333333333333 recallweighted: 0.7533333333333333

Open weights 109M parameters 512 tokens transformers