This model is a fine-tuned version of GeorgeUwaifo/iviegpt2new01cresults on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 4 - evalbatchsize: 8 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 16 - lrschedulertype: linear - lrschedulerwarmupsteps: 387 - numepochs: 5 - mixedprecisiontraining: Native AMP - Transformers 5.16.1 - Pytorch 2.11.0+cu128 - Datasets 4.8.5 - Tokenizers 0.23.1
hasib-ai-chatbot is an open-weight model for text generation from Hasib. It has 124M parameters. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model.
Runs On
What it takes to serve hasib-ai-chatbot (124M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
hasib-ai-chatbot on every accelerator the SAVRN Index prices, at every precision
Model Card
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Excerpt from the card by Hasib.
Configuration
- Architecture
- GPT2LMHeadModel
- Vocabulary size
- 50,257
- Model type
- gpt2
Identity and Version
- Repository
- Jamsed/hasib-ai-chatbot
- Publisher
- Hasib
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 124M parameters
- Languages
- Not stated by the source
- Revision
- 3262dd06db7dc0723fc10f261db6df7f76ba7273
- First published
- 2026-09-28
- Last updated
- 2026-09-28
Files and Weights
7 files, 501.3 MB in total. The weights are 1 file totalling 497.8 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 497.8 MB | 87154e0d412a |
| config.json | Configuration | 834 B | — |
| generation_config.json | Configuration | 228 B | — |
| README.md | Documentation | 5.2 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 3.6 MB | — |
| tokenizer_config.json | Tokenizer | 325 B | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
- Download size
- 497.8 MB
Released by Hasib through its official repository on Hugging Face.
Built From
- Described by arXiv:1910.09700
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 497.8 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About hasib-ai-chatbot
How much GPU memory does hasib-ai-chatbot need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (124M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run hasib-ai-chatbot on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Similar Models
This is the lab4-yoru-AY-482206 checkpoint, trained from scratch on English C4. It was selected by the lowest development loss among nine experimental recipes. The independent seed repeat and reserved audit evaluation were still pending when this checkpoint was published. Reported scores are local evaluation proxies, not an official online-judge result. - Stock Hugging Face GPT2LMHeadModel: 12 layers, 12 attention heads, hidden width 768, context length 1,024, vocabulary 50,257, tied input/output embeddings. - 124,439,808 unique parameters, commonly described as GPT-2 small. The lab uses the historical 117M model-family label. - GPT-2 tokenizer; documents packed with EOS separators. No…
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).