SAVRN
Search Contact SAVRN

Open-weight model

DNABERT-2-117M

by Zhihan Zhou zhihan1996/DNABERT-2-117M

This is the official pre-trained model introduced in DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. We sincerely appreciate the MosaicML team for the MosaicBERT implementation, which serves as the base of DNABERT-2 development.

Parameters
Context512
Weights468.4 MB
License
AccessOpen weights
Monthly Downloads1.3M

Model Card

This is the official pre-trained model introduced in DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. We sincerely appreciate the MosaicML team for the MosaicBERT implementation, which serves as the base of DNABERT-2 development. DNABERT-2 is a transformer-based genome foundation model trained on multi-species genome. To load the model from huggingface: To calculate the embedding of a dna sequence

Excerpt from the card by Zhihan Zhou.

Configuration

Context length (tokens)
512
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
4,096
Stored precision
float32

Identity and Version

Repository
zhihan1996/DNABERT-2-117M
Publisher
Zhihan Zhou
Task
Not stated by the source
Modality
Other
Library
transformers
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
7bce263b15377fc15361f52cfab88f8b586abda0
First published
2023-06-26
Last updated
2025-06-30

Files and Weights

12 files, 468.6 MB in total. The weights are 1 file totalling 468.4 MB in bin.

Weights1 file · 468.4 MB
Configuration6 files · 91.5 KB
Tokenizer2 files · 168.1 KB
Documentation2 files · 12.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
pytorch_model.binWeights468.4 MB 7ff39ec77a48
bert_layers.pyConfiguration40.7 KB
bert_padding.pyConfiguration6.1 KB
config.jsonConfiguration904 B
configuration_bert.pyConfiguration1.0 KB
flash_attn_triton.pyConfiguration42.7 KB
generation_config.jsonConfiguration90 B
LICENSEDocumentation11.4 KB
README.mdDocumentation1.3 KB
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer167.9 KB
tokenizer_config.jsonTokenizer158 B

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
468.4 MB
Download from Zhihan Zhou

Released by Zhihan Zhou through its official repository on Hugging Face.

Built From

  • Described by arXiv:2306.15006

Memory Requirements

PrecisionWeights in memory
As published468.4 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DNABERT-2-117M

What is DNABERT-2-117M's context length?

512 tokens, from the maximum position embeddings in its published configuration.