SAVRN
Search Contact SAVRN

minerva-mlm-8k · Model Card

minerva-mlm-8k: Model Card

Written by Garyk Brixi, published under apache-2.0, revision ee385d0b0b16, read 2026-10-05. Shown as written; SAVRN's own facts about this model are on its page.

Minerva-MLM

Minerva-MLM is a genome language model for coevolutionary mining. The checkpoint includes the masked language model head and three interaction heads for base-pairing, repeat, and structure signals.

Full documentation, finetuning scripts, notebooks, and setup are on the github. Try it in the browser here.

Checkpoints

Model Context length Hugging Face repo
Minerva-MLM 4,096 tokens gbrixi/minerva-mlm
Minerva-MLM-8k 8,192 tokens gbrixi/minerva-mlm-8k

Install

Inference from this repo needs only Transformers and PyTorch:

pip install "transformers>=4.41" torch safetensors
pip install flash-attn      # optional, recommended for production and packed sequences

flash-attn is picked up automatically when present, otherwise Minerva falls back to PyTorch SDPA. The optional minerva-dna package adds plotting helpers, GenBank utilities and finetuning wrappers, and lets you load the model without trust_remote_code:

pip install minerva-dna

Quick Start

import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer

repo = "gbrixi/minerva-mlm"  # or "gbrixi/minerva-mlm-8k" for 8k context
device = "cuda" if torch.cuda.is_available() else "cpu"

model = AutoModelForMaskedLM.from_pretrained(
    repo,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).to(device).eval()
tokenizer = AutoTokenizer.from_pretrained(repo)

sequence = "cgcggggtggagcagcctggtagctcgtcgggctcataacccgaagatcgtcggttcaaatccggcccccgcaacca"
tokens = tokenizer(f"<+>{sequence.lower()}", return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model(**tokens, output_interactions=True)

base_pairing = outputs.interactions["base_pairing"]  # [batch, L, L]
protein = outputs.interactions["protein"]            # [batch, L, L]
repeat = outputs.interactions["repeat"]              # [batch, L, L]
logits = outputs.logits                              # [batch, L, vocab]

With minerva-dna installed, import the class directly and drop trust_remote_code:

from minerva import MinervaForMaskedLM

model = MinervaForMaskedLM.from_pretrained(
    repo, torch_dtype=torch.bfloat16,
).to(device).eval()

Strand Orientation

Minerva was trained with positive-strand DNA. To intepret unknown loci, run inference on both strands and compare results (or combine maps downstream):

def reverse_complement(seq: str) -> str:
    return seq.lower().translate(str.maketrans("acgtn", "tgcan"))[::-1]

strands = {
    "forward": f"<+>{sequence.lower()}",
    "reverse_complement": f"<+>{reverse_complement(sequence)}",
}

results = {}
for name, seq in strands.items():
    tokens = tokenizer(seq, return_tensors="pt").to(model.device)
    with torch.no_grad():
        outputs = model(**tokens, output_interactions=True)
    results[name] = outputs.interactions

For mixed DNA/protein token strings, minerva.sequence_utils.extract_fwd_rc_attention aligns and combines forward and reverse-complement attention maps automatically (requires pip install minerva-dna).

Interaction Outputs

Set output_interactions=True to return all three standard Minerva interaction maps:

  • base_pairing: RNA base-pairing contacts
  • protein: protein contact prediction
  • repeat: repeat element signal

By default, Minerva uses the last-2-layer interaction heads:

outputs = model(**tokens, output_interactions=True)

To plot the interaction-head outputs (requires pip install minerva-dna):

from minerva.visualization import plot_contacts

token_list = tokenizer.convert_ids_to_tokens(tokens["input_ids"][0].tolist())
plot_contacts(outputs, tokens=token_list, track=True)

Six-layer heads are available with interaction_layers=6.

Jacobian Fingerprinting

get_fingerprints returns named interaction-pattern channels for a sequence, computing the categorical Jacobian internally.

fp = model.get_fingerprints(sequence, tokenizer, position_range=(0, 96))
fp.channel_names          # ["basepairing", "repeat", "protein", "other"]
basepairing = fp["basepairing"]   # [L, L]

Raw Jacobians, plotting helpers and the rest of the API are documented in the repository.

Finetuning

Finetuning scripts, LoRA support, and GenBank ingestion live in the GitHub repository. See scripts/finetune.py and the README there for full examples with accelerate and PEFT.

Limitations

  • Model performance is best on the positive DNA strand (<+> prefix). When orientation is unknown, run inference on both strands.
  • trust_remote_code=True is needed only without minerva-dna installed (see Install).

Citation

If you use Minerva, please cite Li & Brixi et al., bioRxiv 2026:

@article{li2026minerva,
  title     = {Coevolutionary mining of prokaryotic non-coding elements with a genome language model},
  author    = {Li, David B. and Brixi, Garyk and Kim, Alexandra S. and Fiamenghi, Mateus B. and Driscoll, Claudia L. and Evans, Simone A. and Gao, Alex and Ivanova, Natalia N. and Kyrpides, Nikos C. and Deisseroth, Karl and Wilkinson, Max E. and Fischbach, Michael A. and Hie, Brian L.},
  journal   = {bioRxiv},
  year      = {2026},
  doi       = {10.64898/2026.09.22.753630},
  url       = {https://www.biorxiv.org/content/10.64898/2026.09.22.753630},
  publisher = {Cold Spring Harbor Laboratory}
}

License

Apache 2.0. Minerva-MLM is initialized from gLM2 650M (Tatta Bio, Apache 2.0) and adopts its mixed-modality tokenization.