SAVRN
Search Contact SAVRN

Open-weight model · Fill mask

ModernBERT-Large-Instruct-Logician-v0.5-ONNX

by John Gorriceta JohnGorri/ModernBERT-Large-Instruct-Logician-v0.5-ONNX

ModernBERT-Large-Instruct-Logician-v0.5-ONNX is an open-weight model for fill mask from John Gorriceta, released under Apache License 2.0. It has 8,192-token context. Its published files total 4.0 GB.

This is an ONNX version of JohnGorri/ModernBERT-Large-Instruct-Logician-v0.5. It was automatically converted and uploaded using this Hugging Face Space.

Parameters—
Context8,192
Weights4.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By John Gorriceta, published under apache-2.0, revision 0adecd3ab46c.

This is an ONNX version of JohnGorri/ModernBERT-Large-Instruct-Logician-v0.5. It was automatically converted and uploaded using this Hugging Face Space. See the pipeline documentation for fill-mask: https://huggingface.co/docs/transformers.js/api/pipelines#modulepipelines.FillMaskPipeline This model is a fine-tuned version of ModernBERT optimized for logical reasoning, deductive analysis, and structure-based token prediction. It relies on the ModernBertForMaskedLM architecture, making it highly effective at filling in missing contextual logic tokens (fill-mask). - As an encoder-based Masked Language Model, it is not designed for long-form generative text (like ChatGPT). It excels at…

Read John Gorriceta's full model card

This is an ONNX version of JohnGorri/ModernBERT-Large-Instruct-Logician-v0.5. It was automatically converted and uploaded using this Hugging Face Space.

Usage with Transformers.js

See the pipeline documentation for fill-mask: https://huggingface.co/docs/transformers.js/api/pipelines#module_pipelines.FillMaskPipeline


ModernBERT-Large-Instruct-Logician-v0

This model is a fine-tuned version of ModernBERT optimized for logical reasoning, deductive analysis, and structure-based token prediction. It relies on the ModernBertForMaskedLM architecture, making it highly effective at filling in missing contextual logic tokens (fill-mask).

Model Details

  • Developed by: JohnGorri
  • Model Type: Masked Language Model (MLM)
  • Base Model: answerdotai/ModernBERT
  • Language: English
  • License: Apache 2.0

Intended Uses & Limitations

Use Cases

  • Logical Deductions: Evaluating context clues to fill in missing arguments or qualifiers.

Limitations

  • As an encoder-based Masked Language Model, it is not designed for long-form generative text (like ChatGPT). It excels at predicting masked tokens inside structured prompts.

Configuration

Architecture
ModernBertForMaskedLM
Context length (tokens)
8,192
Layers
28
Hidden size
1,024
Feed-forward size
2,624
Attention heads
16
Vocabulary size
50,368
Model type
modernbert

Identity and Version

Repository
JohnGorri/ModernBERT-Large-Instruct-Logician-v0.5-ONNX
Publisher
John Gorriceta
Task
Fill mask
Modality
Text
Library
transformers.js
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
0adecd3ab46c0a3324e1380fadcab5b903b3f778
First published
2026-09-21
Last updated
2026-09-21

Files and Weights

11 files, 4.0 GB in total. The weights are 6 files totalling 4.0 GB in onnx.

Weights6 files · 4.0 GB
Configuration1 file · 2.2 KB
Tokenizer2 files · 3.6 MB
Documentation1 file · 1.6 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
onnx/model.onnxWeights1.8 GB e44733eaf113
onnx/model_bnb4.onnxWeights430.2 MB 8fcad38e043b
onnx/model_int8.onnxWeights449.0 MB 1c28badc9355
onnx/model_q4.onnxWeights454.9 MB e5f38630c276
onnx/model_quantized.onnxWeights449.0 MB 1c28badc9355
onnx/model_uint8.onnxWeights449.0 MB 0bbc36b5d5fb
config.jsonConfiguration2.2 KB —
README.mdDocumentation1.6 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer3.6 MB —
tokenizer_config.jsonTokenizer435 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.0 GB
Download from John Gorriceta

Released by John Gorriceta through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published4.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ModernBERT-Large-Instruct-Logician-v0.5-ONNX

Can I use ModernBERT-Large-Instruct-Logician-v0.5-ONNX commercially?

Yes. ModernBERT-Large-Instruct-Logician-v0.5-ONNX is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is ModernBERT-Large-Instruct-Logician-v0.5-ONNX's context length?

8,192 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Fill mask

mdeberta-v3-base

Microsoft

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. With those two improvements, DeBERTa out perform RoBERTa on a majority of NLU tasks with 80GB training data. In DeBERTa V3, we further improved the efficiency of DeBERTa using ELECTRA-Style pre-training with Gradient Disentangled Embedding Sharing. Compared to DeBERTa, our V3 version significantly improves the model performance on downstream tasks. You can find more technique details about the new model from our paper. Please check the official repository for more implementation details and updates. mDeBERTa is multilingual version of DeBERTa which use the same structure as DeBERTa and was…

Open weights mit 512 tokens transformers

Model · Fill mask

deberta-v3-base

Microsoft

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. With those two improvements, DeBERTa out perform RoBERTa on a majority of NLU tasks with 80GB training data. In DeBERTa V3, we further improved the efficiency of DeBERTa using ELECTRA-Style pre-training with Gradient Disentangled Embedding Sharing. Compared to DeBERTa, our V3 version significantly improves the model performance on downstream tasks. You can find more technique details about the new model from our paper. Please check the official repository for more implementation details and updates. The DeBERTa V3 base model comes with 12 layers and a hidden size of 768. It has only 86M…

Open weights mit 512 tokens transformers

Model · Fill mask

Bio_ClinicalBERT

Emily Alsentzer

The Publicly Available Clinical BERT Embeddings paper contains four unique clinicalBERT models: initialized with BERT-Base (casedL-12H-768A-12) or BioBERT (BioBERT-Base v1.0 + PubMed 200K + PMC 270K) & trained on either all MIMIC notes or only discharge summaries. This model card describes the Bio+Clinical BERT model, which was initialized from BioBERT & trained on all MIMIC notes. The BioClinicalBERT model was trained on all notes from MIMIC III, a database containing electronic health records from ICU patients at the Beth Israel Hospital in Boston, MA. For more details on MIMIC, see here. All notes from the NOTEEVENTS table were included (~880M words). Each note in MIMIC was first split…

Open weights mit 512 tokens transformers

This model was previously named "PubMedBERT (abstracts)". You can either adopt the new model name "microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract" or update your transformers library to version 4.22+ if you need to refer to the old name. Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. Recent work shows that for domains with abundant unlabeled text, such as biomedicine, pretraining…

Open weights mit 512 tokens transformers

Model · Fill mask

deberta-v3-large

Microsoft

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. With those two improvements, DeBERTa out perform RoBERTa on a majority of NLU tasks with 80GB training data. In DeBERTa V3, we further improved the efficiency of DeBERTa using ELECTRA-Style pre-training with Gradient Disentangled Embedding Sharing. Compared to DeBERTa, our V3 version significantly improves the model performance on downstream tasks. You can find more technique details about the new model from our paper. Please check the official repository for more implementation details and updates. The DeBERTa V3 large model comes with 24 layers and a hidden size of 1024. It has 304M…

Open weights mit 512 tokens transformers

LEGAL-BERT is a family of BERT models for the legal domain, intended to assist legal NLP research, computational law, and legal technology applications. To pre-train the different variations of LEGAL-BERT, we collected 12 GB of diverse English legal text from several fields (e.g., legislation, court cases, contracts) scraped from publicly available resources. Sub-domain variants (CONTRACTS-, EURLEX-, ECHR-) and/or general LEGAL-BERT perform better than using BERT out of the box for domain-specific tasks. A light-weight model (33% the size of BERT-BASE) pre-trained from scratch on legal data with competitive performance is also available. I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N.…

Open weights cc-by-sa-4.0 512 tokens transformers