SAVRN
Search Contact SAVRN

Open-weight model · Token classification

roberta-large-tweetner7-all

by TNER tner/roberta-large-tweetner7-all

This model is a fine-tuned version of roberta-large on the tner/tweetner7 dataset (trainall split). Model fine-tuning is done via T-NER's hyper-parameter search (see the repository for more detail).

Parameters
Context514
Weights1.4 GB
License
AccessOpen weights
Monthly Downloads163.7k

Model Card

This model is a fine-tuned version of roberta-large on the tner/tweetner7 dataset (trainall split). Model fine-tuning is done via T-NER's hyper-parameter search (see the repository for more detail). It achieves the following results on the test set of 2021: The per-entity breakdown of the F1 score on the test set are below: - creativework: 0.4760582928521859 For F1 scores, the confidence interval is obtained by bootstrap as below: Full evaluation can be found at metric file of NER and metric file of entity span. This model can be used through the tner library. Install the library via pip. TweetNER7 pre-processed tweets where the account name and URLs are converted into special formats (see…

Excerpt from the card by TNER.

Configuration

Architecture
RobertaForTokenClassification
Context length (tokens)
514
Layers
24
Hidden size
1,024
Feed-forward size
4,096
Attention heads
16
Vocabulary size
50,265
Stored precision
float32
Model type
roberta

Identity and Version

Repository
tner/roberta-large-tweetner7-all
Publisher
TNER
Task
Token classification
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
e7fbeec91dc056794ba7853db79686dd9b0acf47
First published
2022-07-02
Last updated
2022-09-27

Files and Weights

14 files, 1.4 GB in total. The weights are 1 file totalling 1.4 GB in bin.

Weights1 file · 1.4 GB
Configuration7 files · 18.2 KB
Tokenizer4 files · 2.6 MB
Documentation1 file · 8.1 KB
Repository1 file · 1.2 KB
Every file
FileTypeSizeSHA-256
pytorch_model.binWeights1.4 GB 212acd582770
config.jsonConfiguration13.3 KB
eval/metric.test_2020.jsonConfiguration1.9 KB
eval/metric.test_2021.jsonConfiguration1.9 KB
eval/metric_span.test_2020.jsonConfiguration252 B
eval/metric_span.test_2021.jsonConfiguration250 B
special_tokens_map.jsonConfiguration239 B
trainer_config.jsonConfiguration333 B
README.mdDocumentation8.1 KB
.gitattributesRepository1.2 KB
merges.txtTokenizer456.4 KB
tokenizer.jsonTokenizer1.4 MB
tokenizer_config.jsonTokenizer328 B
vocab.jsonTokenizer798.3 KB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
1.4 GB
Download from TNER

Released by TNER through its official repository on Hugging Face.

Built From

  • Trained on (disclosed) tner/tweetner7

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
tner/tweetner7 Task Token ClassificationMetric Entity Span F1 (test_2020)Comparison conditions not established 0.764276 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Entity Span F1 (test_2021)Comparison conditions not established 0.788198 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Entity Span Precision (test_2020)Comparison conditions not established 0.798643 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Entity Span Recall (test_2020)Comparison conditions not established 0.732745 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Entity Span Recall (test_2021)Comparison conditions not established 0.804788 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric F1 (test_2020)Comparison conditions not established 0.662879 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric F1 (test_2021)Comparison conditions not established 0.657455 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Macro F1 (test_2020)Comparison conditions not established 0.629722 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Macro F1 (test_2021)Comparison conditions not established 0.612467 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Macro Precision (test_2020)Comparison conditions not established 0.661849 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Macro Precision (test_2021)Comparison conditions not established 0.600517 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Macro Recall (test_2020)Comparison conditions not established 0.601312 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Macro Recall (test_2021)Comparison conditions not established 0.625252 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Precision (test_2020)Comparison conditions not established 0.692482 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Precision (test_2021)Comparison conditions not established 0.644213 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Recall (test_2020)Comparison conditions not established 0.635703 tner
Publisher reported
Evaluated revision not stated
tner/tweetner7 Task Token ClassificationMetric Recall (test_2021)Comparison conditions not established 0.671253 tner
Publisher reported
Evaluated revision not stated

Memory Requirements

PrecisionWeights in memory
As published1.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About roberta-large-tweetner7-all

What is roberta-large-tweetner7-all's context length?

514 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Token classification

stanford-deidentifier-base

Stanford AIMI

Stanford de-identifier was trained on a variety of radiology and biomedical documents with the goal of automatising the de-identification process while reaching satisfactory accuracy for use in production. Manuscript in-proceedings. These model weights are the recommended ones among all available deidentifier weights. This work was supported in part by the Medical Imaging and Data Resource Center (MIDRC), which is funded by the National Institute of Biomedical Imaging and Bioengineering (NIBIB) of the National Institutes of Health under contract 75N92020D00021 and through The Advanced Research Projects Agency for Health (ARPA-H)

Open weights mit 512 tokens transformers

Model · Token classification

bert-portuguese-ner

Luís Filipe Cunha

This model is a fine-tuned version of neuralmind/bert-base-portuguese-cased It achieves the following results on the evaluation set: This model was fine-tunned on token classification task (NER) on Portuguese archival documents. The annotated labels are: Date, Profession, Person, Place, Organization All the training and evaluation data is available at: http://ner.epl.di.uminho.pt/ The following hyperparameters were used during training: - learningrate: 2e-05 - trainbatchsize: 16 - evalbatchsize: 16 - lrschedulertype: linear - numepochs: 4 - Transformers 4.10.0.dev0 - Pytorch 1.9.0+cu111 - Datasets 1.10.2 - Tokenizers 0.10.3

Open weights mit 512 tokens transformers

Model · Token classification

privacy-filter-nemotron-GGUF

LocalAI-io

GGUF conversion of OpenMed/privacy-filter-nemotron, a fine-grained PII token-classification model — a fine-tune of openai/privacy-filter on the nvidia/Nemotron-PII dataset. It labels every token with a BIOES tag over 55 PII categories (221 classes) in a single forward pass, then decodes coherent spans with a constrained Viterbi procedure — so it can be served locally with no Python as the encoder/NER tier of a PII redactor. Where the base openai/privacy-filter covers 8 coarse categories, this fine-tune trades multilingual breadth for category depth: 55 fine-grained English categories (first/last name, government IDs, financial, healthcare, vehicle, digital, …). For the full model…

Open weights apache-2.0 gguf

Model · Token classification

punctuate-all

KREDOR

This is based on Oliver Guhr's work. The difference is that it is a finetuned xlm-roberta-base instead of an xlm-roberta-large and on twelve languages instead of four. The languages are: English, German, French, Spanish, Bulgarian, Italian, Polish, Dutch, Czech, Portugese, Slovak, Slovenian. precision recall f1-score support accuracy 0.98 84425503 macro avg 0.83 0.74 0.77 84425503 weighted avg 0.98 0.98 0.98 84425503

Open weights mit 514 tokens transformers

Model · Token classification

privacy-filter-multilingual-GGUF

LocalAI-io

GGUF conversion of OpenMed/privacy-filter-multilingual, a multilingual PII token-classification model (a fine-tune of openai/privacy-filter). It labels every token with a BIOES tag over 54 PII categories (217 classes) across 16 languages, so it can be served locally with no Python as the encoder/NER tier of a PII redactor. For the full model description, label space, evaluation, limitations, and citations, see the source model card — this card only covers the GGUF packaging and how to run it. This GGUF uses a custom architecture, openai-privacy-filter, that is not (yet) part of 1. privacy-filter.cpp (recommended) — a small standalone GGML engine for exactly this model family, on stock…

Open weights apache-2.0 gguf

Model · Token classification

unbiased-toxic-roberta-onnx

Protect AI

This model is a conversion of unitary/unbiased-toxic-roberta to ONNX format using the Optimum library. Trained models & code to predict toxic comments on 3 Jigsaw challenges: Toxic comment classification, Unintended Bias in Toxic comments, Multilingual toxic comment classification. Built by Laura Hanu at Unitary. The huggingface models currently give different results to the detoxify library (see issue here). All challenges have a toxicity label. The toxicity labels represent the aggregate ratings of up to 10 annotators according the following schema: - Very Toxic (a very hateful, aggressive, or disrespectful comment that is very likely to make you leave a discussion or give up on sharing…

Open weights apache-2.0 514 tokens transformers