SAVRN
Search Contact SAVRN

Open-weight model · Summarization

distilbart-cnn-12-6

by Sam Shleifer sshleifer/distilbart-cnn-12-6

This checkpoint should be loaded into BartForConditionalGeneration.frompretrained. See the BART docs for more information.

Parameters
Context1,024
Weights4.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads838.8k

Model Card

By Sam Shleifer, published under apache-2.0, revision a4f8f3ea906e.

This checkpoint should be loaded into BartForConditionalGeneration.frompretrained. See the BART docs for more information.

Read Sam Shleifer's full model card

Usage

This checkpoint should be loaded into BartForConditionalGeneration.from_pretrained. See the BART docs for more information.

Metrics for DistilBART models

Model Name MM Params Inference Time (MS) Speedup Rouge 2 Rouge-L
distilbart-xsum-12-1 222 90 2.54 18.31 33.37
distilbart-xsum-6-6 230 132 1.73 20.92 35.73
distilbart-xsum-12-3 255 106 2.16 21.37 36.39
distilbart-xsum-9-6 268 136 1.68 21.72 36.61
bart-large-xsum (baseline) 406 229 1 21.85 36.50
distilbart-xsum-12-6 306 137 1.68 22.12 36.99
bart-large-cnn (baseline) 406 381 1 21.06 30.63
distilbart-12-3-cnn 255 214 1.78 20.57 30.00
distilbart-12-6-cnn 306 307 1.24 21.26 30.59
distilbart-6-6-cnn 230 182 2.09 20.17 29.70

Configuration

Architecture
BartForConditionalGeneration
Context length (tokens)
1,024
Layers
12
Vocabulary size
50,264
Model type
bart

Identity and Version

Repository
sshleifer/distilbart-cnn-12-6
Publisher
Sam Shleifer
Task
Summarization
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
en
Revision
a4f8f3ea906ed274767e9906dbaede7531d660ff
First published
2022-03-02
Last updated
2021-06-14

Files and Weights

9 files, 4.1 GB in total. The weights are 3 files totalling 4.1 GB in bin, msgpack, ot.

Weights3 files · 4.1 GB
Configuration1 file · 1.8 KB
Tokenizer3 files · 1.4 MB
Documentation1 file · 1.7 KB
Repository1 file · 391 B
Every file
FileTypeSizeSHA-256
flax_model.msgpackWeights1.2 GB 2e850d264574
pytorch_model.binWeights1.2 GB 3bac65d18c99
rust_model.otWeights1.6 GB 8e589ff34942
config.jsonConfiguration1.8 KB
README.mdDocumentation1.7 KB
.gitattributesRepository391 B
merges.txtTokenizer456.3 KB
tokenizer_config.jsonTokenizer26 B
vocab.jsonTokenizer898.8 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.1 GB
Download from Sam Shleifer

Released by Sam Shleifer through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) cnn_dailymail
  • Trained on (disclosed) xsum

Memory Requirements

PrecisionWeights in memory
As published4.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About distilbart-cnn-12-6

Can I use distilbart-cnn-12-6 commercially?

Yes. distilbart-cnn-12-6 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is distilbart-cnn-12-6's context length?

1,024 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Summarization

pegasus-xsum

Google

Original TF 1 code here Authors: Jingqing Zhang, Yao Zhao, Mohammad Saleh and Peter J. Liu on Dec 18, 2019 The following is copied from the authors' README. We train a pegasus model with sampled gap sentence ratios on both C4 and HugeNews, and stochastically sample important sentences. The updated the results are reported in this table. The "Mixed & Stochastic" model has the following changes: - trained on both C4 and HugeNews (dataset mixture is weighted by their number of examples). - trained for 1.5M instead of 500k (we observe slower convergence on pretraining perplexity). - the model uniformly sample a gap sentence ratio between 15% and 45%. - importance sentences are sampled using a…

Open weights 512 tokens transformers

Model · Summarization

distilbart-xsum-12-6

Sam Shleifer

This checkpoint should be loaded into BartForConditionalGeneration.frompretrained. See the BART docs for more information.

Open weights apache-2.0 1,024 tokens transformers

This repository contains the mT5 checkpoint finetuned on the 45 languages of XL-Sum dataset. For finetuning details and scripts, see the paper and the official repository. Scores on the XL-Sum test sets are as follows: Language | ROUGE-1 / ROUGE-2 / ROUGE-L Amharic | 20.0485 / 7.4111 / 18.0753 Arabic | 34.9107 / 14.7937 / 29.1623 Azerbaijani | 21.4227 / 9.5214 / 19.3331 Bengali | 29.5653 / 12.1095 / 25.1315 Burmese | 15.9626 / 5.1477 / 14.1819 Chinese (Simplified) | 39.4071 / 17.7913 / 33.406 Chinese (Traditional) | 37.1866 / 17.1432 / 31.6184 English | 37.601 / 15.1536 / 29.8817 French | 35.3398 / 16.1739 / 28.2041 Gujarati | 21.9619 / 7.7417 / 19.86 Hausa | 39.4375 / 17.6786 / 31.6667…

Open weights transformers

Model · Summarization

distilbart-cnn-6-6

Sam Shleifer

This checkpoint should be loaded into BartForConditionalGeneration.frompretrained. See the BART docs for more information.

Open weights apache-2.0 1,024 tokens transformers