SAVRN
Search Contact SAVRN

Organization

VinAI Research

vinai

Models in Library1
Datasets in Library0
Models on Hugging Face28
Followers355

Models

Model · Fill mask

bertweet-base

VinAI Research

BERTweet is the first public large-scale language model pre-trained for English Tweets. BERTweet is trained based on the RoBERTa pre-training procedure. The corpus used to pre-train BERTweet consists of 850M English Tweets (16B word tokens ~ 80GB), containing 845M Tweets streamed from 01/2012 to 08/2019 and 5M Tweets related to the COVID-19 pandemic. The general architecture and experimental results of BERTweet can be found in our paper: author = {Dat Quoc Nguyen and Thanh Vu and Anh Tuan Nguyen}, pages = {9--14}, year = {2020} Please CITE our paper when BERTweet is used to help produce published results or is incorporated into other software. For further information or requests, please go…

Open weights mit 130 tokens transformers