SAVRN
Search Contact SAVRN

Organization

Symato Team

Symato

Models in Library0
Datasets in Library1
Models on Hugging Face15
Followers13

Datasets

Dataset

cc

Symato Team

What is Symato CC? To download all WARC data from Common Crawl then filter out Vietnamese in Markdown and Plaintext format. There is 1% of Vietnamse in CC, extract all of them out should be a lot (~10TB of plaintext). Main contributors https://huggingface.co/nampdn-ai https://huggingface.co/binhvq https://huggingface.co/th1nhng0 https://huggingface.co/iambestfeed Simple quality filters To make use of raw data from common crawl, you need to do filtering

Publicly accessible mit 1K<n<10K