A FineWeb-style Greek-language text dataset extracted from Common Crawl, following the FineWeb-2 recipe adapted for Greek (ellGrek). Crawl coverage starts at CC-MAIN-2024-22 rather than Common Crawl's earliest snapshots: this project picks up right where the FineWeb-2 dataset's own Greek (ellGrek) subset leaves off (2013 through April 2024), so it extends FineWeb-2's Greek coverage forward instead of re-extracting and re-deduping ground FineWeb-2 already covers. Each processed crawl (e.g. CC-MAIN-2024-22) contributes two kinds of files, exposed as named configs (loaddataset("alexliap/greek-cc", name="CC-MAIN-2024-22"), etc. — see the viewer's dataset-config dropdown, same convention FineWeb…