Automatically collected, ML-ready snapshots from the public Nextbike Maps feed for the Bajs Zagreb system (city=1172, domain=hd). - latest is the most recently published station snapshot and is the default browser view. - stationstatus is the complete station-level time series. - bikemovements contains changes inferred by comparing bike identifiers between consecutive successful observations. - systemstatus contains city-level totals reported by the source API. All timestamps are UTC. Historical Parquet shards are immutable. Counts reported at city level are preserved as reported and may differ from sums across returned stations. Bike movements are observations inferred from the live feed.…
SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Append-only archives of public Polymarket CLOB WebSocket messages for markets eligible for liquidity rewards. The dataset is intended for market microstructure research, backtests and machine-learning experiments. No wallet data, private keys or authenticated trading information is collected. Files are gzip-compressed JSON Lines and grouped by source and UTC date. The source data comes from Polymarket's public CLOB interfaces and remains subject to the applicable source-platform terms.
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen, Alex Wong, Tomas Hodan, and Kun He CVPR 2026; https://arxiv.org/abs/2603.28760 SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild. It is intended to advance research on egocentric 3D hand–object interaction understanding, and generalization of perception models to real-world
Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines. See manifest.json. Typical sizes: constituents ~20k, sectordaily ~180k, tickerdaily ~4.6M rows. - Daily returns are cap-weighted across constituents; index base level = 100. - Constituents require valid daily bars in dojostockkline and positive market cap from dojoquote.
Pustaka adalah korpus teks terbuka berskala besar yang dikembangkan oleh MEA Ecosystem, dimulai dengan fokus penuh pada Bahasa Indonesia. Corpus ini disusun dari beberapa sumber publik berkualitas — web crawl yang telah difilter, ensiklopedia, berita, forum, hingga lexicon bahasa gaul — lalu diproses ulang melalui pipeline pembersihan (deteksi bahasa, filter panjang dokumen, deteksi boilerplate/spam, dan deduplikasi) sebelum dirilis. Pustaka dibangun untuk digunakan siapa saja — baik untuk pretraining model bahasa, penelitian NLP, maupun eksperimen pribadi — dan dirilis secara terbuka di bawah Hugging Face. Tabel di atas mencakup corpus-bahasa-indonesia/ saja. Untuk corpus lain (kode…
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal pairs across 7 distinct imaging modalities, with comprehensive clinical metadata, expert lesion segmentations, and hierarchical diagnostic taxonomies. All archives are compressed into high-speed Zstandard chunks (.tar.zst) with pre-indexed SHA-256 manifests. A centralized, fully de-identified master index is provided in…
The Laval Objaverse Dataset is a comprehensive dataset designed for multi-view relighting and novel view synthesis tasks. It combines high-quality 3D assets from Objaverse with realistic, diverse illumination conditions from the Laval Indoor and Outdoor HDR datasets. Each render includes synchronized multi-view images, depth maps, and complete lighting metadata. We structured our dataset through a rigorous four-step pipeline: 1. Object Filtering: We source base meshes from the Objaverse dataset. To ensure high visual fidelity, we exclude meshes with poor geometry or materials by adopting the strict object selection criteria from the relitObjaverse dataset [[1]](#references). 2. Lighting…
DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다. 이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음). Python을 몰라도 됩니다. 웹에서 데이터를 고르고 미리본 다음, CSV·엑셀(.xlsx)로 API로 가져갈 수 있습니다. 각 파일은 회사 1곳입니다: {종목코드}.parquet (시세·거시 등 일부는 날짜·시리즈 단위). DART 정기보고서를 회사 단위 패널로 수평화합니다. 서술 본문과 XBRL 연결 표가 한 artifact 에 함께 들어가, 씁니다. DART OpenAPI(fnlttSinglAcntAll) 기반 XBRL 재무 데이터. 지배구조, 보수, 지분, 배당 등을 다루는 DART API 28종. 28종 API: dividend, employee, executive, majorHolder, treasuryStock, capitalChange, auditOpinion, stockTotal, outsideDirector, corporateBond 등. 라이브러리가 필요한 parquet 을 자동으로 내려받고 로컬에 캐시합니다. API 키가 필요…
This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration. Pretraining is executed under the Odyssey Route-B Unified Engine with PyTorch Distributed Data Parallel (DDP) scaling support. $$\text{entropy\weight} = 0.03 \times \left(1.0 - \frac{\text{step}}{20000}\right)$$ This forces the router to explore uniform row distributions, preventing it from locking onto random rows early and guaranteeing convergence to the correct state updates for secret key retrieval. Evaluations are conducted across three independent seeds (42, 100, 2026). Accuracies are reported on MMLU (500…
This dataset contains daily GPS and metadata records of public transport vehicles in Cheboksary, Russia, for the period from YYYY-MM-DD to YYYY-MM-DD. Each file corresponds to a "transport day" (which may start and end at different times depending on the actual end of public transport service, not at midnight). The data was parsed from the website buscheb.ru, which aggregates public transport data for the city of Cheboksary as a contractor. The original data may be owned by buscheb.ru and/or the Cheboksary city transport authority. Please attribute buscheb.ru as the data source. Each row in the CSV files contains the following fields: The lon and lat fields contain raw values that need to…
Hong Kong A&E Waiting Time 2025年10月14日医管局似乎启用了新的api,但直到2025年12月6日我才发现,所以10月13日-12月6日这段时间的数据都是10月13日的数据... - hospCode = 医院名称 - hospTimeEn = 时间点 - topWait = 等候时间 - t1wait = 分流类别 I (危殆) 等候时间 - t2wait = 分流类别 II (危急) 等候时间 - t3median = 分流类别 III (紧急) - 一般等候时间 - t395th = 分流类别 III (紧急) - 参考等候时间 - t45median = 分流类别 IV & V (次紧急及非紧急) - 一般等候时间 - managet1 指急症室正在治理分流类别 I/II 的病人。
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.


