An independent archive of public Kalshi market trade data retrieved from Kalshi market-data API endpoints. This dataset is not affiliated with or endorsed by Kalshi. Each historicalrawN.jsonl file is an immutable numbered shard. Records contain a market ticker, market window, reported volume, and the public trades returned for that market. Trade fields can include timestamps, public trade IDs, prices, quantities, taker direction, and block-trade status. The archive contains no account credentials, private orders, portfolio data, or personal information. Data users are responsible for complying with the Kalshi Developer Agreement and other applicable terms.
SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
This repository stores a standalone, NeuROK-compatible deformation corpus under objaverse/finetunedeformation. It combines processed public animation data with generated MPM solid trajectories. Training consumers should read index/all.jsonl (or its train.jsonl and val.jsonl split files) rather than discovering NPZ files by directory traversal. - 1,998 DynamicObjaverseProcessed geometries; - 500 DyMesh / AnimateAnyMesh geometries; - 2,498 trainable rows in the canonical combined index; The generated target is 5,000 geometries × 2 accepted clips = 10,000 clips. Its deterministic 24-cell cycle balances elastic, plastic, sand, and snow, with every material parameter drawn from one of four…
finetranslations-edu-zhtw 是以 HuggingFaceFW/finetranslations-edu 為來源,將其 translatedchunks(原始多語言教育類網頁內容、先被 pivot 翻譯成英文的版本)進一步翻譯成繁體中文的資料集。 HuggingFaceFW/finetranslations-edu 收錄了原本以英文以外語言(oglanguage,涵蓋約 200 種語言)撰寫、經篩選具教育價值(eduscore)的網頁內容,並將其 pivot 翻譯成英文(translatedtext / translatedchunks)。本資料集在此基礎上,將已經是英文的 translatedchunks 逐段(chunk)翻譯成繁體中文,並依原順序重組回完整段落(translatedtextzhtw),讓來自約 200 種原始語言、已篩選過教育價值的內容,也能以繁體中文呈現。 翻譯由一套 LLM 驅動(模型本身不對外公開資訊,僅說明為 LLM),部署為多個平行推論服務以提升吞吐量。 繁體中文語言模型的持續預訓練(continued pretraining)語料補充,特別是希望涵蓋多元語言/文化來源、且具教育價值的內容。 id / url / oglanguage / oglanguagescore / eduscore / eduscoreraw / translatedtokencount / minhashclustersize:原樣保留自上游 HuggingFaceFW/finetranslations-edu。…
Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. - predictionslatest.parquet is rewritten on scheduled prediction runs - verification updates actual columns once the target hour is in the past - frontendsnapshot.json is regenerated after prediction and verification writes - targettimestamp - referencetime - leadtimehours - leadbucket - predictionmadeat - city - verified - dmitemperature2mpred - dmiwindspeed10mpred - dmiwindgusts10mpred - dmiprecipitationprobabilitypred - dmiprecipitationpred - mltemp - mlwindspeed - mlwindgust - mlrainprob - mlrainamount - actualtemp - actualwindspeed - actualwindgust - actualprecipitation - actualrainevent…
Kiiteitte history Kiiteitte が収集した、今までの選曲履歴。 1時間おきに更新されます。 型 { // 動画ID "videoid": "sm44670499", // タイトル "title": "library->w4nderers / 足立レイ、つくよみちゃん", // 投稿者 "author": "名無し。", // サムネイルのURL "thumbnail": "https://nicovideo.cdn.nimg.jp/thumbnails/44670499/44670499.91820835", // 選曲日時 "date": "2025-02-22 12:51:51", // 新しく増えたお気に入り数。不明の場合は null "newfaves": 5, // 回ったユーザーの数。不明の場合は null "spins": 13, // イチ押しリストのユーザーのURL。イチ押しリスト以外から選曲された場合は null
This dataset hosts the DuckDB database used by the It is derived from the raw benchmark result artifacts in and is packaged for leaderboard, viewer, notebook, and SQL use. The database is produced by the HAKARI-Bench implementation in Because the benchmark code, schema, and build workflow evolve over time, this dataset card intentionally points to the canonical documentation instead of duplicating implementation details here. - duckdb/hakaribench.duckdb: current leaderboard DuckDB database.
This dataset contains the complete text of several novel, fully generated with the assistance of open-source large language models (LLMs). It represents an experiment in large-scale literary creation under a human-AI collaborative paradigm: the human author designed the story architecture, character settings, historical research, and narrative pacing, while open-source LLMs carried out text expansion, dialogue generation, scene depiction, and the weaving of multiple narrative threads based on detailed prompts. In addition to individual files, this dataset is also packaged in JSON format ( )for large‑scale training, strictly following the instruction‑tuning data formats of mainstream LLMs…
Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…
Real-time vehicle detection data collected from Singapore's 90 LTA traffic cameras using CATI (Context-Aware Traffic Intelligence) — a novel FiLM-conditioned YOLOv11 detector that adapts to environmental conditions in real time. This dataset contains per-camera vehicle detection results collected continuously from Singapore's Land Transport Authority (LTA) expressway camera network. Each record captures a full detection sweep of a single camera including vehicle counts, class breakdown, directional split, and environmental context. Detections are produced by CATI — a novel architecture that injects FiLM (Feature-wise Linear Modulation) layers into YOLOv11s, conditioning the backbone on…
Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…
Users should be made aware of the risks, biases and limitations of the dataset.
Incrementally generated unfiltered candidates. This is not a final selected dataset. Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. runconfig.json pins models, source pools, parameters and implementation hashes. For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs. Global completion requires all worker reports to be complete. Assignments are backgroundid modulo numworkers. Per-worker runconfig.json pins the parameters. Workers never share a background ID. Each data/group-NNN/noise-NNNNNN.tar contains one normalized background.wav, 15 rank-NN.mask.wav / rank-NN.mixture.wav pairs…
This repository is a provenance-preserving collection of spatial measurement questions for answer-supervised training. It is built from VSI-590K, SpaceVista-Full, HiSpatial-Data, and CA-VQA. SenseNova-SI-8M is intentionally out of scope for this release. The first deliverable is the complete master collection. Smaller and larger training views will be derived only after all eligible metric examples have been retained and audited; they are not early sampling quotas. - annotations/ /measurement.parquet: canonical question/answer rows. - manifests/ /mediainventory.jsonl: unique required media and usage. - audits/ /: selection counts, validation, checksums, and source policy. - media/ /.tar…
HESRT Human-reviewed, accession-level spatial-omics artifacts with tissue images, expression matrices and source provenance. Samples (GSMs) Studies (GSEs) Expression observations Compressed packages 4,468 507 37,054,005 413.94 GB Each sample is a self-contained, checksummed tar.zst package. Browse the catalog, select the accessions you need, then download those samples. The full collection contains approximately 647.72 GB of uncompressed member data; do not clone
Private working dataset. Conference/journal paper corpus across OpenReview venues (ICLR, NeurIPS incl. D&B/position tracks, ICML incl. position, COLM, TMLR, AISTATS, UAI, ALT, MathAI, and later additions) plus the ACL Anthology family. Coverage, per-venue availability, decisions semantics, and known biases are documented authoritatively in the GitHub repo's data/README.md — read that first; per-venue counts change as the corpus grows, so they are deliberately not duplicated here. - raw/.jsonl — canonical OpenReview snapshots, one line per - papers/ /papers.jsonl — distilled metadata (schema: data/README.md) - reviews/ /reviews.jsonl — one line per forum reply, full text - extractedtext/…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.


