Each baseline directory contains its own data, documentation, and provenance. Gym-Anything also includes environment scripts and task assets. The default configuration is an alias for gym-anything; each configuration uses the test split. Load Benchmark-Agent with loaddataset("assassinlike/b635", "benchmark-agent", split="test"). Benchmark-Agent Qwen and Sol answers and scores: evaluations, configuration benchmark-agent-evaluations (4,290 records). Petri: 160 complete audits (80 DeepSeek Flash targets and 80 Qwen3.8-27B targets), including full trajectories, primary requirement judgments, and all 23 native reference scores. See Petri records; dataset configuration petri.
Independent publisher
Yihe Zang
assassinlike
Models in Library0
Datasets in Library1
Models on Hugging Face3
Followers—