WorldBench is a new benchmark designed to evaluate the physical understanding and prediction of modern world models and vision-language models. There are two components: The video based benchmark can be found in /scenes. There are 4 high-level categories for different physics concepts being tested. Within each, there are 3-5 scenes each with 25-50 variations. The text based benchmark is in /textualquestions. There are 4 JSON files, one per category. Code to run the evaluation for this benchmark along with instructions can be found here: https://drive.google.com/file/d/1TNHfV-mKiidl1eFWJyctBOodWJnCajA/view?usp=sharing
Independent publisher
WorldBench
worldbenchmark
Models in Library0
Datasets in Library1
Models on Hugging Face—
Followers—