1. BigCodeBench-Complete: Code Completion based on the structured docstrings. 1. BigCodeBench-Instruct: Code Generation based on the NL-oriented instructions. The overall statistics of the dataset are as follows: The function-calling (tool use) statistics of the dataset are as follows: BigCodeBench is an easy-to-use benchmark which evaluates LLMs with practical and challenging programming tasks. The dataset was created as part of the BigCode Project, an open scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs). BigCodeBench serves as a fundamental benchmark for LLMs instead of LLM Agents, i.e., code-generating AI systems that enable…