Browse PQC-verified AI models, datasets, and tools
Data Agent Benchmark for Multi-step Reasoning (DABstep) Dataset This repository hosts a HF Dataset the supports the benchmark and leaderboard. For the main entrypoint to the benchmark, see the leaderboard here: https://huggingface.co/spaces/adyen/DABstep This Dataset has 3 splits: tasks submissions task_scores Users of the benchmark would read from the tasks split to run the baseline. The other splits are used to support the leaderboard. The datasets are in the data/context… See the full description on the dataset page: https://huggingface.co/datasets/adyen/DABstep.