Data Agent Benchmark for Multi-step Reasoning (DABstep) Dataset This repository hosts a HF Dataset the supports the benchmark and leaderboard. For the main entrypoint to the benchmark, see the leaderboard here: https://huggingface.co/spaces/adyen/DABstep This Dataset has 3 splits: tasks submissions task_scores Users of the benchmark would read from the tasks split to run the baseline. The other splits are used to support the leaderboard. The datasets are in the data/context… See the full description on the dataset page: https://huggingface.co/datasets/adyen/DABstep.
Use this model
Pull with QuantumShield
quantumshield pull adyen/DABstep Verify integrity
quantumshield verify adyen/DABstep pip install
pip install quantumshield && quantumshield pull adyen/DABstep Unverified Model
This model has not been PQC-verified. File integrity cannot be guaranteed against quantum threats.
README.md
DABstep
Data Agent Benchmark for Multi-step Reasoning (DABstep) Dataset This repository hosts a HF Dataset the supports the benchmark and leaderboard. For the main entrypoint to the benchmark, see the leaderboard here: https://huggingface.co/spaces/adyen/DABstep This Dataset has 3 splits: tasks submissions task_scores Users of the benchmark would read from the tasks split to run the baseline. The other splits are used to support the leaderboard. The datasets are in the data/context… See the full description on the dataset page: https://huggingface.co/datasets/adyen/DABstep.
Intended Uses
This model is registered on the QuantaMrkt quantum-safe registry. This model has not yet been PQC-verified.
Quick Start
# Install the CLI pip install quantumshield # Pull the model quantumshield pull adyen/DABstep # Verify file integrity quantumshield verify adyen/DABstep
About
Data Agent Benchmark for Multi-step Reasoning (DABstep) Dataset This repository hosts a HF Dataset the supports the benchmark and leaderboard. For the main entrypoint to the benchmark, see the leaderboard here: https://huggingface.co/spaces/adyen/DABstep This Dataset has 3 splits: tasks submissions task_scores Users of the benchmark would read from the tasks split to run the baseline. The other splits are used to support the leaderboard. The datasets are in the data/context… See the full description on the dataset page: https://huggingface.co/datasets/adyen/DABstep.
Get this model
Pull with QuantumShield
quantumshield pull adyen/DABstep Verify signatures
quantumshield verify adyen/DABstep