🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.
Use this model
Pull with QuantumShield
quantumshield pull HuggingFaceCode/stack-v3-train Verify integrity
quantumshield verify HuggingFaceCode/stack-v3-train pip install
pip install quantumshield && quantumshield pull HuggingFaceCode/stack-v3-train Unverified Model
This model has not been PQC-verified. File integrity cannot be guaranteed against quantum threats.
README.md
stack-v3-train
🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.
Intended Uses
This model is registered on the QuantaMrkt quantum-safe registry. This model has not yet been PQC-verified.
Quick Start
# Install the CLI pip install quantumshield # Pull the model quantumshield pull HuggingFaceCode/stack-v3-train # Verify file integrity quantumshield verify HuggingFaceCode/stack-v3-train
About
🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.
Get this model
Pull with QuantumShield
quantumshield pull HuggingFaceCode/stack-v3-train Verify signatures
quantumshield verify HuggingFaceCode/stack-v3-train