{"id":780,"slug":"huggingfacecode--stack-v3-train","name":"stack-v3-train","author":"HuggingFaceCode","description":"\n\n\n\t\n\t\t\n\t\n\t\n\t\t🥞 The Stack v3\n\t\n\n\nWhat is it?\nWhat is being released\nHow to download and use it\nDataset statistics\nDataset structure\nDataset creation\nConsiderations for using the data\nAdditional information\n\n\n\t\n\t\t\n\t\n\t\n\t\tWhat is it?\n\t\n\nThe Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.","tags":"[\"Task_categories:text-Generation\",\"Language_creators:crowdsourced\",\"Language_creators:expert-Generated\",\"Multilinguality:multilingual\",\"Language:code\",\"Size_categories:100M<n<1B\"]","license":null,"framework":null,"parameters":null,"downloads":143828,"likes":297,"verified":0,"created_at":"2026-08-04 11:23:28","updated_at":"2026-08-04 11:23:28","source_url":"https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train","source_platform":"huggingface","hf_repo_id":"HuggingFaceCode/stack-v3-train","ollama_name":"","category":"dataset","latest_version":"v1.0.0","version_count":1,"signature_count":1,"risk_level":null,"risk_score":null,"versions":[{"id":779,"model_id":780,"version":"v1.0.0","manifest_hash":"abcc147be15e9214ff9f248d8059dfd5974b2c3951bb418e873dd75a7cf22459","file_count":0,"total_size":0,"r2_manifest_key":"manifests/datasets/huggingfacecode--stack-v3-train/v1.0.0.json","created_at":"2026-08-04 11:23:28"}],"files":[],"signatures":[{"id":1352,"version_id":779,"signer_did":"did:quantamrkt:registry:shield-v1","algorithm":"ML-DSA-65","signature_hex":"ecd7ea22afe1cfa7afcd4c160ccc7cf39926352161c51a230d4dfca4802d6c89","attestation_type":"registry","signed_at":"2026-08-04 11:23:28"}],"hndl":null}