Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.
Use this model
Pull with QuantumShield
quantumshield pull instruction-pretrain/general-instruction-augmented-corpora Verify integrity
quantumshield verify instruction-pretrain/general-instruction-augmented-corpora pip install
pip install quantumshield && quantumshield pull instruction-pretrain/general-instruction-augmented-corpora Unverified Model
This model has not been PQC-verified. File integrity cannot be guaranteed against quantum threats.
README.md
general-instruction-augmented-corpora
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.
Intended Uses
This model is registered on the QuantaMrkt quantum-safe registry. This model has not yet been PQC-verified.
Quick Start
# Install the CLI pip install quantumshield # Pull the model quantumshield pull instruction-pretrain/general-instruction-augmented-corpora # Verify file integrity quantumshield verify instruction-pretrain/general-instruction-augmented-corpora
About
Instruction Pre-Training: Language Models are Supervised Multitask Learners (EMNLP 2024) This repo contains the general instruction-augmented corpora (containing 200M instruction-response pairs covering 40+ task categories) used in our paper Instruction Pre-Training: Language Models are Supervised Multitask Learners. We explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/instruction-pretrain/general-instruction-augmented-corpora.
Get this model
Pull with QuantumShield
quantumshield pull instruction-pretrain/general-instruction-augmented-corpora Verify signatures
quantumshield verify instruction-pretrain/general-instruction-augmented-corpora