mlfoundations / MINT-1T-HTML

Unverified HuggingFace

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.

Task_categories:image-To-Text Task_categories:text-Generation Language:en Size_categories:100M<n<1B Format:parquet Modality:text

94 138,820 1

Unverified Model

This model has not been PQC-verified. File integrity cannot be guaranteed against quantum threats.

README.md

MINT-1T-HTML

Intended Uses

This model is registered on the QuantaMrkt quantum-safe registry. This model has not yet been PQC-verified.

Quick Start

# Install the CLI
pip install quantumshield

# Pull the model
quantumshield pull mlfoundations/MINT-1T-HTML

# Verify file integrity
quantumshield verify mlfoundations/MINT-1T-HTML

About

Created 2026-05-06

Downloads 138,820

Likes 94

Get this model

View on HuggingFace

Pull with QuantumShield

quantumshield pull mlfoundations/MINT-1T-HTML

Verify signatures

quantumshield verify mlfoundations/MINT-1T-HTML

Links

huggingface.co/mlfoundations/MINT-1T-HTML Transparency Log API Endpoint

Signers

did:quantamrkt:regis...hield-v1

mlfoundations / MINT-1T-HTML

README.md

MINT-1T-HTML

Intended Uses

Quick Start

File Manifest

Signature Chain

HNDL Risk Assessment

Transparency Log

About

Get this model

Links

Signers

Tags

mlfoundations / MINT-1T-HTML

Use this model

README.md

MINT-1T-HTML

Intended Uses

Quick Start

File Manifest

Signature Chain

HNDL Risk Assessment

Transparency Log

About

Get this model

Links

Signers

Tags