Model Hub
Browse PQC-verified AI models, datasets, and tools
Nous Research's Hermes 3 built on Llama 3.1. Strong function calling, structured output, and agentic capabilities.
π MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens π MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. π MINT-1T is designed to facilitate research in multimodal pretraining. π MINT-1T is created by a team from the University of Washington inβ¦ See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.
FineNews FineNews is a multilingual news-text dataset for language-model research. It filters and deduplicates a 2021β2025 snapshot of the INFINI-NEWS Corpus, which extracts articles from Common Crawl CC-News. At a glance Measure FineNews Input articles 852,824,802 Infini-News rows from 2021β2025 Output 392,627,654 physical rows Files 294,509 Parquet files Folders 60 publication months (2021-01 to 2025-12), then language Language folders 129β¦ See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news.
Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained analysis of system⦠See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.
I also seperately provide just the prompts in prompts.json keys are the image_id, and the values are the captions generated Captions generated by moondream: vikhyatk/moondream2 Latents generated by SDXL VAE: madebyollin/sdxl-vae-fp16-fix Embeddings generated by SigLIP: hf-hub:timm/ViT-SO400M-14-SigLIP-384 Original dataset: common-canvas/commoncatalog-cc-by Latents f32 and embeddings are f16 bytes Compute cost: 16x3090 for 3 day. Approximately.