Model Hub

Browse PQC-verified AI models, datasets, and tools

HuggingFaceFW/finetranslations HF Unverified

💬 FineTranslations The world's knowledge in 1+1T tokens of parallel text What is it? This dataset contains over 1 trillion tokens of parallel text in English and 500+ languages. It was obtained by translating data from 🥂 FineWeb2 into English using Gemma3 27B. We relied on datatrove's inference runner to deploy a synthetic data pipeline at scale. Its checkpointing and VLLM lifecycle management features allowed us to use leftover compute from the HF cluster… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finetranslations.

Task_categories:text-GenerationTask_categories:translationLanguage:abkLanguage:abqLanguage:absLanguage:acm
Muennighoff/multi_eurlex HF Unverified

MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource). Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU. As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels); this is multi-label classification task (given the text, predict multiple labels).

Size_categories:10M<n<100MModality:textLibrary:datasetsLibrary:mlcroissant
annoymous-1/CC-Bench HF Unverified

CC-Bench: A Cognitive Conflict Benchmark for MLLMs in Safety-Critical Visual Inspection CC-Bench is a joint medical-industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four-way multiple-choice QA format for high-risk visual inspection. This repository currently contains: 4,282 images in total 2,157… See the full description on the dataset page: https://huggingface.co/datasets/annoymous-1/CC-Bench.

Task_categories:visual-Question-AnsweringTask_categories:image-ClassificationSize_categories:1K<n<10KFormat:jsonModality:imageModality:text
CERN/ColliderML-Release-1 HF Unverified

ColliderML: Dataset Release 1 Dataset Description This dataset contains simulated high-energy physics collision events generated using the Open Data Detector (ODD) geometry within the Key4hep and ACTS (A Common Tracking Software) frameworks, representing a generic collider detector similar to those at the HL-LHC. Dataset Summary Collision Energy: 14 TeV (proton-proton) Detector: Open Data Detector (ODD) Simulation: DD4hep + Geant4 + ACTS Format: Apache Parquet… See the full description on the dataset page: https://huggingface.co/datasets/CERN/ColliderML-Release-1.

Task_categories:otherSize_categories:10M<n<100MFormat:parquetModality:timeseriesLibrary:datasetsLibrary:dask
ayayon/papas-nativas-peru-83-variedades HF Unverified

Papas Nativas Peruanas — 83 Variedades (UNSAAC 2024) Colección de imágenes de 83 variedades de papas nativas peruanas para clasificación visual mediante modelos de visión computacional. Descripción del dataset Dataset recopilado de forma colaborativa por estudiantes de Ingeniería Informática de la Universidad Nacional de San Antonio Abad del Cusco (UNSAAC) en el curso de Aprendizaje Automático (2024). Las imágenes fueron capturadas en condiciones variadas (distintos… See the full description on the dataset page: https://huggingface.co/datasets/ayayon/papas-nativas-peru-83-variedades.

Task_categories:image-ClassificationLanguage:esSize_categories:10K<n<100KFormat:imagefolderModality:imageLibrary:datasets
S
Synthefy/Nori-30M HF Unverified

Tabular-RegressionSynthefy-NoriFeatures-TransformerTabularTabular-Foundation-ModelIn-Context-Learning MEDIUM
B
black-forest-labs/FLUX.1-Kontext-dev HF PQC Verified

Image-To-ImageDiffusersSafetensorsImage GenerationFluxDiffusion-Single-File HIGH
D
depth-anything/DA3-LARGE HF Unverified

Depth-EstimationDepth-Anything-3SafetensorsComputer-VisionMonocular-DepthMulti-View-Geometry HIGH
F
facebook/vjepa2-vitg-fpc64-256 HF Unverified

Video-ClassificationTransformersSafetensorsVjepa2Feature ExtractionVideo HIGH
A
Abiray/Sulphur-2-base-GGUF HF Unverified

Text-To-VideoGGUFQuantizedBase_model:SulphurAI/Sulphur-2-BaseBase_model:quantized:SulphurAI/Sulphur-2-Base CRITICAL
legacy-datasets/wikipedia HF Unverified

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).

Task_categories:text-GenerationTask_categories:fill-MaskTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:no-AnnotationLanguage_creators:crowdsourced
L
LiheYoung/depth_anything_vitb14 HF Unverified

Depth-EstimationTransformersPyTorchDepth_anything MEDIUM
sarulab-speech/yodas2_sidon HF Unverified

YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.

Task_categories:text-To-SpeechTask_categories:automatic-Speech-RecognitionLanguage:aaLanguage:abLanguage:afLanguage:ak
Idavidrein/gpqa HF Unverified

Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.

Benchmark:officialBenchmark:eval-YamlTask_categories:question-AnsweringTask_categories:text-GenerationLanguage:enSize_categories:1K<n<10K
S
Synthefy/Nori HF Unverified

Tabular-RegressionSynthefy-NoriFeatures-TransformerTabularTabular-Foundation-ModelIn-Context-Learning MEDIUM
truthfulqa/truthful_qa HF Unverified

Dataset Card for truthful_qa Dataset Summary TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.

Task_categories:multiple-ChoiceTask_categories:text-GenerationTask_categories:question-AnsweringTask_ids:multiple-Choice-QaTask_ids:language-ModelingTask_ids:open-Domain-Qa
HuggingFaceM4/FineVision HF Unverified

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision Load the data from datasets import load_dataset, get_dataset_config_names # Get all subset names and load the first one available_subsets =… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/FineVision.

Size_categories:10M<n<100MFormat:parquetModality:imageModality:textLibrary:datasetsLibrary:dask
MMMU/MMMU HF Unverified

MMMU (A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI) 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 📖 arXiv | GitHub 🔔News 🛠️[2026-04-21]: Fixed option issue in test_Psychology_15. ‼️[2026-02-12]: We have released the answers for the test set! You can now evaluate your models on the test set locally! 🎉 🛠️[2024-05-30]: Fixed duplicate option issues in Materials dataset items (validation_Materials_25;… See the full description on the dataset page: https://huggingface.co/datasets/MMMU/MMMU.

Task_categories:question-AnsweringTask_categories:visual-Question-AnsweringTask_categories:multiple-ChoiceLanguage:enSize_categories:10K<n<100KFormat:parquet
P
philschmid/bart-large-cnn-samsum HF Unverified

SummarizationTransformersPyTorchBartText2text-GenerationSagemaker HIGH
M
MCG-NJU/videomae-base-finetuned-kinetics HF Unverified

Video-ClassificationTransformersPyTorchSafetensorsVideomaeVision MEDIUM
Showing 20 of 806 items (page 30 of 41)