Model Hub
Browse PQC-verified AI models, datasets, and tools
S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.
Dataset Card for "rotten_tomatoes" Dataset Summary Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.
Dataset Card for Office-Home This is a FiftyOne dataset with 15588 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Office-Home") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Office-Home.
Horama/animal-200 Raw wildlife image collection covering 199 species (mammals, birds, reptiles), scraped from multiple web sources. Images are organized by species folder and can be used as-is for image classification (species identification) or as input for downstream annotation pipelines (object detection, etc.). For animal detection, see Horama/animal-200-detection dataset. Sources Images were collected from three web sources using dedicated scrapers: Source… See the full description on the dataset page: https://huggingface.co/datasets/Horama/animal-200.
-- 2nd International Chinese Word Segmentation Bakeoff - Data Release Release 1, 2005-11-18 Introduction This directory contains the training, test, and gold-standard data used in the 2nd International Chinese Word Segmentation Bakeoff. Also included is the script used to score the results submitted by the bakeoff participants and the simple segmenter used to generate the baseline and topline data. File List gold/ Contains the gold standard… See the full description on the dataset page: https://huggingface.co/datasets/zeroMN/hanlp_date-zh.
Military Aircraft Detection Dataset Military aircraft detection dataset in COCO and YOLO format. This dataset is synchronized from the original Kaggle dataset:https://www.kaggle.com/datasets/a2015003713/militaryaircraftdetectiondataset
MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging This is the MR-RATE-coreg repository, part of the MR-RATE dataset release. It contains co-registered MRI volumes in which all imaging volumes within each study have been spatially aligned to a common T1-weighted reference frame. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-coreg.
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/saffatgazi/fish-vista.
Dataset Card for mmla-mpala Dataset Details This is a dataset containing annotated video frames of giraffes, Grevy's zebras, and Plains zebras collected at the Mpala Research Center in Kenya. The dataset is intended for use in training and evaluating computer vision models for animal detection and classification from drone imagery. The annotations indicate the presence of animals in the images in YOLO format. The dataset is designed to facilitate research in wildlife… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/mmla_mpala.
LongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.
FLARE 2026: Multimodal Model for 3D Medical Image Parsing The task is to train one multimodal model for report generation and vision QA. Data Description The dataset contains two subsets for abdomen and lung CT report generation and VQA. FLARE-Task5-MLLM-3D/ ├── README.md ├── train # training set │ ├── CT-AMOS-1290 # source: https://era-ai-biomed.github.io/amos/ │ ├── CT-AMOS-Tr.json │ ├── CT-RATE-2000 # source:… See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE26-MLLM-3D.
MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging This is the MR-RATE-atlas repository, part of the MR-RATE dataset release. It contains atlas-registered MRI volumes in which all imaging sequences within each study have been spatially normalized to a standard atlas-space. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-atlas.
REPID: Rendering Evaluation of Photographic Image Dataset REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA). Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic preferences for different rendering styles of the same scene. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.
Papas Nativas Peruanas — 83 Variedades (UNSAAC 2024) Colección de imágenes de 83 variedades de papas nativas peruanas para clasificación visual mediante modelos de visión computacional. Descripción del dataset Dataset recopilado de forma colaborativa por estudiantes de Ingeniería Informática de la Universidad Nacional de San Antonio Abad del Cusco (UNSAAC) en el curso de Aprendizaje Automático (2024). Las imágenes fueron capturadas en condiciones variadas (distintos… See the full description on the dataset page: https://huggingface.co/datasets/ayayon/papas-nativas-peru-83-variedades.