Datasets

Training datasets with quantum-safe provenance

DWKPartners/trading-data HF Unverified

DartLab Data Structured company data from DART & EDGAR disclosure filings DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사 What is this? Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map. 한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다. This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/DWKPartners/trading-data.

Task_categories:table-Question-AnsweringTask_categories:text-ClassificationLanguage:koLanguage:enSize_categories:1K<n<10KFinance
muset-ai/DeepResearch-Bench-II-Dataset HF Unverified

Task_categories:text-GenerationTask_categories:text-ClassificationLanguage:zhLanguage:enSize_categories:n<1KModality:document
DWKPartners/dartlab-data HF Unverified

DartLab Data Structured company data from DART & EDGAR disclosure filings DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사 What is this? Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map. 한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다. This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/DWKPartners/dartlab-data.

Task_categories:table-Question-AnsweringTask_categories:text-ClassificationLanguage:koLanguage:enSize_categories:1K<n<10KFinance
RevolutionCrossroads/loc_chronicling_america_1770-1810 HF Unverified

Dataset Card for Chronicling America: Historic American Newspapers 1770–1810 Dataset Summary A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset includes page-level records with images, original Chronicling America OCR, AI-generated OCR, and publication metadata for newspapers published between 1770 and 1810. It provides a foundation for research, machine learning… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810.

Task_categories:image-To-TextTask_categories:text-RetrievalTask_categories:text-ClassificationTask_categories:feature-ExtractionLanguage:enSize_categories:10K<n<100K
echodict/KakologArchives_duplicate HF Unverified

ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/echodict/KakologArchives_duplicate.

Task_categories:text-ClassificationLanguage:ja
AlgorithmicResearchGroup/s2orc_full HF Unverified

S2ORC Full — Semantic Scholar Open Research Corpus A complete redistribution of the S2ORC dataset in Parquet format on Hugging Face, containing 14.5 million academic papers with full text, structured metadata, and citation information. Dataset Description S2ORC (Semantic Scholar Open Research Corpus) is a general-purpose corpus for NLP and text mining research over scientific papers, originally developed by the Allen Institute for AI. This version provides the full… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc_full.

Task_categories:text-GenerationTask_categories:feature-ExtractionTask_categories:text-ClassificationLanguage:enSize_categories:10M<n<100MModality:text
cornell-movie-review-data/rotten_tomatoes HF Unverified

Dataset Card for "rotten_tomatoes" Dataset Summary Movie Review Dataset. This is a dataset of containing 5,331 positive and 5,331 negative processed sentences from Rotten Tomatoes movie reviews. This data was first used in Bo Pang and Lillian Lee, ``Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.'', Proceedings of the ACL, 2005. Supported Tasks and Leaderboards More Information Needed Languages… See the full description on the dataset page: https://huggingface.co/datasets/cornell-movie-review-data/rotten_tomatoes.

Task_categories:text-ClassificationTask_ids:sentiment-ClassificationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingualSource_datasets:original
Voxel51/Office-Home HF Unverified

Dataset Card for Office-Home This is a FiftyOne dataset with 15588 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Office-Home") # Launch the App session = fo.launch_app(dataset) Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Office-Home.

Task_categories:image-ClassificationLanguage:enSize_categories:10K<n<100KFormat:imagefolderModality:imageLibrary:datasets
Horama/animal-200 HF Unverified

Horama/animal-200 Raw wildlife image collection covering 199 species (mammals, birds, reptiles), scraped from multiple web sources. Images are organized by species folder and can be used as-is for image classification (species identification) or as input for downstream annotation pipelines (object detection, etc.). For animal detection, see Horama/animal-200-detection dataset. Sources Images were collected from three web sources using dedicated scrapers: Source… See the full description on the dataset page: https://huggingface.co/datasets/Horama/animal-200.

Task_categories:image-ClassificationTask_categories:object-DetectionLanguage:enLanguage:frSize_categories:10K<n<100KModality:image
zeroMN/hanlp_date-zh HF Unverified

-- 2nd International Chinese Word Segmentation Bakeoff - Data Release Release 1, 2005-11-18 Introduction This directory contains the training, test, and gold-standard data used in the 2nd International Chinese Word Segmentation Bakeoff. Also included is the script used to score the results submitted by the bakeoff participants and the simple segmenter used to generate the baseline and topline data. File List gold/ Contains the gold standard… See the full description on the dataset page: https://huggingface.co/datasets/zeroMN/hanlp_date-zh.

Task_categories:text-ClassificationLanguage:zhSize_categories:100M<n<1BCode
a2015003713/military-aircraft-detection-dataset HF Unverified

Military Aircraft Detection Dataset Military aircraft detection dataset in COCO and YOLO format. This dataset is synchronized from the original Kaggle dataset:https://www.kaggle.com/datasets/a2015003713/militaryaircraftdetectiondataset

Task_categories:object-DetectionTask_categories:image-ClassificationTask_categories:image-Feature-ExtractionSize_categories:10K<n<100KFormat:textModality:image
Forithmus/MR-RATE-coreg HF Unverified

MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging This is the MR-RATE-coreg repository, part of the MR-RATE dataset release. It contains co-registered MRI volumes in which all imaging volumes within each study have been spatially aligned to a common T1-weighted reference frame. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-coreg.

Task_categories:image-To-TextTask_categories:text-To-ImageTask_categories:image-ClassificationTask_categories:question-AnsweringTask_categories:visual-Question-AnsweringTask_categories:zero-Shot-Classification
saffatgazi/fish-vista HF Unverified

Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images. See Example Code to Use the Segmentation Dataset Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset. Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/saffatgazi/fish-vista.

Task_categories:image-ClassificationTask_categories:image-SegmentationLanguage:enSize_categories:10K<n<100KFormat:csvModality:image
imageomics/mmla_mpala HF Unverified

Dataset Card for mmla-mpala Dataset Details This is a dataset containing annotated video frames of giraffes, Grevy's zebras, and Plains zebras collected at the Mpala Research Center in Kenya. The dataset is intended for use in training and evaluating computer vision models for animal detection and classification from drone imagery. The annotations indicate the presence of animals in the images in YOLO format. The dataset is designed to facilitate research in wildlife… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/mmla_mpala.

Task_categories:image-ClassificationLanguage:enSize_categories:10K<n<100KModality:imageModality:textDoi:10.57967/hf/7377
zai-org/LongBench HF Unverified

LongBench is a comprehensive benchmark for multilingual and multi-task purposes, with the goal to fully measure and evaluate the ability of pre-trained language models to understand long text. This dataset consists of twenty different tasks, covering key long-text application scenarios such as multi-document QA, single-document QA, summarization, few-shot learning, synthetic tasks, and code completion.

Task_categories:question-AnsweringTask_categories:text-GenerationTask_categories:summarizationTask_categories:text-ClassificationLanguage:enLanguage:zh
FLARE-MedFM/FLARE26-MLLM-3D HF Unverified

FLARE 2026: Multimodal Model for 3D Medical Image Parsing The task is to train one multimodal model for report generation and vision QA. Data Description The dataset contains two subsets for abdomen and lung CT report generation and VQA. FLARE-Task5-MLLM-3D/ ├── README.md ├── train # training set │ ├── CT-AMOS-1290 # source: https://era-ai-biomed.github.io/amos/ │ ├── CT-AMOS-Tr.json │ ├── CT-RATE-2000 # source:… See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE26-MLLM-3D.

Task_categories:image-ClassificationLanguage:enMedical
vsevolodpl/REPID HF Unverified

REPID: Rendering Evaluation of Photographic Image Dataset REPID (officially introduced as the Rendering Evaluation of Photographic Image Dataset) is a large-scale benchmark designed for Image Rendering Quality Assessment (IRQA). Unlike traditional Image Quality Assessment (IQA) which focuses on technical degradations like noise or blur, REPID aims to model subjective human aesthetic preferences for different rendering styles of the same scene. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/vsevolodpl/REPID.

Task_categories:image-ClassificationTask_categories:image-To-ImageTask_categories:otherSize_categories:10K<n<100KFormat:csvModality:image
ayayon/papas-nativas-peru-83-variedades HF Unverified

Papas Nativas Peruanas — 83 Variedades (UNSAAC 2024) Colección de imágenes de 83 variedades de papas nativas peruanas para clasificación visual mediante modelos de visión computacional. Descripción del dataset Dataset recopilado de forma colaborativa por estudiantes de Ingeniería Informática de la Universidad Nacional de San Antonio Abad del Cusco (UNSAAC) en el curso de Aprendizaje Automático (2024). Las imágenes fueron capturadas en condiciones variadas (distintos… See the full description on the dataset page: https://huggingface.co/datasets/ayayon/papas-nativas-peru-83-variedades.

Task_categories:image-ClassificationLanguage:esSize_categories:10K<n<100KFormat:imagefolderModality:imageLibrary:datasets
Forithmus/MR-RATE-atlas HF Unverified

MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging This is the MR-RATE-atlas repository, part of the MR-RATE dataset release. It contains atlas-registered MRI volumes in which all imaging sequences within each study have been spatially normalized to a standard atlas-space. For full dataset details, native-space MRI volumes, radiology reports, metadata, and data splits, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/Forithmus/MR-RATE-atlas.

Task_categories:image-To-TextTask_categories:text-To-ImageTask_categories:image-ClassificationTask_categories:question-AnsweringTask_categories:visual-Question-AnsweringTask_categories:zero-Shot-Classification
ufldl-stanford/svhn HF Unverified

Dataset Card for Street View House Numbers Dataset Summary SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and formatting. It can be seen as similar in flavor to MNIST (e.g., the images are of small cropped digits), but incorporates an order of magnitude more labeled data (over 600,000 digit images) and comes from a significantly harder, unsolved, real world problem… See the full description on the dataset page: https://huggingface.co/datasets/ufldl-stanford/svhn.

Task_categories:image-ClassificationTask_categories:object-DetectionAnnotations_creators:machine-GeneratedAnnotations_creators:expert-GeneratedLanguage_creators:machine-GeneratedMultilinguality:monolingual
Showing 20 of 277 datasets (page 12 of 14)