Model Hub

Browse PQC-verified AI models, datasets, and tools

ksolovev/fine-news HF Unverified

FineNews FineNews is a multilingual news-text dataset for language-model research. It filters and deduplicates a 2021–2025 snapshot of the INFINI-NEWS Corpus, which extracts articles from Common Crawl CC-News. At a glance Measure FineNews Input articles 852,824,802 Infini-News rows from 2021–2025 Output 392,627,654 physical rows Files 294,509 Parquet files Folders 60 publication months (2021-01 to 2025-12), then language Language folders 129… See the full description on the dataset page: https://huggingface.co/datasets/ksolovev/fine-news.

Task_categories:text-GenerationSize_categories:100M<n<1BNewsJournalismMediaCommon-Crawl
shash42/forecast-news HF Unverified

Forecast News Deduplicated daily news corpus used by forecast-sim and future-sim. Snapshot 31,696,086 articles 3,401 daily partitions Coverage: 2016-08-26 through 2026-06-30 Snapshot published: 2026-07-13 Stored data size: approximately 168.8 GB The Parquet files are the canonical complete representation. The repository also contains daily JSONL files where available and compact headline JSON files used by article-browsing workflows. Layout Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.

Size_categories:10M<n<100MFormat:parquetModality:textLibrary:datasetsLibrary:daskLibrary:polars
fancyzhx/ag_news HF Unverified

Dataset Card for "ag_news" Dataset Summary AG is a collection of more than 1 million news articles. News articles have been gathered from more than 2000 news sources by ComeToMyHead in more than 1 year of activity. ComeToMyHead is an academic news search engine which has been running since July, 2004. The dataset is provided by the academic comunity for research purposes in data mining (clustering, classification, etc), information retrieval (ranking, search, etc), xml… See the full description on the dataset page: https://huggingface.co/datasets/fancyzhx/ag_news.

Task_categories:text-ClassificationTask_ids:topic-ClassificationAnnotations_creators:foundLanguage_creators:foundMultilinguality:monolingualSource_datasets:original
ruggsea/infini-news-corpus HF Unverified

INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.

Task_categories:text-GenerationTask_categories:text-ClassificationTask_categories:text-RetrievalAnnotations_creators:machine-GeneratedMultilinguality:multilingualSource_datasets:original
Showing 4 of 4 items (page 1 of 1)
Prev Next