Model Hub

Browse PQC-verified AI models, datasets, and tools

T
trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 HF PQC Verified

Text GenerationTransformersSafetensorsQwen2TrlConversational HIGH
L
laion/clap-htsat-fused HF PQC Verified

Audio-ClassificationTransformersPyTorchSafetensorsClapFeature Extraction HIGH
P
ProsusAI/finbert HF PQC Verified

Text ClassificationTransformersPyTorchTfJAXBERT HIGH
H
hmellor/tiny-random-LlamaForCausalLM HF PQC Verified

Text GenerationTransformersSafetensorsLLaMAConversationalText Generation HIGH
F
Falconsai/nsfw_image_detection HF PQC Verified

Image-ClassificationTransformersPyTorchSafetensorsVit HIGH
F
facebook/sam3 HF PQC Verified

Mask-GenerationTransformersSafetensorsSam3_videoFeature ExtractionSam3 HIGH
Salesforce/wikitext HF PQC Verified

Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.

Task_categories:text-GenerationTask_categories:fill-MaskTask_ids:language-ModelingTask_ids:masked-Language-ModelingAnnotations_creators:no-AnnotationLanguage_creators:crowdsourced
OpenAssistant/OpenAssistant Conversations v2 HF PQC Verified

Human-generated, human-annotated conversation trees. 91K messages across 35+ languages. RLHF training data.

DatasetConversationRLHF91K messages CRITICAL
T
TrustSafeAI/RADAR-Vicuna-7B HF PQC Verified

Text ClassificationTransformersPyTorchRoberta HIGH
K
k2-fsa/OmniVoice HF PQC Verified

Text-To-SpeechOmnivoiceSafetensorsZero-ShotMultilingualVoice-Cloning HIGH
S
Salesforce/SFR-Embedding-2_R HF PQC Verified

State-of-the-art text embedding model. Top of MTEB leaderboard with strong retrieval and clustering.

TransformerEmbeddings7BRetrieval HIGH
P
PaddlePaddle/PP-DocLayoutV3_safetensors HF PQC Verified

Object-DetectionTransformersSafetensorsPp_doclayout_v3PaddleOCRPaddlePaddle MEDIUM
S
segment-any-text/sat-3l-sm HF Unverified

Token ClassificationTransformersONNXSafetensorsXlm-TokenMultilingual HIGH
F
FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree HF Unverified

Text-To-VideoDiffusersSafetensorsVideoAudioText-To-Audio-Video CRITICAL
nvidia/SAGE-10k HF Unverified

SAGE-10k SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects. 🔑 Key Features SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.

Task_categories:text-To-3dLanguage:enSize_categories:10K<n<100KScene-GenerationInteractive-ScenesEmbodied-AI
K
krpraveen/sanskrit-en-custom-transformer HF Unverified

TranslationPyTorchSanskritNmtFrom-ScratchSa MEDIUM
Salesforce/GiftEvalPretrain HF Unverified

GIFT-Eval Pre-training Datasets Pretraining dataset aligned with GIFT-Eval that has 71 univariate and 17 multivariate datasets, spanning seven domains and 13 frequencies, totaling 4.5 million time series and 230 billion data points. Notably this collection of data has no leakage issue with the train/test split and can be used to pretrain foundation models that can be fairly evaluated on GIFT-Eval. 📄 Paper 🖥️ Code 📔 Blog Post 🏎️ Leader Board Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/GiftEvalPretrain.

Task_categories:time-Series-ForecastingSize_categories:1M<n<10MModality:timeseriesTimeseriesForecastingBenchmark
isaacus/open-australian-legal-corpus HF Unverified

Open Australian Legal Corpus ‍⚖️ The Open Australian Legal Corpus by Isaacus, a foundational legal AI research company, is the first and only multijurisdictional open corpus of Australian legislative and judicial documents. Comprised of 229,122 texts totalling over 60 million lines and 1.4 billion tokens, the Corpus includes every in force statute and regulation in the Commonwealth, New South Wales, Queensland, Western Australia, South Australia, Tasmania and Norfolk Island, in… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/open-australian-legal-corpus.

Task_categories:text-GenerationTask_categories:fill-MaskTask_categories:text-RetrievalTask_ids:language-ModelingTask_ids:masked-Language-ModelingTask_ids:document-Retrieval
labofsahil/pypi-packages-metadata-dataset HF Unverified

Size_categories:10M<n<100MModality:text
sarulab-speech/yodas2_sidon HF Unverified

YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling. YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.

Task_categories:text-To-SpeechTask_categories:automatic-Speech-RecognitionLanguage:aaLanguage:abLanguage:afLanguage:ak
Showing 20 of 29 items (page 1 of 2)
Prev Next