Model Hub

Browse PQC-verified AI models, datasets, and tools

T
theainerd/Wav2Vec2-large-xlsr-hindi HF PQC Verified

Speech RecognitionTransformersPyTorchSafetensorsWav2vec2Base_model:facebook/wav2vec2-Large-Xlsr-53 HIGH
S
Salesforce/SFR-Embedding-2_R HF PQC Verified

State-of-the-art text embedding model. Top of MTEB leaderboard with strong retrieval and clustering.

TransformerEmbeddings7BRetrieval HIGH
fineinstructions/fineinstructions_nemotron HF Unverified

✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline. The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details. Each .parquet file in the data folderhas a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.

Language:enSize_categories:1B<n<10BFormat:parquetModality:tabularModality:textLibrary:datasets
L
lucas-leme/FinBERT-PT-BR HF PQC Verified

Text ClassificationTransformersPyTorchBERTPt MEDIUM
m-a-p/FineFineWeb HF PQC Verified

FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.

Task_categories:text-ClassificationTask_categories:text-GenerationLanguage:enSize_categories:1B<n<10BModality:tabularModality:text
S
Systran/faster-whisper-tiny.en HF PQC Verified

Speech RecognitionCtranslate2AudioEnglish MEDIUM
H
Helsinki-NLP/opus-mt-nl-en HF Unverified

TranslationTransformersPyTorchTfRustMarian HIGH
F
facebook/mms-1b-all HF PQC Verified

Speech RecognitionTransformersPyTorchSafetensorsWav2vec2Mms HIGH
N
nlptown/bert-base-multilingual-uncased-sentiment HF PQC Verified

Text ClassificationTransformersPyTorchTfJAXSafetensors HIGH
C
CIDAS/clipseg-rd64-refined HF PQC Verified

Image-SegmentationTransformersPyTorchSafetensorsClipsegVision HIGH
K
kresnik/wav2vec2-large-xlsr-korean HF PQC Verified

Speech RecognitionTransformersPyTorchSafetensorsWav2vec2Speech HIGH
A
audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim HF PQC Verified

Audio-ClassificationTransformersPyTorchSafetensorsWav2vec2Speech HIGH
J
jonatasgrosman/wav2vec2-large-xlsr-53-chinese-zh-cn HF PQC Verified

Speech RecognitionTransformersPyTorchJAXWav2vec2Audio HIGH
openai/gsm8k HF PQC Verified

Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.

Benchmark:officialBenchmark:eval-YamlTask_categories:text-GenerationAnnotations_creators:crowdsourcedLanguage_creators:crowdsourcedMultilinguality:monolingual
genrobot2025/10Kh-RealOmin-OpenData HF Unverified

Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.

Task_categories:roboticsTask_categories:reinforcement-LearningLanguage:enLanguage:zhSize_categories:n>1TModality:video
S
Systran/faster-whisper-tiny HF Unverified

Speech RecognitionCtranslate2AudioHawJw MEDIUM
M
microsoft/table-transformer-detection HF PQC Verified

Object-DetectionTransformersPyTorchSafetensorsTable-Transformer MEDIUM
T
Tongyi-MAI/Z-Image-Turbo HF PQC Verified

Text-to-ImageDiffusersSafetensorsDiffusers:ZImagePipelineEnglish CRITICAL
A
amunchet/rorshark-vit-base HF Unverified

Image-ClassificationTransformersTensorboardSafetensorsVitVision MEDIUM
N
NousResearch/Hermes-3-Llama-3.1-8B HF Ollama PQC Verified

Nous Research's Hermes 3 built on Llama 3.1. Strong function calling, structured output, and agentic capabilities.

TransformerText Generation8BInstructAgentic HIGH
Showing 20 of 801 items (page 14 of 41)