RADAR//huggingface

A daily sweep of what's moving on the Hugging Face Hub — trending, top, and fast-rising models, datasets, and spaces. Every entry gets a one-line take, not just a like count.

last swept Aug 18, 2026 20:01 UTC 210 entries tracked in this sweep

All-time leaderboard by top downloads.

sentence-similarity NEW ♥ 5.2k⬇ 258.2m

sentence-transformers/all-MiniLM-L6-v2

Compact sentence embedding model for semantic similarity tasks across multiple frameworks.

sentence-transformerspytorchtfrustonnx
fill-mask NEW ♥ 2.7k⬇ 111.4m

google-bert/bert-base-uncased

Original uncased BERT base model for masked language modeling and fine-tuning tasks.

transformerspytorchtfjaxrust
text-ranking NEW ♥ 300⬇ 89.2m

cross-encoder/ms-marco-MiniLM-L6-v2

Cross-encoder trained on MS MARCO for passage reranking and relevance scoring.

sentence-transformerspytorchjaxonnxsafetensors
feature-extraction NEW ♥ 534⬇ 73.5m

BAAI/bge-small-en-v1.5

Small English embedding model from BAAI optimized for retrieval and feature extraction.

sentence-transformerspytorchonnxsafetensorsbert
NEW ♥ 154⬇ 56.1m

google/electra-base-discriminator

ELECTRA base discriminator for efficient pre-training via token replacement detection instead of masking.

transformerspytorchtfjaxrust
time-series-forecasting NEW ♥ 400⬇ 39.0m

amazon/chronos-2

T5-based foundation model for zero-shot time series forecasting across diverse domains and frequencies.

chronos-forecastingsafetensorst5time seriesforecasting
sentence-similarity NEW ♥ 3.4k⬇ 35.6m

BAAI/bge-m3

Multilingual embedding model supporting dense retrieval, sparse vectors, and multi-vector representations.

sentence-transformerspytorchonnxxlm-robertafeature-extraction
NEW ♥ 50⬇ 34.9m

lpiccinelli/unidepth-v2-vitl14

Vision Transformer for monocular metric depth estimation from single RGB images.

UniDepthpytorchsafetensorsmodel_hub_mixinmonocular-metric-depth-estimation
text-generation NEW ♥ 1.5k⬇ 28.8m

Qwen/Qwen3-0.6B

Compact 0.6B parameter Qwen3 language model for text generation and conversation.

transformerssafetensorsqwen3text-generationconversational
sentence-similarity NEW ♥ 1.3k⬇ 25.2m

sentence-transformers/all-mpnet-base-v2

General-purpose sentence embedding model based on MPNet, strong all-around semantic similarity performance.

sentence-transformerspytorchonnxsafetensorsopenvino
translation NEW ♥ 592⬇ 22.8m

google-t5/t5-small

Small T5 encoder-decoder for text-to-text tasks like translation, summarization, and QA.

transformerspytorchtfjaxrust
zero-shot-image-classification NEW ♥ 1.0k⬇ 20.8m

openai/clip-vit-base-patch32

Vision-language model aligning images and text for zero-shot image classification and retrieval.

transformerspytorchtfjaxclip
text-classification NEW ♥ 1.1k⬇ 18.8m

BAAI/bge-reranker-v2-m3

Multilingual cross-encoder for reranking search results with fine-grained relevance scoring.

sentence-transformerssafetensorsxlm-robertatext-classificationtransformers
image-classification NEW ♥ 101⬇ 18.5m

timm/mobilenetv3_small_100.lamb_in1k

Efficient MobileNetV3-small trained on ImageNet-1k for mobile image classification.

timmpytorchsafetensorsimage-classificationtransformers
fill-mask NEW ♥ 882⬇ 17.9m

FacebookAI/xlm-roberta-base

Cross-lingual RoBERTa trained on 100 languages for multilingual masked language modeling.

transformerspytorchtfjaxonnx
text-generation NEW ♥ 292⬇ 17.0m

facebook/opt-125m

Small 125M parameter OPT decoder for open-source language generation research.

transformerspytorchtfjaxopt
sentence-similarity NEW ♥ 892⬇ 16.4m

nomic-ai/nomic-embed-text-v1.5

Open embedding model with long context (8k tokens) and strong retrieval performance.

sentence-transformersonnxsafetensorsnomic_bertfeature-extraction
text-generation NEW ♥ 1.3k⬇ 15.8m

Qwen/Qwen3-8B

8B parameter Qwen3 language model for general text generation and conversational AI.

transformerssafetensorsqwen3text-generationconversational
NEW ♥ 1.4k⬇ 14.6m

Comfy-Org/MiniMax-H3

ComfyUI-compatible diffusion checkpoint fine-tuned from MiniMax H3 base model.

diffusion-single-filecomfyuibase_model:MiniMaxAI/MiniMax-H3base_model:finetune:MiniMaxAI/MiniMax-H3license:other
text-generation NEW ♥ 3.4k⬇ 13.9m

openai-community/gpt2

Original GPT-2 124M model for autoregressive text generation, widely used baseline.

transformerspytorchtfjaxtflite
image-text-to-text NEW ♥ 1.8k⬇ 13.8m

Qwen/Qwen3.5-9B

9B multimodal model handling both image and text inputs for visual question answering.

transformerssafetensorsqwen3_5image-text-to-textconversational
fill-mask NEW ♥ 638⬇ 13.1m

FacebookAI/roberta-base

RoBERTa base encoder optimized for masked language modeling and downstream NLP tasks.

transformerspytorchtfjaxrust
feature-extraction NEW ♥ 714⬇ 12.9m

BAAI/bge-large-en-v1.5

Large English embedding model from BAAI for high-quality semantic search and retrieval.

sentence-transformerspytorchonnxsafetensorsbert
sentence-similarity NEW ♥ 386⬇ 12.6m

intfloat/multilingual-e5-small

Compact multilingual E5 embeddings for cross-lingual sentence similarity and retrieval.

sentence-transformerspytorchonnxsafetensorsopenvino
text-to-speech NEW ♥ 6.7k⬇ 12.5m

hexgrad/Kokoro-82M

Lightweight 82M parameter text-to-speech model fine-tuned from StyleTTS2 for English voice synthesis.

text-to-speechenarxiv:2306.07691arxiv:2203.02395base_model:yl4579/StyleTTS2-LJSpeech
text-generation NEW ♥ 557⬇ 12.3m

nvidia/Qwen3.6-35B-A3B-NVFP4

NVIDIA-optimized Qwen 35B MoE quantized to FP4 for efficient inference on NVIDIA hardware.

Model Optimizersafetensorsqwen3_5_moenvidiaModelOpt
text-generation NEW ♥ 1.5k⬇ 12.3m

Qwen/Qwen2.5-7B-Instruct

Instruction-tuned 7B Qwen model optimized for chat and task-following interactions.

transformerssafetensorsqwen2text-generationchat
image-text-to-text NEW ♥ 355⬇ 12.2m

Qwen/Qwen3.6-35B-A3B-FP8

FP8-quantized 35B MoE multimodal model balancing quality and inference efficiency.

transformerssafetensorsqwen3_5_moeimage-text-to-textconversational

Already-popular models created in the last ~6 months — the fast risers.

image-text-to-text NEW ♥ 11.1k⬇ 665.5k

Qwen/Qwen3.8-27B

27B parameter vision-language model that accepts images and text for conversational tasks.

transformerssafetensorsqwen3_5image-text-to-textconversational
image-text-to-text NEW ♥ 10.8k⬇ 2.2m

moonshotai/Kimi-K3

Vision-language model with compressed tensors for efficient multimodal feature extraction.

transformerssafetensorskimi_k3feature-extractioncompressed-tensors
text-generation NEW ♥ 5.5k⬇ 1.2m

deepseek-ai/DeepSeek-V4-Pro

Large conversational text generation model from DeepSeek, V4 Pro variant with transformers support.

transformerssafetensorsdeepseek_v4text-generationconversational
text-generation NEW ♥ 5.0k⬇ 2.7m

zai-org/GLM-5.2

GLM-5.2 MoE model with DSA architecture for conversational text generation in English.

transformerssafetensorsglm_moe_dsatext-generationconversational
image-text-to-video NEW ♥ 4.1k⬇ 2.9m

MiniMaxAI/MiniMax-H3

Image and text to video generation model producing dynamic videos from multimodal prompts.

minimax-h3diffuserssafetensorstext-to-videoimage-to-video
image-text-to-text NEW ♥ 4.1k⬇ 3.2m

baidu/Unlimited-OCR

Baidu's vision-language model for OCR tasks, extracting text from images with feature embeddings.

transformerssafetensorsunlimited-ocrfeature-extractionbaidu
image-text-to-text NEW ♥ 3.6k⬇ 9.4m

google/gemma-4-31B-it

Google's 31B instruction-tuned Gemma 4 supporting multimodal image-text conversations.

transformerssafetensorsgemma4image-text-to-textconversational
text-generation NEW ♥ 3.5k⬇ 2.1m

deepseek-ai/DeepSeek-V4-Flash-0731

Faster inference variant of DeepSeek V4 from July 2024 for low-latency generation.

transformerssafetensorsdeepseek_v4text-generationconversational
image-text-to-text NEW ♥ 2.9k⬇ 108.4k

nvidia/LocateAnything-3B

NVIDIA's 3B vision model for object localization and spatial understanding tasks.

transformerssafetensorslocateanythingfeature-extractionnvidia
image-text-to-text NEW ♥ 2.7k⬇ 6.0m

Qwen/Qwen3.6-35B-A3B

Qwen 3.6 35B MoE model for multimodal image-text conversations, Apache 2.0 licensed.

transformerssafetensorsqwen3_5_moeimage-text-to-textconversational
image-text-to-text NEW ♥ 2.3k⬇ 6.8m

Qwen/Qwen3.6-27B

Qwen 3.6 27B model for multimodal image-text conversations, Apache 2.0 licensed.

transformerssafetensorsqwen3_5image-text-to-textconversational
text-generation NEW ♥ 2.1k⬇ 1.9m

deepseek-ai/DeepSeek-V4-Flash

DeepSeek V4 Flash variant optimized for faster conversational text generation inference.

transformerssafetensorsdeepseek_v4text-generationconversational
text-to-video NEW ♥ 2.0k⬇ 337.7k

SulphurAI/Sulphur-2-base

Text-to-video diffusion model based on quantized LTX-2.3, supports GGUF format.

diffuserssafetensorsgguftext-to-videobase_model:Lightricks/LTX-2.3
text-generation NEW ♥ 1.8k⬇ 78.1k

zai-org/GLM-5.1

GLM-5.1 MoE model with DSA architecture for conversational text generation tasks.

transformerssafetensorsglm_moe_dsatext-generationconversational

Top trending models within a few curated task areas.

text-generation NEW ♥ 1.1k⬇ 11.2k

Qwen/Qwen3.8-2.4T-A95B

Mixture-of-experts text generation model trained on 2.4T tokens with 95B active parameters.

transformerssafetensorsqwen3_5_moe_texttext-generationconversational
text-generation NEW ♥ 598⬇ 31.0k

deepseek-ai/DeepSeek-V4-Pro-0813

DeepSeek's V4 Pro checkpoint from August 2024 for conversational text generation tasks.

transformerssafetensorsdeepseek_v4text-generationconversational
text-generation NEW ♥ 3.5k⬇ 2.1m

deepseek-ai/DeepSeek-V4-Flash-0731

Faster inference variant of DeepSeek V4 from July 2024 for low-latency generation.

transformerssafetensorsdeepseek_v4text-generationconversational
text-generation NEW ♥ 224⬇ 13.3k

Qwen/Qwen3.8-2.4T-A95B-FP8

FP8-quantized MoE variant of the 2.4T token Qwen model with 95B active parameters.

transformerssafetensorsqwen3_5_moe_texttext-generationconversational
text-generation NEW ♥ 317⬇ 10.0k

inclusionAI/Ling-3.0-tiny

Tiny conversational LLM using custom hybrid architecture, MIT licensed for fast text generation.

safetensorsbailing_hybridtext-generationconversationalcustom_code
text-generation NEW ♥ 129⬇ 1.8k

empero-ai/Qwen3.8-9B

9B parameter Qwen multimodal model supporting image-text-to-text generation tasks.

transformerssafetensorsqwen3_5image-text-to-textempero-ai
text-to-image NEW ♥ 246⬇ 24.9k

Gazingstars123/Anima-2.9B

2.9B parameter text-to-image diffusion model compatible with ComfyUI workflows.

diffusion-single-fileanimacomfyuitext-to-imageen
text-to-image NEW ♥ 14.2k⬇ 585.9k

black-forest-labs/FLUX.1-dev

Development version of FLUX diffusion model for high-quality text-to-image generation.

diffuserssafetensorstext-to-imageimage-generationflux
text-to-image NEW ♥ 63⬇ 1.1k

realrebelai/Rebels_w4a8s

4-bit weight, 8-bit activation quantized image generation model optimized for ComfyUI.

w4a8quantizedint4comfyuicomfy-kitchen
text-to-image NEW ♥ 882⬇ 73.4k

krea/Krea-2-Turbo

Fast inference variant of Krea-2 image generator, fine-tuned for speed over quality.

diffuserssafetensorstext-to-imageenbase_model:krea/Krea-2-Raw
text-to-image NEW ♥ 497⬇ 85.9k

krea/Krea-2-Raw

Base Krea-2 diffusion model with custom pipeline for text-to-image synthesis.

diffuserssafetensorstext-to-imageenlicense:other
text-to-image NEW ♥ 280⬇ 0

lodestones/Kroma

Krea2-based text-to-image model packaged for ComfyUI workflows with custom licensing.

krea2kreatext-to-imagecomfyuilicense:other
text-to-image NEW ♥ 5.1k⬇ 863.5k

Tongyi-MAI/Z-Image-Turbo

Fast diffusion model from Tongyi research with academic paper backing (arXiv 2511.22699).

diffuserssafetensorstext-to-imageenarxiv:2511.22699
text-to-image NEW ♥ 5.6k⬇ 365.1k

black-forest-labs/FLUX.1-schnell

Schnell (fast) variant of FLUX diffusion optimized for rapid text-to-image generation.

diffuserssafetensorstext-to-imageimage-generationflux
text-to-image NEW ♥ 8.0k⬇ 1.6m

stabilityai/stable-diffusion-xl-base-1.0

Stable Diffusion XL base model for high-resolution text-to-image generation (1024px).

diffusersonnxsafetensorstext-to-imagestable-diffusion
automatic-speech-recognition NEW ♥ 3.1k⬇ 9.7m

pyannote/speaker-diarization-3.1

Speaker diarization pipeline identifying who spoke when in multi-speaker audio recordings.

pyannote-audiopyannotepyannote-audio-pipelineaudiovoice
automatic-speech-recognition NEW ♥ 1.0k⬇ 1.2m

nvidia/nemotron-3.5-asr-streaming-0.6b

Compact 0.6B streaming ASR model from NVIDIA for real-time speech transcription.

nemosafetensorsggufnemotron3_5_asrfeature-extraction
automatic-speech-recognition NEW ♥ 6.2k⬇ 4.9m

openai/whisper-large-v3

OpenAI's large Whisper v3 for robust multilingual speech recognition across 99 languages.

transformerspytorchjaxsafetensorswhisper
automatic-speech-recognition NEW ♥ 1.1k⬇ 5.4m

pyannote/speaker-diarization-community-1

Community-trained speaker diarization model for distinguishing speakers in audio streams.

pyannote-audiopyannotepyannote-audio-pipelineaudiovoice
automatic-speech-recognition NEW ♥ 42⬇ 2.3k

ARTPARK-IISc/SraVaani-1.0

Indic language ASR supporting Hindi, Kannada, and Malayalam with custom architecture.

sravaani_tdtautomatic-speech-recognitioncustom_codehikn
automatic-speech-recognition NEW ♥ 1.0k⬇ 4.4m

Qwen/Qwen3-ASR-1.7B

1.7B parameter speech recognition model from Qwen with published research backing.

safetensorsqwen3_asrautomatic-speech-recognitionarxiv:2601.21337license:apache-2.0
automatic-speech-recognition NEW ♥ 3.2k⬇ 8.0m

openai/whisper-large-v3-turbo

Faster turbo variant of Whisper v3 large for low-latency speech transcription.

transformerssafetensorswhisperautomatic-speech-recognitionaudio
automatic-speech-recognition NEW ♥ 611⬇ 138.1k

nvidia/nemotron-speech-streaming-en-0.6b

English-only streaming ASR from NVIDIA at 0.6B parameters for real-time applications.

nemosafetensorsggufnemotron_asr_streamingfeature-extraction
automatic-speech-recognition NEW ♥ 1.3k⬇ 696.5k

microsoft/VibeVoice-ASR

Microsoft ASR with integrated speaker diarization for transcription with speaker labels.

transformerssafetensorsvibevoiceASRTranscriptoin
automatic-speech-recognition NEW ♥ 1.5k⬇ 0

ggerganov/whisper.cpp

C++ implementation of Whisper for efficient on-device speech recognition inference.

automatic-speech-recognitionlicense:mitregion:us
text-to-speech NEW ♥ 134⬇ 4.8k

IndexTeam/IndexTTS-2.5

Zero-shot text-to-speech with voice cloning from minimal reference audio samples.

indexttssafetensorstext-to-speechttszero-shot
text-to-speech NEW ♥ 51⬇ 0

javawock7618/comfy-MiniMax-H3-workflows

ComfyUI workflows for MiniMax H3 reference-to-video and first-last-frame generation.

minimax-h3comfyuifirst-last-frameflfreference-to-video
text-to-speech NEW ♥ 1.3k⬇ 914.7k

k2-fsa/OmniVoice

Zero-shot multilingual TTS with voice cloning and design capabilities across languages.

omnivoicesafetensorszero-shotmultilingualvoice-cloning
text-to-speech NEW ♥ 6.7k⬇ 12.5m

hexgrad/Kokoro-82M

Lightweight 82M parameter text-to-speech model fine-tuned from StyleTTS2 for English voice synthesis.

text-to-speechenarxiv:2306.07691arxiv:2203.02395base_model:yl4579/StyleTTS2-LJSpeech
text-to-speech NEW ♥ 1.3k⬇ 512.3k

fishaudio/s2-pro

Instruction-following multilingual TTS model with Chinese language support.

safetensorsfish_qwen3_omnitext-to-speechinstruction-followingmultilingual
text-to-speech NEW ♥ 1.9k⬇ 2.3m

Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

1.7B parameter TTS model with 12kHz output supporting custom voice synthesis.

safetensorsqwen3_ttstext-to-speecharxiv:2601.15621license:apache-2.0
text-to-speech NEW ♥ 926⬇ 27.3k

Supertone/supertonic-3

ONNX-optimized multilingual speech synthesis model for deployment.

supertoniconnxtext-to-speechspeech-synthesistts
text-to-speech NEW ♥ 1.7k⬇ 2.1m

ResembleAI/chatterbox

Multilingual TTS with voice cloning for speech generation across languages.

chatterboxtext-to-speechspeechspeech-generationvoice-cloning
text-to-speech NEW ♥ 1.5k⬇ 430.2k

openbmb/VoxCPM2

Multilingual TTS with voice cloning capabilities for diverse language support.

voxcpmsafetensorstext-to-speechttsmultilingual
image-text-to-text NEW ♥ 11.1k⬇ 665.5k

Qwen/Qwen3.8-27B

27B parameter vision-language model that accepts images and text for conversational tasks.

transformerssafetensorsqwen3_5image-text-to-textconversational
image-text-to-text NEW ♥ 1.7k⬇ 384.1k

meta-models/Muse-Glimmer-30B

30B vision-language model from Meta for multimodal understanding and conversation.

transformerssafetensorsmuse_glimmerimage-text-to-textconversational
image-text-to-text NEW ♥ 557⬇ 741.0k

Qwen/Qwen3.8-27B-FP8

FP8-quantized Qwen3.8-27B for efficient vision-language inference with lower precision.

transformerssafetensorsqwen3_5image-text-to-textconversational
image-text-to-text NEW ♥ 513⬇ 45.5k

orcarouter/Qwen3.8-27B-Uncensored-FP8

Abliterated FP8 variant of Qwen3.8-27B with content restrictions removed for open responses.

transformerssafetensorsqwen3_5image-text-to-textabliterated
image-text-to-text NEW ♥ 10.8k⬇ 2.2m

moonshotai/Kimi-K3

Vision-language model with compressed tensors for efficient multimodal feature extraction.

transformerssafetensorskimi_k3feature-extractioncompressed-tensors
image-text-to-text NEW ♥ 219⬇ 1.1k

dots-studio/dots3-note-prev

Preview checkpoint from the Dots3-Note series for vision-language text generation.

transformerssafetensorsdots3_notetext-generationdots3
image-text-to-text NEW ♥ 209⬇ 0

orcarouter/Qwen3.8-27B-Uncensored-MLX

27B uncensored vision-language model optimized for Apple MLX framework.

mlxsafetensorsabliteratedqwen3.8qwen3_5
image-text-to-text NEW ♥ 480⬇ 787.3k

unsloth/Muse-Glimmer-30B-GGUF

30B vision-language model in GGUF format for image-text understanding tasks.

transformersggufunslothmetaimage-text-to-text

All-time leaderboard by top downloads.

NEW ♥ 75⬇ 2.6m

KakologArchives/KakologArchives

Japanese text classification dataset, likely from Kakolog (2ch/5ch archives)

task_categories:text-classificationlanguage:jalicense:mitregion:us
NEW ♥ 178⬇ 2.0m

huggingface/documentation-images

Official HuggingFace documentation images (non-commercial license)

license:cc-by-nc-sa-4.0size_categories:n<1Kformat:imagefoldermodality:imagelibrary:datasets
NEW ♥ 49⬇ 1.6m

ryanmarten/OpenThoughts-1k-sample

1K sample of reasoning traces or chain-of-thought examples for training or analysis

size_categories:1K<n<10Kformat:parquetmodality:textlibrary:datasetslibrary:pandas
NEW ♥ 759⬇ 1.5m

Salesforce/wikitext

Classic Wikipedia-based language modeling benchmark for pretraining and masked LM tasks

task_categories:text-generationtask_categories:fill-masktask_ids:language-modelingtask_ids:masked-language-modelingannotations_creators:no-annotation
NEW ♥ 32⬇ 1.5m

ayuo/hd_tmp

Temporary or private dataset with no public metadata

region:us
NEW ♥ 15⬇ 1.4m

IPEC-COMMUNITY/language_table_lerobot

Language-conditioned tabletop manipulation dataset in LeRobot format from RLDS

task_categories:roboticslicense:apache-2.0region:usLeRobotlanguage_table
NEW ♥ 630⬇ 1.4m

allenai/c4

Colossal Clean Crawled Corpus: massive web text for language model pretraining

task_categories:text-generationtask_categories:fill-masktask_ids:language-modelingtask_ids:masked-language-modelingannotations_creators:no-annotation
NEW ♥ 46⬇ 1.3m

xlangai/ubuntu_osworld_file_cache

Cached Ubuntu filesystem states for OS-World agent benchmark environments

license:apache-2.0arxiv:2404.07972region:us
NEW ♥ 52⬇ 1.3m

hf-doc-build/doc-build-dev

Internal HuggingFace documentation build artifacts and development data

license:mitregion:usdocumentation
NEW ♥ 167⬇ 1.1m

m-a-p/FineFineWeb

Billion-scale refined web corpus with quality annotations for improved pretraining

task_categories:text-classificationtask_categories:text-generationlanguage:enlicense:apache-2.0size_categories:1B<n<10B
NEW ♥ 243⬇ 1.1m

genrobot2025/10Kh-RealOmin-OpenData

Massive 10K+ hours of real-world omni-robot data for embodied AI (multi-language)

task_categories:roboticstask_categories:reinforcement-learninglanguage:enlanguage:zhlicense:cc-by-sa-4.0
NEW ♥ 1.6k⬇ 1.1m

openai/gsm8k

Official grade school math word problems benchmark, crowdsourced for reasoning evaluation

benchmark:officialbenchmark:eval-yamltask_categories:text-generationannotations_creators:crowdsourcedlanguage_creators:crowdsourced
NEW ♥ 0⬇ 990.4k

picbreeder-vlm/picbreeder-vlm-archive

100K+ machine-generated image captions from evolutionary image breeding experiments

task_categories:image-to-textannotations_creators:machine-generatedsource_datasets:originallanguage:enlicense:cc-by-nc-4.0
NEW ♥ 24⬇ 893.5k

Benjy/typed_digital_signatures

10K-100K images of typed digital signatures for classification and feature extraction tasks.

task_categories:image-classificationtask_categories:zero-shot-image-classificationtask_categories:image-feature-extractionlanguage:enlicense:mit
NEW ♥ 1⬇ 883.8k

happyhackingspace/dit

Dataset with minimal metadata; purpose unclear from available tags.

region:us
NEW ♥ 31⬇ 845.6k

anisoleai/fineweb-tokenized

Massive 1T+ token pre-tokenized text corpus from FineWeb for language model pretraining.

task_categories:text-generationlanguage:enlicense:odc-bysize_categories:n>1Tformat:parquet
NEW ♥ 0⬇ 822.2k

drssth/ModelNet-simscan

100B+ 3D point cloud classification dataset, likely synthetic scans from ModelNet objects.

license:ccsize_categories:100B<n<1Tregion:usPointCloud3D
NEW ♥ 35⬇ 753.1k

HennyPr/ps2_hf2

Small text dataset with under 1K samples; purpose unclear from metadata.

size_categories:n<1Kformat:textmodality:textlibrary:datasetslibrary:mlcroissant
NEW ♥ 24⬇ 717.4k

applied-ai-018/pretraining_v1-omega_books

100M-1B scale book corpus for language model pretraining, distributed via Dask.

size_categories:100M<n<1Bformat:parquetmodality:tabularmodality:textlibrary:datasets
NEW ♥ 13⬇ 709.8k

updatebao/geonamebase_1

Image dataset related to geographic names or locations.

modality:imageregion:us
NEW ♥ 24⬇ 656.3k

Dagonulca/figofigofigofigo

Small video dataset with fewer than 1K samples; purpose unclear.

size_categories:n<1Kmodality:videolibrary:datasetslibrary:mlcroissantregion:us
NEW ♥ 23⬇ 609.2k

jat-project/jat-dataset-tokenized

10M-100M timeseries samples pre-tokenized for JAT (likely Jack of All Trades) agent training.

size_categories:10M<n<100Mformat:parquetmodality:timeserieslibrary:datasetslibrary:dask
NEW ♥ 103⬇ 590.7k

XDOF/ABC-130k

1T+ scale robotics dataset (arXiv 2606.27375) with 130K scenarios for embodied AI.

task_categories:roboticslanguage:enlicense:apache-2.0size_categories:n>1Tarxiv:2606.27375
NEW ♥ 29⬇ 577.3k

Maximilians/ps2_hf1

Dataset with minimal metadata; purpose unclear from available tags.

region:us

Already-popular datasets created in the last ~6 months — the fast risers.

NEW ♥ 720⬇ 59.0k

Glint-Research/Fable-5-traces

1K-10K machine-generated agent reasoning traces for text generation training.

task_categories:text-generationannotations_creators:machine-generatedlanguage:enlicense:agpl-3.0size_categories:1K<n<10K
NEW ♥ 542⬇ 8.5k

nvidia/Nemotron-Personas-Korea

1M-10M Korean persona-based conversations for instruction tuning and dialogue training.

task_categories:text-generationlanguage:kolicense:cc-by-4.0size_categories:1M<n<10Mformat:parquet
NEW ♥ 502⬇ 18.8k

FlyRank/internship-warehouse

10-100M record tabular dataset of internship listings, likely for search or ranking tasks.

language:enlicense:othersize_categories:10M<n<100Mmodality:tabularmodality:text
NEW ♥ 440⬇ 2.1k

angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k

8.7K reasoning traces from Claude Opus 4.6-4.7 for training step-by-step inference.

task_categories:text-generationtask_categories:question-answeringlanguage:enlicense:apache-2.0size_categories:10K<n<100K
NEW ♥ 393⬇ 502

Roman1111111/claude-opus-4.6-10000x

1K-10K samples from Claude Opus 4.6, possibly augmented or filtered.

license:mitsize_categories:1K<n<10Kformat:jsonmodality:textlibrary:datasets
NEW ♥ 384⬇ 16.9k

openbmb/UltraData-SFT-2605

10M-100M bilingual (EN/ZH) instruction tuning samples for supervised fine-tuning.

task_categories:text-generationtask_categories:question-answeringlanguage:enlanguage:zhlicense:apache-2.0
NEW ♥ 380⬇ 4.9k

lambda/hermes-agent-reasoning-traces

10K-100K reasoning traces from Hermes agent for training chain-of-thought capabilities.

task_categories:text-generationlanguage:enlicense:apache-2.0size_categories:10K<n<100Kformat:parquet
NEW ♥ 369⬇ 1.9k

XYZAILab/XYZ-Aquila-SFT

1K-10K bilingual (EN/ZH) QA pairs for supervised fine-tuning of Aquila models.

task_categories:text-generationtask_categories:question-answeringlanguage:enlanguage:zhlicense:apache-2.0
NEW ♥ 359⬇ 4.3k

armand0e/claude-fable-5-claude-code

Small collection of Claude-generated agent traces with code, formatted for training.

task_categories:text-generationsize_categories:n<1Kformat:jsonformat:agent-tracesmodality:tabular
NEW ♥ 348⬇ 257.6k

HuggingFaceCode/stack-v3-train

Stack v3 multilingual code corpus for training code generation models, ODC-BY licensed.

task_categories:text-generationlanguage_creators:crowdsourcedlanguage_creators:expert-generatedmultilinguality:multilinguallanguage:code
NEW ♥ 346⬇ 4.1k

stepfun-ai/Step-3.5-Flash-SFT

1M-10M multilingual instruction tuning samples for Step 3.5 Flash model training.

task_categories:text-generationlanguage:multilinguallicense:apache-2.0license:cc-by-nc-2.0size_categories:1M<n<10M
NEW ♥ 341⬇ 30.4k

open-index/hacker-news

Hacker News posts and comments for text generation, classification, and retrieval tasks.

task_categories:text-generationtask_categories:feature-extractiontask_categories:text-classificationtask_categories:question-answeringlanguage:en

All-time leaderboard by top likes.

docker NEW ♥ 16.6k

enzostvs/deepsite

AI-powered website generation or analysis tool

dockerregion:us
docker NEW ♥ 14.1k

open-llm-leaderboard/open_llm_leaderboard

Community benchmark ranking open-source language models on standardized English text tasks

dockerleaderboardmodality:textsubmission:automatictest:public
docker NEW ♥ 7.6k

mteb/leaderboard

Massive Text Embedding Benchmark leaderboard for comparing embedding models

dockerleaderboardregion:us
static NEW ♥ 5.7k

dalle-mini/dalle-mini

Text-to-image generation demo, predecessor to DALL-E 2, generates images from prompts.

staticregion:us
gradio NEW ♥ 5.4k

AP123/IllusionDiffusion

Generates optical illusion images using diffusion models with hidden patterns.

gradioregion:us
gradio NEW ♥ 5.1k

facebook/MusicGen

Meta's music generation model that creates audio from text descriptions.

gradiomusic generationlanguage modelsLLMsregion:us
static NEW ♥ 5.0k

lmarena-ai/arena-leaderboard

Crowdsourced leaderboard ranking language models by human preference voting.

staticleaderboardregion:us
gradio NEW ♥ 4.8k

microsoft/TRELLIS

Microsoft research demo, likely for 3D or multimodal generation tasks.

gradioregion:us
gradio NEW ♥ 4.3k

KingNish/OpenGPT-4o

Open alternative attempting GPT-4-like multimodal conversation capabilities.

gradioregion:us
static NEW ♥ 4.0k

nanotron/ultrascale-playbook

Documentation or guide for training models at extreme scale with Nanotron framework.

staticregion:us
gradio NEW ♥ 3.8k

KlingTeam/LivePortrait

Animates still portraits with motion control, turning images into video.

gradioMultimodalMotion controlImage-to-VideoVideo-to-Video
gradio NEW ♥ 3.7k

mrfakename/Z-Image-Turbo

Fast image generation or manipulation tool with MCP server integration.

gradiomcp-serverregion:us
gradio NEW ♥ 3.6k

InstantX/InstantID

Identity-preserving image generation, swaps faces while maintaining likeness.

gradioregion:us
gradio NEW ♥ 3.4k

hexgrad/Kokoro-TTS

Text-to-speech synthesis model generating natural voice audio from text.

gradioregion:us
gradio NEW ♥ 3.4k

tencent/Hunyuan3D-2

Tencent's 3D asset generation model, creates 3D objects from text or images.

gradioregion:us
docker NEW ♥ 3.3k

akhaliq/anycoder

Code generation or programming assistance tool, containerized for isolation.

dockerregion:us
docker NEW ♥ 3.3k

HuggingFaceTB/smol-training-playbook

Research guide for training small language models efficiently with visualizations.

dockerresearch-article-templateresearch paperscientific paperdata visualization
gradio NEW ♥ 3.0k

r3gm/wan2-2-fp8da-aoti-preview

Preview of WAN2 model with FP8 quantization and ahead-of-time compilation.

gradiomcp-serverregion:us
gradio NEW ♥ 3.0k

mukaist/Midjourney

Unofficial Midjourney-style text-to-image generation interface or alternative.

gradioregion:us
gradio NEW ♥ 2.9k

not-lain/background-removal

Removes backgrounds from images automatically, outputs transparent PNGs.

gradiomcp-serverregion:us
gradio NEW ♥ 2.9k

mrfakename/E2-F5-TTS

End-to-end text-to-speech with F5 model architecture for voice synthesis.

gradioregion:us
docker NEW ♥ 2.8k

sanchit-gandhi/whisper-jax

JAX-optimized Whisper speech recognition for faster transcription inference.

dockerregion:us
gradio NEW ♥ 2.8k

openai/whisper

OpenAI's speech recognition model transcribing audio to text across languages.

gradioregion:us
gradio NEW ♥ 2.8k

HumanAIGC/OutfitAnyone

Virtual try-on system that swaps clothing on people in images realistically.

gradioregion:us

Already-popular spaces created in the last ~6 months — the fast risers.

gradio NEW ♥ 1.2k

kulkas2pintu/wan555

Gradio app with MCP server integration for model inference

gradiomcp-serverregion:us
gradio NEW ♥ 1.2k

k2-fsa/OmniVoice

Multi-language voice synthesis or speech processing from the K2 project.

gradioregion:us