Qwen/Qwen3.8-27B
27B parameter vision-language model that accepts images and text for conversational tasks.
A daily sweep of what's moving on the Hugging Face Hub — trending, top, and fast-rising models, datasets, and spaces. Every entry gets a one-line take, not just a like count.
Trending models right now, by Hub trending score.
27B parameter vision-language model that accepts images and text for conversational tasks.
GGUF-quantized version of Qwen3.8-27B optimized for faster inference with reduced memory.
Mixture-of-experts text generation model trained on 2.4T tokens with 95B active parameters.
Text-to-music diffusion model that generates audio from natural language descriptions.
Multimodal video generation model supporting text, image, video, and audio as conditioning inputs.
DeepSeek's V4 Pro checkpoint from August 2024 for conversational text generation tasks.
30B vision-language model from Meta for multimodal understanding and conversation.
FP8-quantized Qwen3.8-27B for efficient vision-language inference with lower precision.
Image and text to video generation model producing dynamic videos from multimodal prompts.
Abliterated FP8 variant of Qwen3.8-27B with content restrictions removed for open responses.
GGUF-quantized uncensored Qwen3.8 optimized for llama.cpp with speculative decoding support.
Faster inference variant of DeepSeek V4 from July 2024 for low-latency generation.
Accelerated image-to-video diffusion model supporting text, image, and region-based conditioning.
NVIDIA FP4-quantized Qwen3.8-27B for extreme compression and fast inference on compatible hardware.
2.9B parameter text-to-image diffusion model compatible with ComfyUI workflows.
Vision-language model with compressed tensors for efficient multimodal feature extraction.
Preview checkpoint from the Dots3-Note series for vision-language text generation.
NVIDIA's 30B Nemotron Lightning model in FP4 precision for ultra-fast text generation.
FP8-quantized MoE variant of the 2.4T token Qwen model with 95B active parameters.
Heavily fine-tuned uncensored GGUF variant of Qwen for creative and unrestricted generation.
All-time leaderboard by top downloads.
Compact sentence embedding model for semantic similarity tasks across multiple frameworks.
Original uncased BERT base model for masked language modeling and fine-tuning tasks.
Cross-encoder trained on MS MARCO for passage reranking and relevance scoring.
Small English embedding model from BAAI optimized for retrieval and feature extraction.
Multilingual sentence embeddings for cross-lingual semantic similarity in 50+ languages.
ELECTRA base discriminator for efficient pre-training via token replacement detection instead of masking.
T5-based foundation model for zero-shot time series forecasting across diverse domains and frequencies.
Multilingual embedding model supporting dense retrieval, sparse vectors, and multi-vector representations.
Vision Transformer for monocular metric depth estimation from single RGB images.
Compact 0.6B parameter Qwen3 language model for text generation and conversation.
General-purpose sentence embedding model based on MPNet, strong all-around semantic similarity performance.
Small T5 encoder-decoder for text-to-text tasks like translation, summarization, and QA.
Vision-language model aligning images and text for zero-shot image classification and retrieval.
Multilingual cross-encoder for reranking search results with fine-grained relevance scoring.
Efficient MobileNetV3-small trained on ImageNet-1k for mobile image classification.
Cross-lingual RoBERTa trained on 100 languages for multilingual masked language modeling.
Small 125M parameter OPT decoder for open-source language generation research.
Tiny Qwen2 test model for TRL library development and unit testing.
Open embedding model with long context (8k tokens) and strong retrieval performance.
8B parameter Qwen3 language model for general text generation and conversational AI.
ComfyUI-compatible diffusion checkpoint fine-tuned from MiniMax H3 base model.
Original GPT-2 124M model for autoregressive text generation, widely used baseline.
9B multimodal model handling both image and text inputs for visual question answering.
RoBERTa base encoder optimized for masked language modeling and downstream NLP tasks.
Large English embedding model from BAAI for high-quality semantic search and retrieval.
Compact multilingual E5 embeddings for cross-lingual sentence similarity and retrieval.
Lightweight 82M parameter text-to-speech model fine-tuned from StyleTTS2 for English voice synthesis.
NVIDIA-optimized Qwen 35B MoE quantized to FP4 for efficient inference on NVIDIA hardware.
Instruction-tuned 7B Qwen model optimized for chat and task-following interactions.
FP8-quantized 35B MoE multimodal model balancing quality and inference efficiency.
Already-popular models created in the last ~6 months — the fast risers.
27B parameter vision-language model that accepts images and text for conversational tasks.
Vision-language model with compressed tensors for efficient multimodal feature extraction.
Large conversational text generation model from DeepSeek, V4 Pro variant with transformers support.
GLM-5.2 MoE model with DSA architecture for conversational text generation in English.
Image and text to video generation model producing dynamic videos from multimodal prompts.
Baidu's vision-language model for OCR tasks, extracting text from images with feature embeddings.
Google's 31B instruction-tuned Gemma 4 supporting multimodal image-text conversations.
Faster inference variant of DeepSeek V4 from July 2024 for low-latency generation.
Uncensored Qwen 3.6 35B MoE variant with aggressive tuning, GGUF quantized for vision tasks.
Qwen 3.5 27B distilled from Claude Opus 4.6 to enhance reasoning capabilities for vision-language.
NVIDIA's 3B vision model for object localization and spatial understanding tasks.
Gemma 4 12B tuned for code generation and reasoning, GGUF quantized for efficient inference.
Qwen 3.6 35B MoE model for multimodal image-text conversations, Apache 2.0 licensed.
Uncensored 9B Qwen 3.5 variant trained on 1M tokens, GGUF quantized for reasoning tasks.
Qwen 3.6 27B model for multimodal image-text conversations, Apache 2.0 licensed.
Heavily fine-tuned uncensored GGUF variant of Qwen for creative and unrestricted generation.
DeepSeek V4 Flash variant optimized for faster conversational text generation inference.
Text-to-video diffusion model based on quantized LTX-2.3, supports GGUF format.
Uncensored Qwen 3.5 9B with aggressive tuning, GGUF format for English/Chinese generation.
GLM-5.1 MoE model with DSA architecture for conversational text generation tasks.
Top trending models within a few curated task areas.
Mixture-of-experts text generation model trained on 2.4T tokens with 95B active parameters.
DeepSeek's V4 Pro checkpoint from August 2024 for conversational text generation tasks.
GGUF-quantized uncensored Qwen3.8 optimized for llama.cpp with speculative decoding support.
Faster inference variant of DeepSeek V4 from July 2024 for low-latency generation.
NVIDIA's 30B Nemotron Lightning model in FP4 precision for ultra-fast text generation.
FP8-quantized MoE variant of the 2.4T token Qwen model with 95B active parameters.
Tiny conversational LLM using custom hybrid architecture, MIT licensed for fast text generation.
Uncensored 27B Qwen variant with abliterated safety filters in quantized GGUF format.
9B parameter Qwen multimodal model supporting image-text-to-text generation tasks.
27B uncensored Qwen vision-language model in BF16 precision for multimodal generation.
2.9B parameter text-to-image diffusion model compatible with ComfyUI workflows.
Development version of FLUX diffusion model for high-quality text-to-image generation.
4-bit weight, 8-bit activation quantized image generation model optimized for ComfyUI.
Fast inference variant of Krea-2 image generator, fine-tuned for speed over quality.
Base Krea-2 diffusion model with custom pipeline for text-to-image synthesis.
Uncensored text encoder component for FLUX2 Klein 9B image generation pipeline.
Krea2-based text-to-image model packaged for ComfyUI workflows with custom licensing.
Fast diffusion model from Tongyi research with academic paper backing (arXiv 2511.22699).
Schnell (fast) variant of FLUX diffusion optimized for rapid text-to-image generation.
Stable Diffusion XL base model for high-resolution text-to-image generation (1024px).
Speaker diarization pipeline identifying who spoke when in multi-speaker audio recordings.
Compact 0.6B streaming ASR model from NVIDIA for real-time speech transcription.
OpenAI's large Whisper v3 for robust multilingual speech recognition across 99 languages.
Community-trained speaker diarization model for distinguishing speakers in audio streams.
Indic language ASR supporting Hindi, Kannada, and Malayalam with custom architecture.
1.7B parameter speech recognition model from Qwen with published research backing.
Faster turbo variant of Whisper v3 large for low-latency speech transcription.
English-only streaming ASR from NVIDIA at 0.6B parameters for real-time applications.
Microsoft ASR with integrated speaker diarization for transcription with speaker labels.
C++ implementation of Whisper for efficient on-device speech recognition inference.
Zero-shot text-to-speech with voice cloning from minimal reference audio samples.
ComfyUI workflows for MiniMax H3 reference-to-video and first-last-frame generation.
Zero-shot multilingual TTS with voice cloning and design capabilities across languages.
Lightweight 82M parameter text-to-speech model fine-tuned from StyleTTS2 for English voice synthesis.
Instruction-following multilingual TTS model with Chinese language support.
1.7B parameter TTS model with 12kHz output supporting custom voice synthesis.
ONNX-optimized multilingual speech synthesis model for deployment.
Multilingual TTS with voice cloning for speech generation across languages.
357M parameter NeMo TTS model with GGUF quantization for multilingual speech.
Multilingual TTS with voice cloning capabilities for diverse language support.
27B parameter vision-language model that accepts images and text for conversational tasks.
30B vision-language model from Meta for multimodal understanding and conversation.
FP8-quantized Qwen3.8-27B for efficient vision-language inference with lower precision.
Abliterated FP8 variant of Qwen3.8-27B with content restrictions removed for open responses.
Vision-language model with compressed tensors for efficient multimodal feature extraction.
Preview checkpoint from the Dots3-Note series for vision-language text generation.
Heavily fine-tuned uncensored GGUF variant of Qwen for creative and unrestricted generation.
27B uncensored vision-language model optimized for Apple MLX framework.
GGUF-quantized uncensored multimodal Qwen variant for efficient vision-text tasks.
30B vision-language model in GGUF format for image-text understanding tasks.
Trending datasets right now, by Hub trending score.
Multilingual distillation dataset from Qwen, GLM, and Kimi models for text generation.
Massive 10B+ token web text corpus for pretraining, ODC-BY licensed and tabular formatted.
1K video-text pairs dataset, likely for video understanding or generation tasks.
Anthropic's human preference dataset with 100K+ samples for RLHF alignment training.
Visual QA dataset for chart understanding and interpretation, 1-10K image samples.
Stack v3 multilingual code corpus for training code generation models, ODC-BY licensed.
1-10K unsolved math problems dataset for question answering and reasoning benchmarks.
10-100K text generation dataset with token classification annotations, MIT licensed.
NVIDIA's 10-100K sample RL dataset for training agentic text generation models.
10-100M record tabular dataset of internship listings, likely for search or ranking tasks.
50K distilled text examples for training smaller models from Qwen 3.8B outputs
Persona-based text generation dataset with character traits or dialogue patterns
Large-scale image-text archive with 1-10M samples for multimodal pretraining
Official vision-language benchmark for evaluating image understanding (1-10K samples)
100K+ math problems for supervised fine-tuning from NVIDIA's Nemotron series
Official grade school math word problems benchmark, crowdsourced for reasoning evaluation
CC0-licensed collection of 1-10K prompt examples for chat and QA use cases
Official benchmark for evaluating document information extraction capabilities
Software engineering training data (1-10K samples) for code generation fine-tuning
Billion-scale educational web text corpus filtered for high-quality pretraining data
All-time leaderboard by top downloads.
Japanese text classification dataset, likely from Kakolog (2ch/5ch archives)
Video-to-audio tokenization data, likely for multimodal generation training
Official HuggingFace documentation images (non-commercial license)
1K sample of reasoning traces or chain-of-thought examples for training or analysis
Classic Wikipedia-based language modeling benchmark for pretraining and masked LM tasks
Temporary or private dataset with no public metadata
Simulated robotics data from NVIDIA's GR00T project for embodied AI training
Language-conditioned tabletop manipulation dataset in LeRobot format from RLDS
Colossal Clean Crawled Corpus: massive web text for language model pretraining
Small image collection of censored or restricted historical documents
Cached Ubuntu filesystem states for OS-World agent benchmark environments
Internal HuggingFace documentation build artifacts and development data
Billion-scale refined web corpus with quality annotations for improved pretraining
Massive 10K+ hours of real-world omni-robot data for embodied AI (multi-language)
Official grade school math word problems benchmark, crowdsourced for reasoning evaluation
100K+ machine-generated image captions from evolutionary image breeding experiments
10K-100K images of typed digital signatures for classification and feature extraction tasks.
Dataset with minimal metadata; purpose unclear from available tags.
Massive 1T+ token pre-tokenized text corpus from FineWeb for language model pretraining.
100B+ 3D point cloud classification dataset, likely synthetic scans from ModelNet objects.
Collection of arXiv computer science papers in PDF format from 2020-2025.
Small text dataset with under 1K samples; purpose unclear from metadata.
100M-1B scale book corpus for language model pretraining, distributed via Dask.
85M sample mid-stage training data for LLaVA-OneVision multimodal model (arXiv 2509.23661).
Image dataset related to geographic names or locations.
Small video dataset with fewer than 1K samples; purpose unclear.
10M-100M timeseries samples pre-tokenized for JAT (likely Jack of All Trades) agent training.
Dataset for video-to-audio tokenization; likely supports multimodal generation.
1T+ scale robotics dataset (arXiv 2606.27375) with 130K scenarios for embodied AI.
Dataset with minimal metadata; purpose unclear from available tags.
Already-popular datasets created in the last ~6 months — the fast risers.
1K-10K machine-generated agent reasoning traces for text generation training.
1M-10M Korean persona-based conversations for instruction tuning and dialogue training.
10-100M record tabular dataset of internship listings, likely for search or ranking tasks.
8.7K reasoning traces from Claude Opus 4.6-4.7 for training step-by-step inference.
1K-10K samples from Claude Opus 4.6, possibly augmented or filtered.
10M-100M bilingual (EN/ZH) instruction tuning samples for supervised fine-tuning.
10K-100K reasoning traces from Hermes agent for training chain-of-thought capabilities.
1K-10K bilingual (EN/ZH) QA pairs for supervised fine-tuning of Aquila models.
Small collection of Claude-generated agent traces with code, formatted for training.
Stack v3 multilingual code corpus for training code generation models, ODC-BY licensed.
1M-10M multilingual instruction tuning samples for Step 3.5 Flash model training.
Hacker News posts and comments for text generation, classification, and retrieval tasks.
Trending spaces right now, by Hub trending score.
Gradio space for MiniMax Music3 model; likely generates or processes music.
Interactive demo for MiniMax H3 Turbo with LoRA fine-tuning capabilities
Benchmark comparing AI agents on long-term memory and recall tasks
Fast image editing interface using Qwen vision model with LoRA adapters
Gradio interface for video generation or editing experiments
Gradio app with MCP server integration for model inference
Adult content generation interface using Krea model
Docker-based tool for vulnerability scanning or security analysis
Experimental all-in-one Qwen image editor with multiple LoRA variants
Custom video generator supporting both text-to-video and image-to-video workflows
Free tool to detect AI-generated text from GPT and ChatGPT models
Image editing studio focused on photorealistic manipulation
Free API endpoint for Qwen 27B language model inference
Image-to-video generation interface with VIP model
High-speed inference demo for MiniMax H3 model with MCP integration
Wan 2.2 14B image-to-video model with custom LoRA support
LTX 2.3 image-to-video model with Eros variant fine-tuning
Lightricks' LTX 2.5 generative model demo with MCP server support
Microsoft's TRELLIS 2 model interactive demonstration
Manual submission leaderboard for evaluating text generation models with private test sets.
All-time leaderboard by top likes.
AI-powered website generation or analysis tool
Community benchmark ranking open-source language models on standardized English text tasks
Automated comic strip generator using AI for panels and dialogue
Virtual clothing try-on using Kwai's Kolors diffusion model
Development version of Black Forest Labs' FLUX.1 image generation model
Massive Text Embedding Benchmark leaderboard for comparing embedding models
Text-to-image generation demo, predecessor to DALL-E 2, generates images from prompts.
Generates optical illusion images using diffusion models with hidden patterns.
Fast text-to-image synthesis from Black Forest Labs, optimized for speed.
Meta's music generation model that creates audio from text descriptions.
Crowdsourced leaderboard ranking language models by human preference voting.
Microsoft research demo, likely for 3D or multimodal generation tasks.
Open alternative attempting GPT-4-like multimodal conversation capabilities.
Documentation or guide for training models at extreme scale with Nanotron framework.
Animates still portraits with motion control, turning images into video.
Fast image generation or manipulation tool with MCP server integration.
Identity-preserving image generation, swaps faces while maintaining likeness.
Text-to-speech synthesis model generating natural voice audio from text.
Optimized inference for WAN2 model with FP8 precision and AOTI compilation.
Tencent's 3D asset generation model, creates 3D objects from text or images.
Code generation or programming assistance tool, containerized for isolation.
Research guide for training small language models efficiently with visualizations.
Preview of WAN2 model with FP8 quantization and ahead-of-time compilation.
Unofficial Midjourney-style text-to-image generation interface or alternative.
Reverse-engineers prompts from images using CLIP for prompt discovery.
Removes backgrounds from images automatically, outputs transparent PNGs.
End-to-end text-to-speech with F5 model architecture for voice synthesis.
JAX-optimized Whisper speech recognition for faster transcription inference.
OpenAI's speech recognition model transcribing audio to text across languages.
Virtual try-on system that swaps clothing on people in images realistically.
Already-popular spaces created in the last ~6 months — the fast risers.
Gradio app with MCP server integration for model inference
Multi-language voice synthesis or speech processing from the K2 project.