Models
Datasets
Spaces
Docs
Enterprise
Pricing
Log In
Sign Up

Collections

Discover the best community collections!

Collections including paper arxiv:2505.16839

Discrete Diffusion LLM & MLLM

An collection of research/models in discrete diffusion large language and multimodal models

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Paper • 2506.13759 • Published Jun 16 • 43
GSAI-ML/LLaDA-8B-Instruct

Text Generation • 8B • Updated 27 days ago • 283k • 332
Dream-org/Dream-v0-Base-7B

Text Generation • 8B • Updated Jul 15 • 99.6k • 51
Dream-org/Dream-v0-Instruct-7B

Text Generation • 8B • Updated Jul 15 • 55.2k • 142

Large Language Diffusion Models

Paper • 2502.09992 • Published Feb 14 • 122
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Paper • 2503.09573 • Published Mar 12 • 73
MMaDA: Multimodal Large Diffusion Language Models

Paper • 2505.15809 • Published May 21 • 97
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective

Paper • 2505.15045 • Published May 21 • 54

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Paper • 2402.04252 • Published Feb 6, 2024 • 28
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

Paper • 2402.03749 • Published Feb 6, 2024 • 14
ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Paper • 2402.04615 • Published Feb 7, 2024 • 44
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss

Paper • 2402.05008 • Published Feb 7, 2024 • 23

LArge VIsion-language Diffusion moDel with mAsking

jacklishufan/lavida-llada-v1.0-instruct

8B • Updated May 22 • 479 • 6
jacklishufan/lavida-llada-1.0-fim

8B • Updated May 22 • 10 • 1
hbXNov/lavida-llada-reason

8B • Updated May 10 • 774 • 2
jacklishufan/lavida-llada-1.0-lowres

8B • Updated May 22 • 7 • 1

Image-Video MultiModal Understanding

Apollo: An Exploration of Video Understanding in Large Multimodal Models

Paper • 2412.10360 • Published Dec 13, 2024 • 147
SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization

Paper • 2501.01245 • Published Jan 2 • 5
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Paper • 2501.00599 • Published Dec 31, 2024 • 47
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

Paper • 2501.08326 • Published Jan 14 • 33

Discrete Diffusion LLM & MLLM

An collection of research/models in discrete diffusion large language and multimodal models

Discrete Diffusion in Large Language and Multimodal Models: A Survey

Paper • 2506.13759 • Published Jun 16 • 43
GSAI-ML/LLaDA-8B-Instruct

Text Generation • 8B • Updated 27 days ago • 283k • 332
Dream-org/Dream-v0-Base-7B

Text Generation • 8B • Updated Jul 15 • 99.6k • 51
Dream-org/Dream-v0-Instruct-7B

Text Generation • 8B • Updated Jul 15 • 55.2k • 142

LArge VIsion-language Diffusion moDel with mAsking

jacklishufan/lavida-llada-v1.0-instruct

8B • Updated May 22 • 479 • 6
jacklishufan/lavida-llada-1.0-fim

8B • Updated May 22 • 10 • 1
hbXNov/lavida-llada-reason

8B • Updated May 10 • 774 • 2
jacklishufan/lavida-llada-1.0-lowres

8B • Updated May 22 • 7 • 1

Large Language Diffusion Models

Paper • 2502.09992 • Published Feb 14 • 122
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Paper • 2503.09573 • Published Mar 12 • 73
MMaDA: Multimodal Large Diffusion Language Models

Paper • 2505.15809 • Published May 21 • 97
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective

Paper • 2505.15045 • Published May 21 • 54

Image-Video MultiModal Understanding

Apollo: An Exploration of Video Understanding in Large Multimodal Models

Paper • 2412.10360 • Published Dec 13, 2024 • 147
SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization

Paper • 2501.01245 • Published Jan 2 • 5
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Paper • 2501.00599 • Published Dec 31, 2024 • 47
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

Paper • 2501.08326 • Published Jan 14 • 33

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Paper • 2402.04252 • Published Feb 6, 2024 • 28
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

Paper • 2402.03749 • Published Feb 6, 2024 • 14
ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Paper • 2402.04615 • Published Feb 7, 2024 • 44
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss

Paper • 2402.05008 • Published Feb 7, 2024 • 23

Company

TOS Privacy About Jobs

Website

Models Datasets Spaces Pricing Docs