Collections
Discover the best community collections!
Collections including paper arxiv:2408.08872
-
sentence-transformers/all-mpnet-base-v2
Sentence Similarity • 0.1B • Updated • 19M • • 1.19k -
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Paper • 1910.10683 • Published • 14 -
google-t5/t5-base
Translation • 0.2B • Updated • 1.62M • • 755 -
Attention Is All You Need
Paper • 1706.03762 • Published • 96
-
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
Paper • 2408.08872 • Published • 100 -
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Paper • 2408.11039 • Published • 63 -
Building and better understanding vision-language models: insights and future directions
Paper • 2408.12637 • Published • 133
-
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Paper • 2408.10188 • Published • 52 -
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
Paper • 2408.08872 • Published • 100 -
Building and better understanding vision-language models: insights and future directions
Paper • 2408.12637 • Published • 133 -
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Paper • 2408.12528 • Published • 51
-
sentence-transformers/all-mpnet-base-v2
Sentence Similarity • 0.1B • Updated • 19M • • 1.19k -
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Paper • 1910.10683 • Published • 14 -
google-t5/t5-base
Translation • 0.2B • Updated • 1.62M • • 755 -
Attention Is All You Need
Paper • 1706.03762 • Published • 96
-
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Paper • 2408.10188 • Published • 52 -
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
Paper • 2408.08872 • Published • 100 -
Building and better understanding vision-language models: insights and future directions
Paper • 2408.12637 • Published • 133 -
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Paper • 2408.12528 • Published • 51
-
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
Paper • 2408.08872 • Published • 100 -
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Paper • 2408.11039 • Published • 63 -
Building and better understanding vision-language models: insights and future directions
Paper • 2408.12637 • Published • 133