-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
CodeGoat24/UnifiedReward-Think-7b
8B • Updated • 25 • 10 -
CodeGoat24/UnifiedReward-7b-v1.5
8B • Updated • 2.51k • 7
Collections
Discover the best community collections!
Collections including paper arxiv:2505.03318
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
mradermacher/UnifiedReward-qwen-32b-i1-GGUF
33B • Updated • 436 • 1 -
mradermacher/UnifiedReward-Think-qwen-7b-i1-GGUF
8B • Updated • 308
-
JudgeLRM: Large Reasoning Models as a Judge
Paper • 2504.00050 • Published • 62 -
RM-R1: Reward Modeling as Reasoning
Paper • 2505.02387 • Published • 80 -
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Paper • 2505.01441 • Published • 39 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
CodeGoat24/UnifiedReward-Think-qwen-7b
8B • Updated • 318 • 3 -
CodeGoat24/UnifiedReward-qwen-32b
33B • Updated • 20 • 1
-
OmniGen2: Exploration to Advanced Multimodal Generation
Paper • 2506.18871 • Published • 77 -
OmniGen: Unified Image Generation
Paper • 2409.11340 • Published • 115 -
Show-o Turbo: Towards Accelerated Unified Multimodal Understanding and Generation
Paper • 2502.05415 • Published • 22 -
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Paper • 2408.12528 • Published • 51
-
microsoft/bitnet-b1.58-2B-4T
Text Generation • 0.8B • Updated • 5.83k • 1.22k -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Paper • 2504.10449 • Published • 15 -
nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct
Text Generation • 8B • Updated • 103 • 15 -
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Paper • 2504.11536 • Published • 63
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
CodeGoat24/UnifiedReward-2.0-T2X-score-data
Viewer • Updated • 337k • 782 -
CodeGoat24/ImageGen-CoT-Reward-5K
Viewer • Updated • 5.54k • 118 • 1
-
InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
Paper • 2502.11573 • Published • 9 -
Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking
Paper • 2502.02339 • Published • 22 -
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
Paper • 2502.11775 • Published • 9 -
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Paper • 2412.18319 • Published • 39
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
CodeGoat24/UnifiedReward-Think-7b
8B • Updated • 25 • 10 -
CodeGoat24/UnifiedReward-7b-v1.5
8B • Updated • 2.51k • 7
-
OmniGen2: Exploration to Advanced Multimodal Generation
Paper • 2506.18871 • Published • 77 -
OmniGen: Unified Image Generation
Paper • 2409.11340 • Published • 115 -
Show-o Turbo: Towards Accelerated Unified Multimodal Understanding and Generation
Paper • 2502.05415 • Published • 22 -
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Paper • 2408.12528 • Published • 51
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
mradermacher/UnifiedReward-qwen-32b-i1-GGUF
33B • Updated • 436 • 1 -
mradermacher/UnifiedReward-Think-qwen-7b-i1-GGUF
8B • Updated • 308
-
microsoft/bitnet-b1.58-2B-4T
Text Generation • 0.8B • Updated • 5.83k • 1.22k -
M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Paper • 2504.10449 • Published • 15 -
nvidia/Llama-3.1-Nemotron-8B-UltraLong-2M-Instruct
Text Generation • 8B • Updated • 103 • 15 -
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Paper • 2504.11536 • Published • 63
-
JudgeLRM: Large Reasoning Models as a Judge
Paper • 2504.00050 • Published • 62 -
RM-R1: Reward Modeling as Reasoning
Paper • 2505.02387 • Published • 80 -
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Paper • 2505.01441 • Published • 39 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
CodeGoat24/UnifiedReward-2.0-T2X-score-data
Viewer • Updated • 337k • 782 -
CodeGoat24/ImageGen-CoT-Reward-5K
Viewer • Updated • 5.54k • 118 • 1
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 123 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 93 -
CodeGoat24/UnifiedReward-Think-qwen-7b
8B • Updated • 318 • 3 -
CodeGoat24/UnifiedReward-qwen-32b
33B • Updated • 20 • 1
-
InfiR : Crafting Effective Small Language Models and Multimodal Small Language Models in Reasoning
Paper • 2502.11573 • Published • 9 -
Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking
Paper • 2502.02339 • Published • 22 -
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
Paper • 2502.11775 • Published • 9 -
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Paper • 2412.18319 • Published • 39