Picture for Min Dou

Min Dou

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Add code
Dec 06, 2024
Viaarxiv icon

ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving

Add code
Nov 08, 2024
Figure 1 for ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving
Figure 2 for ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving
Figure 3 for ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving
Figure 4 for ZOPP: A Framework of Zero-shot Offboard Panoptic Perception for Autonomous Driving
Viaarxiv icon

DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes

Add code
Sep 06, 2024
Figure 1 for DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
Figure 2 for DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
Figure 3 for DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
Figure 4 for DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
Viaarxiv icon

DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving

Add code
Aug 01, 2024
Figure 1 for DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
Figure 2 for DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
Figure 3 for DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
Figure 4 for DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
Viaarxiv icon

DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models

Add code
Jun 17, 2024
Figure 1 for DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
Figure 2 for DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
Figure 3 for DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
Figure 4 for DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
Viaarxiv icon

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Add code
Jun 13, 2024
Figure 1 for OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Figure 2 for OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Figure 3 for OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Figure 4 for OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Viaarxiv icon

OmniCorpus: An Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Add code
Jun 12, 2024
Figure 1 for OmniCorpus: An Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Figure 2 for OmniCorpus: An Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Figure 3 for OmniCorpus: An Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Figure 4 for OmniCorpus: An Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
Viaarxiv icon

Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving

Add code
May 24, 2024
Figure 1 for Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
Figure 2 for Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
Figure 3 for Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
Figure 4 for Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
Viaarxiv icon

Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond

Add code
May 06, 2024
Viaarxiv icon

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Add code
Apr 29, 2024
Viaarxiv icon