Picture for Fahad Shahbaz Khan

Fahad Shahbaz Khan

UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities

Add code
Dec 13, 2024
Viaarxiv icon

Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

Add code
Nov 29, 2024
Figure 1 for Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
Figure 2 for Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
Figure 3 for Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
Figure 4 for Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
Viaarxiv icon

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Add code
Nov 28, 2024
Figure 1 for GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Figure 2 for GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Figure 3 for GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Figure 4 for GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Viaarxiv icon

ALOcc: Adaptive Lifting-based 3D Semantic Occupancy and Cost Volume-based Flow Prediction

Add code
Nov 12, 2024
Viaarxiv icon

Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis

Add code
Nov 11, 2024
Viaarxiv icon

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Add code
Nov 07, 2024
Figure 1 for VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Figure 2 for VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Figure 3 for VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Figure 4 for VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Viaarxiv icon

How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?

Add code
Oct 23, 2024
Figure 1 for How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
Figure 2 for How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
Figure 3 for How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
Figure 4 for How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
Viaarxiv icon

Frontiers in Intelligent Colonoscopy

Add code
Oct 22, 2024
Viaarxiv icon

DB-SAM: Delving into High Quality Universal Medical Image Segmentation

Add code
Oct 05, 2024
Viaarxiv icon

Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking

Add code
Oct 02, 2024
Viaarxiv icon