Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatially Relation Matching

Nov 21, 2023

Meng Chu, Zhedong Zheng, Wei Ji, Tat-Seng Chua

Figure 1 for Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatially Relation Matching

Figure 2 for Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatially Relation Matching

Figure 3 for Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatially Relation Matching

Figure 4 for Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatially Relation Matching

Share this with someone who'll enjoy it:

Abstract:Drone navigation through natural language commands remains a significant challenge due to the lack of publicly available multi-modal datasets and the intricate demands of fine-grained visual-text alignment. In response to this pressing need, we present a new human-computer interaction annotation benchmark called GeoText-1652, meticulously curated through a robust Large Language Model (LLM)-based data generation framework and the expertise of pre-trained vision models. This new dataset seamlessly extends the existing image dataset, \ie, University-1652, with spatial-aware text annotations, encompassing intricate image-text-bounding box associations. Besides, we introduce a new optimization objective to leverage fine-grained spatial associations, called blending spatial matching, for region-level spatial relation matching. Extensive experiments reveal that our approach maintains an exceptional recall rate under varying description complexities. This underscores the promising potential of our approach in elevating drone control and navigation through the seamless integration of natural language commands in real-world scenarios.

* 10 pages, 6 figures

View paper on

Share this with someone who'll enjoy it:

Title:Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatially Relation Matching

Paper and Code