AI & Autonomy

AGT-CV: Drexel's Air-Ground Perception Dataset Targets the Off-Road Blind Spot in Multi-Robot AI

Drexel's AGT-CV dataset pairs a Clearpath Husky UGV with an Autel EVO II UAV for real-world cross-view collaborative perception research.

AGT-CV: Drexel's Air-Ground Perception Dataset Targets the Off-Road Blind Spot in Multi-Robot AI
Researchers at Drexel University have released AGT-CV, a real-world multi-modal dataset pairing a Clearpath Husky UGV with an Autel EVO II UAV across five unstructured terrain types, yielding over 13,000 synchronized frames. The dataset is the first of its kind built specifically for cross-view collaborative perception in off-road environments, integrating LiDAR, stereo vision, thermal imaging, and Meta's SAM 3 zero-shot segmentation.

Main Story

Collaborative perception between aerial and ground robots has long been constrained by a fundamental data problem: virtually no real-world datasets exist that capture overlapping, multi-modal observations from heterogeneous platforms navigating genuinely unstructured terrain. A new open dataset from Drexel University's robotics group directly addresses this gap.

The AGT-CV (Aerial-Ground Team Cross-View) dataset was collected using a Clearpath Husky UGV and an Autel EVO II UAV across diverse unstructured environments, including forest trails, rocky paths, muddy terrain, snow piles, and grass-covered fields.

The research team chose this pairing deliberately. Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex outdoor environments. Yet progress in multi-robot collaborative perception has been constrained by the lack of real-world datasets featuring overlapping multi-modal observations from platforms operating in unstructured terrain.

AGT-CV is designed to be the corrective. The dataset is collected across four unique environments, with over 13,000 synchronized frames spanning approximately 29 minutes of operation, and includes both SAM 3-based zero-shot segmentation and almost 8,000 manually labeled images.

A defining methodological choice was the collection timing. A unique aspect of the dataset is its early-spring collection period, during which sparse tree canopies allow the aerial robot to partially observe the ground robot and terrain through the trees, enabling occlusion-aware collaborative perception. This deliberate seasonal selection gives researchers a rare natural laboratory for studying partial occlusion — one of the hardest failure modes to replicate synthetically.

Unlike prior multi-robot datasets that primarily focus on SLAM or simulated cooperative driving, AGT-CV is specifically designed to support research on cross-view perception, air-ground viewpoint fusion, terrain-aware perception, and collaborative scene understanding in real off-road environments.

On the annotation side, the team deployed Meta's SAM 3, a zero-shot model released in November 2025 that detects, segments, and tracks objects in images and videos from concept prompts. SAM 3 introduces Promptable Concept Segmentation: given a short noun phrase, it returns unique masks and IDs for every matching instance at once — a capability where SAM 1 and 2 predicted only one object per prompt. Using SAM 3 for automated pre-annotation, alongside nearly 8,000 hand-labeled frames, gives AGT-CV a dual-track ground truth that future benchmarks can leverage at both ends of the labeling effort spectrum.

Technical Breakdown

Ground Platform — Clearpath Husky UGV The Clearpath Husky is a medium-sized unmanned ground vehicle designed for research and industrial applications in outdoor environments, with a steel chassis, all-wheel-drive electric drivetrain, and a payload capacity of up to 75 kg. It weighs 50 kg, reaches a maximum speed of 1.0 m/s, and offers 130 mm of ground clearance. In the AGT-CV configuration, the ground platform provides 3D LiDAR, stereo camera, IMU, and GPS data. The platform includes an onboard computer running Ubuntu and ROS, with pre-configured drivers for common sensors including Velodyne LiDAR and various camera systems.

Aerial Platform — Autel EVO II UAV The aerial platform contributes RGB imagery, thermal/infrared observations, and GPS from a complementary overhead viewpoint, enabling rich cross-modal and cross-view perception. The Autel EVO II Dual thermal variant combines an infrared imaging camera with an 8K visible-light camera. The infrared sensor features a vanadium oxide uncooled focal plane detector with 640×512 resolution, a 13 mm focal length, and 1–8× zoom. Maximum horizontal flight speed reaches 72 km/h and maximum flight time is 38 minutes.

Dataset Composition

  • Environments: 4 distinct terrain types (forest trails, rocky paths, muddy terrain, snow, grass)
  • Frames: >13,000 synchronized multi-modal frames
  • Duration: ~29 minutes of operational recording
  • Labels: ~8,000 manually annotated images + SAM 3 zero-shot segmentation masks
  • Sensor modalities: 3D LiDAR, stereo camera, IMU, GPS (ground); RGB, thermal/IR, GPS (aerial)

Autonomy Level The dataset targets collaborative perception research rather than closed-loop autonomy. The platforms operate as a coordinated observational team, with cross-view and cross-modal data fusion being the primary research target.

Industry Impact

For Robotics Researchers and Dataset Curators AGT-CV fills a documented vacuum. Very few datasets are designed for heterogeneous multi-robot applications, and prior sensor modality setups were generally unsuitable for multi-robot collaborative perception — making AGT-CV among the first real-world heterogeneous air-ground collaborative perception datasets explicitly designed for multi-modal collaborative robot perception research. Teams working on field robotics, search-and-rescue, and autonomous inspection now have a benchmark that reflects the sensor diversity of operational deployments.

For UAV and UGV Manufacturers The pairing of a Clearpath Husky — the most widely deployed ROS-based outdoor UGV in research, used at hundreds of universities and research institutions worldwide — with the Autel EVO II's dual-sensor payload validates a practical, commercially available hardware stack. Manufacturers building multi-robot platforms for off-road applications can reference AGT-CV as a benchmark configuration.

For AI and Perception Engineers The dataset's dual annotation strategy — combining SAM 3 zero-shot masks with ~8,000 hand-labeled images — establishes a practical workflow for scaling perception ground truth at low cost. SAM 3 generally performs well in zero-shot settings but benefits from fine-tuning for niche domains — precisely the kind of domain-specific gap that AGT-CV's manual labels are positioned to close. Engineers developing terrain-aware segmentation and cross-view localization models gain a ready-made evaluation split that does not rely on simulation.

For Defence and Critical-Infrastructure Operators Off-road multi-robot perception is a core capability requirement for perimeter monitoring, disaster response, and infrastructure inspection in GPS-degraded, canopy-occluded environments. AGT-CV's explicit focus on occlusion-aware collaborative perception and early-spring canopy conditions makes it directly relevant to operational scenario planning in these sectors.

For Standards and Certification Bodies As regulators begin to consider certification frameworks for autonomous multi-robot systems operating beyond visual line of sight in complex terrain, curated real-world datasets like AGT-CV will be essential inputs to evidence-based performance standards. The dataset's synchronized, multi-modal structure and dual ground-truth methodology offer a reproducible evaluation baseline that standards bodies can reference.

#multi-robot perception#collaborative ai#off-road autonomy#dataset#lidar#thermal imaging