NATOPS Aircraft Handling Signals Database

Yale Song, David Demirdjian, and Randall Davis
MIT Computer Science and Artificial Intelligence Laboratory
32 Vassar Street, Cambridge, MA 02139, USA
{yalesong,demirdj,davis}_at_csail.mit.edu

 

The Naval Air Training and Operating Procedures Standardization (NATOPS) manual standardizes general flight and operating procedures for the US naval aircraft. There are a large number of publications that belong to NATOPS; among the many, an aircraft signals manual (NAVAIR 00-80T-113) contains information about all aircraft systems including general aircraft and carrier flight deck handling signals.

We selected 24 NATOPS aircraft handling signals, the gestures most often used in routine practice on the deck environment. A stereo camera was used to collect this database, producing 320 x 240 pixel resolution images at 20 FPS. Twenty subjects repeated each of 24 gestures 20 times, resulting in 400 samples for each gesture class. Each sample had a unique duration; the average length of all samples was 2.34 sec (std.dev=0.62). Videos were recorded in a closed room environment with a constant illuminating condition, and with positions of cameras and subjects fixed throughout the recording. We use this controlled circumstance as our first step towards developing a proof-of-concept for NATOPS gesture recognition, and discovered that even this somewhat artificial environment still posed substantial challenges for our vocabulary.

Update (July 2016)

To get access to the dataset, please visit our git repository at https://github.com/yalesong/natops

Publication

Yale Song, David Demirdjian, and Randall Davis. Tracking Body and Hands For Gesture Recognition: NATOPS Aircraft Handling Signals Database. In Proceedings of the 9th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2011). Santa Barbara, CA, March 2011.

Feature Detaset (Matlab format)

This feature dataset includes body and hand features we estimated.
Matlab file

Feature Detaset (CSV format)

This feature dataset includes body and hand features we estimated.
Body and Hand Features (70 MB)      Segmentation Label (19 KB)

Video Dataset

This video dataset includes stereo camera-recorded images, depth maps, and mask images.
Note: The dataset only contains video clips from one subject (out of 20). Contact yalesong if you want the entire dataset (19 GB).
Video Clips (860 MB)     Ground-Truth Body and Hand Pose Labels (60 MB)