IEEE Access 2024UAV action recognition

U-ActionNet: Dual-Pathway Fourier Networks with Region-of-Interest Module for Efficient Action Recognition in UAV Surveillance

  • Abdul Monaf Chowdhury1
  • Ahsan Imran1
  • Md Mehedi Hasan1
  • Riad Ahmed2
  • Akm Azad3
  • Salem A. Alyami3
  • 1University of Dhaka
  • 2BRAC University
  • 3Imam Mohammad Ibn Saud Islamic University
TL;DR

From a drone, a person fills less than a tenth of the frame and the camera never stops moving, so action recognisers built for ground cameras fall apart: I3D scores 98% on UCF-101 but 24% on UAV-Human. U-ActionNet first crops each clip to the people in it, extracts C3D features, and then works in the frequency domain: a temporal FFT separates moving actors from the static background, and Fourier self-attention captures long-range space–time context at FFT cost. A light version, with keyframe sampling, MobileNetV2 and a BiLSTM, has 8.4× fewer parameters and is built to run on the edge.

95.05%
Top-1 on ten UAV-Human actions
+6.35 over C3D + FC
94.94%
Top-1 on ten Drone Action actions
+6.17 over FFT-UAVNet
8.4×
Fewer parameters in U-ActionNet Light: 7.5M against 63.3M
0.16s
Light-model inference on UAV-Human, against 0.29 s server-side
Pipeline: a drone records video; frames go through human actor detection, feature extraction, and action classification into labels such as punching, kicking and walking.

Figure 1. Action recognition from a drone. The drone records the scene, the people in each frame are located, features are extracted from the clip, and a classifier names the action.

The problem

Why drone footage breaks ground-camera models

Action recognition has been built on footage from cameras at eye level, where the person is large, upright and centred. Seen from a drone, the same person is a few dozen pixels in a frame dominated by ground, and a network trained on it learns more about the background than about the actor. The drop is not subtle:

The same recogniser, from the ground and from the air I3D top-1 accuracy (%)
UCF-101 98.0
HMDB-51 80.9
UAV-Human 23.86
TinyVIRAT 28.7
UCF-Aerial 16.8
Ground cameras Drone footage
Gap 1Tiny actors

People occupy less than 10% of a UAV frame. Without being told where to look, a model spends its capacity on grass, roads and roofs.

Gap 2A moving camera

The drone drifts and turns, so the viewpoint changes within a clip, and motion blur and occlusion blur the line between actor and scene.

Gap 3No compute on board

Strong video models, and full space–time self-attention in particular, need desktop or cloud GPUs. A drone or an edge box has neither.

Method

Crop to the people, then look for motion in frequency

U-ActionNet: YOLO object localisation, m-C3D feature extraction, then Fourier substance separation and Fourier self-attention in parallel, fully connected layers, and the classification label.

Figure 2. The server-side U-ActionNet. Object localisation (green) crops each clip to its people, a modified C3D (red) extracts features, two Fourier blocks (blue) run in parallel on them, and two dense layers (yellow) classify.

1 · Region of interest

Every clip goes through YOLOv8x trained on COCO. Only person detections are kept, and a person becomes the region of interest when they are detected in 32 consecutive frames with a confidence above 0.60. The crops are resized to 112 × 112 and stitched back into a clip, so the network sees the 10–15% of the frame where the action happens. This alone cuts the memory each video takes by up to 80%.

The region-of-interest module Figure 3
Region-of-interest module: input video, YOLO object detector, detection filter, ROI selection, resize frames to 112 by 112 by 3, generated video.

Figure 3. YOLOv8 detects objects, a filter drops everything but people, the region of interest is selected, and the frames are resized to 112 × 112 × 3 to form a new clip.

2 · Modified C3D features

The backbone is C3D cut after its sixth layer: those more general features transfer better and overfit less on small datasets. The Fourier module works on them, followed by two redesigned dense layers with dropout 0.5. Both Fourier branches read the same features in parallel.

3 · Fourier substance separation

Take the C3D features \(f \in \mathbb{R}^{C \times T' \times H' \times W'}\) and run a 1D FFT along time at every channel and location. A static region puts its energy at low temporal frequencies; motion shows up at high ones. Weighting the power by the squared frequency gives a dynamic mask \(D_{FS}\), and multiplying by the activations keeps what both moves and matters:

\[ \begin{aligned} \mathcal{A}_T(f)(k) &= \textstyle\sum_{n=0}^{T'-1} f(c,n,h,w)\, e^{-2\pi i kn/T'}, \\ D_{FS} &= \lVert \mathcal{A}_T(f)(k) \rVert_2^2 \cdot \lVert fr_k \rVert_2^2, \\ \nabla_{FS} &= f \odot D_{FS}. \end{aligned} \]
substance separation

The separated features rank the four kinds of region in the order an action classifier needs, from weakest to strongest:

  • 1 · static, not salient
  • 2 · moving, not salient
  • 3 · static, salient
  • 4 · moving, salient
InteractiveWhat the frequency mask keeps
A synthetic 16-frame clip, computed live; not model activations
Feature map f at frame t
Dynamic mask D, whole clip
Separated f ⊙ D
1 / 16
Share of the separated signal where the person moves –
On the parked car –
On the grass –

Pick an action, toggle the camera shake, and play the clip.

4 · Fourier self-attention

Self-attention over every space–time position costs \(O\big(C\,(HWT)^2\big)\). U-ActionNet instead correlates the features with themselves in the frequency domain, the way an autocorrelation does: a 2D FFT over time and space, a product with the complex conjugate, an inverse FFT, and a small scaled residual with \(\lambda_{FA} = 0.01\):

\[ \begin{aligned} \mathcal{T}_{ST} &= \mathrm{FFT}_{t,\,hw}(f), \qquad \mathcal{F}_{ST} = \mathcal{T}_{ST} \odot \mathcal{T}_{ST}^{*}, \\ \psi_{FA} &= \lambda_{FA}\, \mathrm{IFFT}(\mathcal{F}_{ST}) + f. \end{aligned} \]
Fourier attention

That costs \(O(C \cdot HWT \log HWT)\) and relates distant parts of the clip, such as an actor and the object they act on, without a single attention matrix. The server-side model trains with Adam at a learning rate of 0.001 for 50 epochs, with batches of 16 clips.

Edge model

A version small enough for the drone

The server-side model has 63.3 million parameters, far too many for an edge device. U-ActionNet Light keeps the region-of-interest crop and the Fourier module, and replaces the rest with parts built for small hardware.

U-ActionNet Light: input frames, ROI extraction and frame sampling, MobileNet and Fourier module for feature extraction, a BiLSTM temporal information retainer, and dense layers to the classification label.

Figure 4. U-ActionNet Light. Pre-processing (region-of-interest crop and frame sampling), feature extraction (MobileNetV2 and the Fourier module), a BiLSTM that keeps temporal context, and dense layers.

  1. KeyframesAfter the person crop, FFmpeg's scene-change detection with a threshold of 0.01 keeps only frames that differ from the last, cutting the frame count by up to 30%.
  2. FeaturesNormalised 112 × 112 frames go through MobileNetV2, one frame at a time, with only its last ten layers fine-tuned. The same Fourier module then separates motion and adds context.
  3. TimeA BiLSTM reads the frame features forwards and backwards, and dense layers with dropout and L2 regularisation (1e-4) classify. It trains with SGD at a learning rate of 0.01 for 50 epochs.
Average inference time seconds, lower is faster
UAV-Human · server 0.291
UAV-Human · light 0.164
Drone Action · server 0.395
Drone Action · light 0.216
Server-side U-ActionNet U-ActionNet Light
Parameters Server-side Light
Total 63,317,258 7,528,522
Trainable 10,102,794 2,003,018
Non-trainable 53,214,464 5,525,504

The light model has 8.4× fewer parameters (5× fewer trainable ones) and runs about 1.8× faster: 0.164 s against 0.291 s on UAV-Human, and 0.216 s against 0.395 s on Drone Action.

What it gives up in accuracy is in the results below.

Results

Two aerial datasets, ten actions each

Both datasets are used as ten-class subsets chosen for surveillance: fights, threats and calls for help from UAV-Human, which has 155 classes in full, and ten of the thirteen actions in Drone Action. Each is split 80/20 by video. The scores below are on these subsets and are not comparable to results on the full benchmarks.

10 videos per class, 8 for training and 2 for testing.

  • Walking
  • Jogging
  • Running
  • Hitting with a bottle
  • Hitting with a stick
  • Stabbing
  • Punching
  • Kicking
  • Clapping
  • Waving hands
Drone Action · 10 classes top-1 accuracy (%)
C3D + FC 78.52
FFT-UAVNet 88.77
U-ActionNet 94.94
Method Loss Top-1 (%) Top-3 (%)
C3D + FC 0.3774 78.52 86.71
FFT-UAVNet 0.2554 88.77 98.19
U-ActionNet 0.1592 94.94 98.87

U-ActionNet reaches 94.94% top-1, 6.17 points above the conference version, FFT-UAVNet, and 16.42 above C3D with fully connected layers. Top-3 is nearly saturated for both Fourier models (98.87% and 98.19%).

The remaining errors are between actions that look alike from above: jogging is taken for running 17% of the time, and stabbing and punching for hitting with a bottle 12% and 11% of the time.

F1 per class · Drone Action server-side and light models
Action U-ActionNet U-ActionNet Light
Clapping 0.96 0.82
Hitting with a bottle 0.87 0.79
Hitting with a stick 1.00 0.90
Jogging 0.90 0.67
Kicking 0.98 0.94
Running 0.88 0.40
Stabbing 0.94 0.69
Walking 1.00 0.72
Waving hands 0.98 0.86
Punching 0.94 0.86

F1 on the test split, from the paper's per-class classification reports.

30 videos per class, 24 for training and 6 for testing.

  • Punching someone
  • Kicking someone
  • Pushing someone
  • Slap on the back
  • Threat with a knife
  • Hold someone hostage
  • Threat with a gun
  • Drag someone
  • Call for help
  • Stab with a knife
UAV-Human · 10 classes top-1 accuracy (%)
X3D 31.33
FFT-UAVNet 64.86
C3D + FC 88.7
U-ActionNet 95.05
Method Loss Top-1 (%) Top-3 (%)
X3D 1.977 31.33 45.60
FFT-UAVNet 0.992 64.86 83.37
C3D + FC 0.3497 88.70 94.44
U-ActionNet 0.1924 95.05 98.64

U-ActionNet reaches 95.05% top-1, 6.35 points above C3D with fully connected layers, the closest baseline, and 4.20 points higher in top-3.

Every class has an F1 of at least 0.91. The most common mistakes are punching taken for a threat with a knife, and pushing for kicking, each 6% of the time.

F1 per class · UAV-Human server-side and light models
Action U-ActionNet U-ActionNet Light
Punching someone 0.91 0.81
Kicking someone 0.92 0.86
Pushing someone 0.91 0.72
Slap on the back 0.99 0.84
Threat with a knife 0.93 0.81
Hold someone hostage 0.97 0.77
Threat with a gun 0.96 0.91
Drag someone 0.97 0.84
Call for help 0.99 0.87
Stab with a knife 0.96 0.93

F1 on the test split, from the paper's per-class classification reports.

Confusion matrices of the server-side model on UAV-Human and Drone Action; most mass is on the diagonal, with jogging confused for running and stabbing and punching for hitting with a bottle.

Figure 5. Where the server-side model errs. Confusion matrices on UAV-Human (left) and Drone Action (right), in percent of each true class.

The light model against its baseline

The paper's baseline is a MobileNet and BiLSTM pipeline without keyframe sampling. U-ActionNet Light improves on it on both datasets:

Dataset Loss Top-1 (%) Top-3 (%)
Drone Action 0.6037 → 0.5081 74.19 → 80.43 +6.24 86.91 → 90.93 +4.02
UAV-Human 0.5218 → 0.4988 82.38 → 84.74 +2.36 92.05 → 93.41 +1.36

Baseline → U-ActionNet Light; changes in percentage points.

Against the server-side model, the light model gives up 10.3 points of top-1 on UAV-Human and 14.5 on Drone Action, in exchange for 8.4× fewer parameters and about half the inference time.

Its weakest classes are the ones that differ only in speed: on Drone Action, running reaches an F1 of 0.40 and jogging 0.67, and the model often takes running for walking.

From FFT-UAVNet

From a conference paper to a journal paper

Both papers grew out of a BSc thesis in Robotics and Mechatronics Engineering at the University of Dhaka, Enhancing UAV Based Human Action Recognition: A Deep Learning Approach, supervised by Md Mehedi Hasan.

STI 2023FFT-UAVNet

The first version: C3D with Fourier object disentanglement and Fourier space–time attention, on the same ten UAV-Human classes. It more than doubles the top-1 accuracy of C3D and X3D trained the same way.

IEEE Access 2024U-ActionNet

Adds the YOLOv8 region-of-interest crop, a second dataset (Drone Action) and the light edge model. On the same UAV-Human classes, top-1 goes from 64.86% to 95.05%.

FFT-UAVNet paper · UAV-Human, 10 classes Loss Top-1 (%) Top-3 (%)
C3D 2.145 28.05 58.65
X3D 1.977 31.33 45.60
FFT-UAVNet 0.992 64.86 83.37

From Table I of the conference paper; the baselines use two fully connected layers on unmodified networks.

Scope: both datasets are used as ten-class subsets with small test splits (60 and 20 videos), and the method assumes at most two people acting in a clip. Extending it to many actors doing different things at once, and to better frame sampling for the light model, is left for future work.

Cite

Citation

If this work is useful to you, please cite it as:

@article{chowdhury2024u,
  title   = {{U-ActionNet}: Dual-Pathway Fourier Networks With Region-of-Interest
             Module for Efficient Action Recognition in {UAV} Surveillance},
  author  = {Chowdhury, Abdul Monaf and Imran, Ahsan and Hasan, Md Mehedi and
             Ahmed, Riad and Azad, Akm and Alyami, Salem A.},
  journal = {IEEE Access},
  volume  = {12},
  year    = {2024},
  doi     = {10.1109/ACCESS.2024.3516586}
}