From a drone, a person fills less than a tenth of the frame and the camera never stops moving, so action recognisers built for ground cameras fall apart: I3D scores 98% on UCF-101 but 24% on UAV-Human. U-ActionNet first crops each clip to the people in it, extracts C3D features, and then works in the frequency domain: a temporal FFT separates moving actors from the static background, and Fourier self-attention captures long-range space–time context at FFT cost. A light version, with keyframe sampling, MobileNetV2 and a BiLSTM, has 8.4× fewer parameters and is built to run on the edge.
Figure 1. Action recognition from a drone. The drone records the scene, the people in each frame are located, features are extracted from the clip, and a classifier names the action.
Why drone footage breaks ground-camera models
Action recognition has been built on footage from cameras at eye level, where the person is large, upright and centred. Seen from a drone, the same person is a few dozen pixels in a frame dominated by ground, and a network trained on it learns more about the background than about the actor. The drop is not subtle:
People occupy less than 10% of a UAV frame. Without being told where to look, a model spends its capacity on grass, roads and roofs.
The drone drifts and turns, so the viewpoint changes within a clip, and motion blur and occlusion blur the line between actor and scene.
Strong video models, and full space–time self-attention in particular, need desktop or cloud GPUs. A drone or an edge box has neither.
Crop to the people, then look for motion in frequency
Figure 2. The server-side U-ActionNet. Object localisation (green) crops each clip to its people, a modified C3D (red) extracts features, two Fourier blocks (blue) run in parallel on them, and two dense layers (yellow) classify.
1 · Region of interest
Every clip goes through YOLOv8x trained on COCO. Only person detections are kept, and a person becomes the region of interest when they are detected in 32 consecutive frames with a confidence above 0.60. The crops are resized to 112 × 112 and stitched back into a clip, so the network sees the 10–15% of the frame where the action happens. This alone cuts the memory each video takes by up to 80%.
The region-of-interest module Figure 3
Figure 3. YOLOv8 detects objects, a filter drops everything but people, the region of interest is selected, and the frames are resized to 112 × 112 × 3 to form a new clip.
2 · Modified C3D features
The backbone is C3D cut after its sixth layer: those more general features transfer better and overfit less on small datasets. The Fourier module works on them, followed by two redesigned dense layers with dropout 0.5. Both Fourier branches read the same features in parallel.
3 · Fourier substance separation
Take the C3D features \(f \in \mathbb{R}^{C \times T' \times H' \times W'}\) and run a 1D FFT along time at every channel and location. A static region puts its energy at low temporal frequencies; motion shows up at high ones. Weighting the power by the squared frequency gives a dynamic mask \(D_{FS}\), and multiplying by the activations keeps what both moves and matters:
The separated features rank the four kinds of region in the order an action classifier needs, from weakest to strongest:
- 1 · static, not salient
- 2 · moving, not salient
- 3 · static, salient
- 4 · moving, salient
4 · Fourier self-attention
Self-attention over every space–time position costs \(O\big(C\,(HWT)^2\big)\). U-ActionNet instead correlates the features with themselves in the frequency domain, the way an autocorrelation does: a 2D FFT over time and space, a product with the complex conjugate, an inverse FFT, and a small scaled residual with \(\lambda_{FA} = 0.01\):
That costs \(O(C \cdot HWT \log HWT)\) and relates distant parts of the clip, such as an actor and the object they act on, without a single attention matrix. The server-side model trains with Adam at a learning rate of 0.001 for 50 epochs, with batches of 16 clips.
A version small enough for the drone
The server-side model has 63.3 million parameters, far too many for an edge device. U-ActionNet Light keeps the region-of-interest crop and the Fourier module, and replaces the rest with parts built for small hardware.
Figure 4. U-ActionNet Light. Pre-processing (region-of-interest crop and frame sampling), feature extraction (MobileNetV2 and the Fourier module), a BiLSTM that keeps temporal context, and dense layers.
- KeyframesAfter the person crop, FFmpeg's scene-change detection with a threshold of 0.01 keeps only frames that differ from the last, cutting the frame count by up to 30%.
- FeaturesNormalised 112 × 112 frames go through MobileNetV2, one frame at a time, with only its last ten layers fine-tuned. The same Fourier module then separates motion and adds context.
- TimeA BiLSTM reads the frame features forwards and backwards, and dense layers with dropout and L2 regularisation (1e-4) classify. It trains with SGD at a learning rate of 0.01 for 50 epochs.
| Parameters | Server-side | Light |
|---|---|---|
| Total | 63,317,258 | 7,528,522 |
| Trainable | 10,102,794 | 2,003,018 |
| Non-trainable | 53,214,464 | 5,525,504 |
The light model has 8.4× fewer parameters (5× fewer trainable ones) and runs about 1.8× faster: 0.164 s against 0.291 s on UAV-Human, and 0.216 s against 0.395 s on Drone Action.
What it gives up in accuracy is in the results below.
Two aerial datasets, ten actions each
Both datasets are used as ten-class subsets chosen for surveillance: fights, threats and calls for help from UAV-Human, which has 155 classes in full, and ten of the thirteen actions in Drone Action. Each is split 80/20 by video. The scores below are on these subsets and are not comparable to results on the full benchmarks.
10 videos per class, 8 for training and 2 for testing.
- Walking
- Jogging
- Running
- Hitting with a bottle
- Hitting with a stick
- Stabbing
- Punching
- Kicking
- Clapping
- Waving hands
| Method | Loss | Top-1 (%) | Top-3 (%) |
|---|---|---|---|
| C3D + FC | 0.3774 | 78.52 | 86.71 |
| FFT-UAVNet | 0.2554 | 88.77 | 98.19 |
| U-ActionNet | 0.1592 | 94.94 | 98.87 |
U-ActionNet reaches 94.94% top-1, 6.17 points above the conference version, FFT-UAVNet, and 16.42 above C3D with fully connected layers. Top-3 is nearly saturated for both Fourier models (98.87% and 98.19%).
The remaining errors are between actions that look alike from above: jogging is taken for running 17% of the time, and stabbing and punching for hitting with a bottle 12% and 11% of the time.
F1 per class · Drone Action server-side and light models
| Action | U-ActionNet | U-ActionNet Light |
|---|---|---|
| Clapping | 0.96 | 0.82 |
| Hitting with a bottle | 0.87 | 0.79 |
| Hitting with a stick | 1.00 | 0.90 |
| Jogging | 0.90 | 0.67 |
| Kicking | 0.98 | 0.94 |
| Running | 0.88 | 0.40 |
| Stabbing | 0.94 | 0.69 |
| Walking | 1.00 | 0.72 |
| Waving hands | 0.98 | 0.86 |
| Punching | 0.94 | 0.86 |
F1 on the test split, from the paper's per-class classification reports.
30 videos per class, 24 for training and 6 for testing.
- Punching someone
- Kicking someone
- Pushing someone
- Slap on the back
- Threat with a knife
- Hold someone hostage
- Threat with a gun
- Drag someone
- Call for help
- Stab with a knife
| Method | Loss | Top-1 (%) | Top-3 (%) |
|---|---|---|---|
| X3D | 1.977 | 31.33 | 45.60 |
| FFT-UAVNet | 0.992 | 64.86 | 83.37 |
| C3D + FC | 0.3497 | 88.70 | 94.44 |
| U-ActionNet | 0.1924 | 95.05 | 98.64 |
U-ActionNet reaches 95.05% top-1, 6.35 points above C3D with fully connected layers, the closest baseline, and 4.20 points higher in top-3.
Every class has an F1 of at least 0.91. The most common mistakes are punching taken for a threat with a knife, and pushing for kicking, each 6% of the time.
F1 per class · UAV-Human server-side and light models
| Action | U-ActionNet | U-ActionNet Light |
|---|---|---|
| Punching someone | 0.91 | 0.81 |
| Kicking someone | 0.92 | 0.86 |
| Pushing someone | 0.91 | 0.72 |
| Slap on the back | 0.99 | 0.84 |
| Threat with a knife | 0.93 | 0.81 |
| Hold someone hostage | 0.97 | 0.77 |
| Threat with a gun | 0.96 | 0.91 |
| Drag someone | 0.97 | 0.84 |
| Call for help | 0.99 | 0.87 |
| Stab with a knife | 0.96 | 0.93 |
F1 on the test split, from the paper's per-class classification reports.
Figure 5. Where the server-side model errs. Confusion matrices on UAV-Human (left) and Drone Action (right), in percent of each true class.
The light model against its baseline
The paper's baseline is a MobileNet and BiLSTM pipeline without keyframe sampling. U-ActionNet Light improves on it on both datasets:
| Dataset | Loss | Top-1 (%) | Top-3 (%) |
|---|---|---|---|
| Drone Action | 0.6037 → 0.5081 | 74.19 → 80.43 +6.24 | 86.91 → 90.93 +4.02 |
| UAV-Human | 0.5218 → 0.4988 | 82.38 → 84.74 +2.36 | 92.05 → 93.41 +1.36 |
Baseline → U-ActionNet Light; changes in percentage points.
Against the server-side model, the light model gives up 10.3 points of top-1 on UAV-Human and 14.5 on Drone Action, in exchange for 8.4× fewer parameters and about half the inference time.
Its weakest classes are the ones that differ only in speed: on Drone Action, running reaches an F1 of 0.40 and jogging 0.67, and the model often takes running for walking.
From a conference paper to a journal paper
Both papers grew out of a BSc thesis in Robotics and Mechatronics Engineering at the University of Dhaka, Enhancing UAV Based Human Action Recognition: A Deep Learning Approach, supervised by Md Mehedi Hasan.
The first version: C3D with Fourier object disentanglement and Fourier space–time attention, on the same ten UAV-Human classes. It more than doubles the top-1 accuracy of C3D and X3D trained the same way.
Adds the YOLOv8 region-of-interest crop, a second dataset (Drone Action) and the light edge model. On the same UAV-Human classes, top-1 goes from 64.86% to 95.05%.
| FFT-UAVNet paper · UAV-Human, 10 classes | Loss | Top-1 (%) | Top-3 (%) |
|---|---|---|---|
| C3D | 2.145 | 28.05 | 58.65 |
| X3D | 1.977 | 31.33 | 45.60 |
| FFT-UAVNet | 0.992 | 64.86 | 83.37 |
From Table I of the conference paper; the baselines use two fully connected layers on unmodified networks.
Scope: both datasets are used as ten-class subsets with small test splits (60 and 20 videos), and the method assumes at most two people acting in a clip. Extending it to many actors doing different things at once, and to better frame sampling for the light model, is left for future work.
Citation
If this work is useful to you, please cite it as:
@article{chowdhury2024u,
title = {{U-ActionNet}: Dual-Pathway Fourier Networks With Region-of-Interest
Module for Efficient Action Recognition in {UAV} Surveillance},
author = {Chowdhury, Abdul Monaf and Imran, Ahsan and Hasan, Md Mehedi and
Ahmed, Riad and Azad, Akm and Alyami, Salem A.},
journal = {IEEE Access},
volume = {12},
year = {2024},
doi = {10.1109/ACCESS.2024.3516586}
}