Constarium
← Search

Data · dataset · 2026

Towards effective and efficient video object detection

Listed in ZivaHub and Deakin Research Online and DMU Figshare — shown once because both records carry DOI 10.17034/32633805.v1

<br>The past decade has witnessed great progress in object detection on still images.

Description

As one of the fundamental computer vision tasks, object detection aims to locate and classify given objects simultaneously. However, in many real-world computer vision applications, e.g., video surveillance and autonomous driving, data are obtained in the format of video rather than image.

Till now, video object detection is still very challenging. <br><br>This thesis focuses on addressing core challenges for video object detection. Specifically, we summarise three research questions. Firstly, how to achieve accurate video object detection performance on low-quality frames?

Read the rest (4 more)

Secondly, how to achieve accurate video object detection algorithms with real-time speed? Finally, how to extract spatio-temporal features more efficiently and effectively? <br><br>In this thesis, we present three novel methods to address the aforementioned challenges and achieve effective and efficient video object detection. Firstly, we propose a memory bank structure to enhance the low-quality frame using features from the previous frames.

Using the memory bank, we can easily enhance the quality of features extracted from poor video frames. Therefore, our method outperforms many state-of-the-art methods on the large-scale ImageNet VID benchmark. Secondly, we propose a general framework that can be applied to various one-stage detectors for video object detection.

We design a location prior network and a size prior network to skip unnecessary computations on background regions. Thirdly, we carefully design a temporal dilated transformer block (TDTB) to build a simple yet effective backbone for dense video tasks, named temporal dilated video transformer (TDViT). TDViT achieves excellent performance on two widely used dense video benchmarks, ImageNet VID for video object detection and YouTube VIS for video instance segmentation.

The excellent performance and generalisation ability demonstrate that our TDViT can be served as a general backbone for dense video tasks.<br>

Links

Where it is published

Catalogue records · 1

Topics

Inferred from text
Image 75% · Video 75%
Provenance · 3 source records, 9 field assertions
SourceKeyLast seenRaw
ZivaHuboai:figshare.com:article/3263380510 d agoJSON v1
Deakin Research Onlineoai:figshare.com:article/3263380510 d agoJSON v1
DMU Figshareoai:figshare.com:article/3263380510 d agoJSON v1
FieldAssertionExtractorEvidence
concepts[field].local:field:earth-environmentalmapping · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
concepts[field].local:field:earth-environmentalmapping · figshare dmu ac ukconnector:figshare_dmu_ac_uk@1.0.0
concepts[field].local:field:earth-environmentalmapping · dro deakin edu auconnector:dro_deakin_edu_au@1.0.0
concepts[modality].local:modality:imageenrichment · zivahub uct ac zakeyword-concept-rules@1.0.0title+description (75%)
concepts[modality].local:modality:videoenrichment · zivahub uct ac zakeyword-concept-rules@1.0.0title+description (75%)
descriptionsource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0/metadata/dc/description
license_textsource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
publication_datesource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0
titlesource · zivahub uct ac zaconnector:zivahub_uct_ac_za@1.0.0/metadata/dc/title