Abstract
Human action understanding is a fundamental and challenging task in computer
vision. Although there exists tremendous research on this area, most works
focus on action recognition, while action retrieval has received less
attention. In this paper, we focus on the neglected but important task of
image-based action retrieval which aims to find images that depict the same
action as a query image. We establish benchmarks for this task and set up
important baseline methods for fair comparison. We present an end-to-end model
that learns rich action representations from three aspects: the anchored
person, contextual regions, and the global image. A novel fusion transformer
module is designed to model the relationships among different features and
effectively fuse them into an action representation. Experiments on the
Stanford-40 and PASCAL VOC 2012 Action datasets show that the proposed method
significantly outperforms previous approaches for image-based action retrieval.
Citation
ID:
282449
Ref Key:
gui2024regionaware