Spatiotemporal Residual Attentive Networks

Spatiotemporal Residual Attentive Networks predicts dynamic eye-fixation maps (video saliency) by integrating appearance and motion information to model spatiotemporal visual attention.


Key Features:

  • Spatiotemporal Feature Integration: Integrates appearance and motion streams via dense residual cross-connections to fuse multi-layer motion features with appearance information and enable early multi-path information exchange.
  • Multi-Scale Saliency Learning: Employs a composite attention mechanism to learn local attentions at multiple scales and global attention priors end-to-end, enhancing fused spatiotemporal features for saliency prediction.
  • Temporal Characteristics Modeling: Uses a lightweight convolutional Gated Recurrent Unit (convGRU) to model long-term temporal characteristics and temporal dependencies in video sequences.

Scientific Applications:

  • Video Saliency Prediction: Predicts spatially and temporally varying saliency maps for dynamic visual content.
  • Eye-Tracking Studies: Provides model-based estimates of viewer fixation distributions for comparison with eye-tracking data.
  • Human–Computer Interaction: Informs interfaces and systems that rely on predicted user visual attention in dynamic scenes.
  • Multimedia Content Analysis: Supports analysis and optimization of video content by identifying salient regions over time.
  • Cognitive Science: Offers computational predictions of attentional allocation useful for studying visual attention mechanisms.

Methodology:

The model integrates appearance and motion via a dense residual network with dense residual cross-connections, applies a composite attention mechanism to learn multi-scale local attentions and global attention priors end-to-end, and employs a lightweight convGRU to model long-term temporal dependencies.

Topics

Details

Programming Languages:
Python
Added:
11/14/2019
Last Updated:
12/24/2020

Operations

Publications

Lai Q, Wang W, Sun H, Shen J. Video Saliency Prediction Using Spatiotemporal Residual Attentive Networks. IEEE Transactions on Image Processing. 2020;29:1113-1126. doi:10.1109/tip.2019.2936112. PMID:31449021.

PMID: 31449021
Funding: - Natural Science Foundation of Beijing Municipality: 4182056 - National Natural Science Foundation of China: 61602183