Repository logo
Log In(current)
  1. Home
  2. Colleges & Schools
  3. Graduate School
  4. Doctoral Dissertations
  5. Attention Mechanism for Recognition in Computer Vision
Details

Attention Mechanism for Recognition in Computer Vision

Date Issued
August 1, 2019
Author(s)
Rahimpour, Alireza
Advisor(s)
Hairong Qi
Additional Advisor(s)
Jens Gregor
Russell Zaretzki
Seddik M. Djouadi
Permanent URI
https://trace.tennessee.edu/handle/20.500.14382/26825
Abstract

It has been proven that humans do not focus their attention on an entire scene at once when they perform a recognition task. Instead, they pay attention to the most important parts of the scene to extract the most discriminative information. Inspired by this observation, in this dissertation, the importance of attention mechanism in recognition tasks in computer vision is studied by designing novel attention-based models. In specific, four scenarios are investigated that represent the most important aspects of attention mechanism. First, an attention-based model is designed to reduce the visual features' dimensionality by selectively processing only a small subset of the data. We study this aspect of the attention mechanism in a framework based on object recognition in distributed camera networks. Second, an attention-based image retrieval system (i.e., person re-identification) is proposed which learns to focus on the most discriminative regions of the person's image and process those regions with higher computation power using a deep convolutional neural network. Furthermore, we show how visualizing the attention maps can make deep neural networks more interpretable. Third, a model for estimating the importance of the objects in a scene based on a given task is proposed. More specifically, the proposed model estimates the importance of the road users that a driver (or an autonomous vehicle) should pay attention to in a driving scenario in order to have safe navigation. In this scenario, the attention estimation is the final output of the model. Fourth, an attention-based module and a new loss function in a meta-learning based few-shot learning system is proposed in order to incorporate the context of the task into the feature representations of the samples and increasing the few-shot recognition accuracy. In this dissertation, we showed that attention can be multi-facet and studied the attention mechanism from the perspectives of feature selection, reducing the computational cost, interpretable deep learning models, task-driven importance estimation, and context incorporation. Through the study of four scenarios, we further advanced the field of where ''attention is all you need''.

Subjects

Attention mechanism

Machine learning

Deep Learning

Computer Vision

Meta-learning

Image retrieval and m...

Person Re-identificat...

Object detection

Feature selection.

Degree
Doctor of Philosophy
Major
Electrical Engineering
File(s)
Thumbnail Image
Name

utk.ir.td_12362.pdf

Size

31.56 MB

Format

Adobe PDF

Checksum (MD5)

bb67da796f3a358851c188a0b42d0dc8


University Libraries

1015 Volunteer Boulevard
Knoxville, TN 37996
865-974-4351

Map & Directions
Donate to the Libraries
  • About
  • John C. Hodges Society
  • Speaking Volumes magazine
  • Outreach
  • Directory
  • Employment
  • Policies
  • Library Intranet
University of Tennessee power T logo

The University of Tennessee, Knoxville
Knoxville, Tennessee 37996
865-974-1000

Events
A-Z
Apply
Privacy
Map
Directory
Give to UT
Accessibility

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science