Repository logo
Log In(current)
  1. Home
  2. Colleges & Schools
  3. Graduate School
  4. Doctoral Dissertations
  5. Binary Representation Learning for Large Scale Visual Data
Details

Binary Representation Learning for Large Scale Visual Data

Date Issued
May 12, 2018
Author(s)
Liu, Liu  
Advisor(s)
Hairong Qi
Additional Advisor(s)
Jens Gregor
Husheng Li
Russell L. Zaretzki
Permanent URI
https://trace.tennessee.edu/handle/20.500.14382/26196
Abstract

The exponentially growing modern media created large amount of multimodal or multidomain visual data, which usually reside in high dimensional space. And it is crucial to provide not only effective but also efficient understanding of the data.In this dissertation, we focus on learning binary representation of visual dataset, whose primary use has been hash code for retrieval purpose. Simultaneously it serves as multifunctional feature that can also be used for various computer vision tasks. Essentially, this is achieved by discriminative learning that preserves the supervision information in the binary representation.By using deep networks such as convolutional neural networks (CNNs) as backbones, and effective binary embedding algorithm that is seamlessly integrated into the learning process, we achieve state-of-the art performance on several settings. First, we study the supervised binary representation learning problem by using label information directly instead of pairwise similarity or triplet loss. By considering images and associated textual information, we study the cross-modal representation learning. CNNs are used in both image and text embedding, and we are able to perform retrieval and prediction across these modalities. Furthermore, by utilizing unlabeled images from a different domain, we propose to use adversarial learning to connect these domains. Finally, we also consider progressive learning for more efficient learning and instance-level representation learning to provide finer granularity understanding. This dissertation demonstrates that binary representation is versatile and powerful under various circumstances with different tasks.

Subjects

Image Hashing

Discriminative Learni...

Natural Language Proc...

Adversarial Learning

Object Detection

Computer Vision

Degree
Doctor of Philosophy
Major
Computer Engineering
File(s)
Thumbnail Image
Name

utk.ir.td_851.pdf

Size

3.11 MB

Format

Adobe PDF

Checksum (MD5)

73acedb6ff81a5cc2df248d95387b1f3


University Libraries

1015 Volunteer Boulevard
Knoxville, TN 37996
865-974-4351

Map & Directions
Donate to the Libraries
  • About
  • John C. Hodges Society
  • Speaking Volumes magazine
  • Outreach
  • Directory
  • Employment
  • Policies
  • Library Intranet
University of Tennessee power T logo

The University of Tennessee, Knoxville
Knoxville, Tennessee 37996
865-974-1000

Events
A-Z
Apply
Privacy
Map
Directory
Give to UT
Accessibility

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science