Repository logo
Log In(current)
  1. Home
  2. Colleges & Schools
  3. Graduate School
  4. Masters Theses
  5. Koopman-Inspired Proximal Policy Optimization (KIPPO)
Details

Koopman-Inspired Proximal Policy Optimization (KIPPO)

Date Issued
August 1, 2024
Author(s)
Cozma, Andrei  
Advisor(s)
Hairong Qi
Additional Advisor(s)
Hairong Qi
Catherine Schuman
Dan Wilson
Amir Sadovnik
Permanent URI
https://trace.tennessee.edu/handle/20.500.14382/33130
Abstract

Reinforcement Learning (RL) has made significant strides in various domains, yet developing effective control policies for environments with complex, nonlinear dynamics remains a challenge, particularly for policy gradient methods. These methods often struggle due to high-variance in gradient estimates, non-convex optimization landscapes, and sample inefficiency, resulting in unstable learning, suboptimal policies, and trade-offs between performance and reproducibility. The quest for more robust, stable, and effective methods has led to numerous innovations and remains a critical area of research. Proximal Policy Optimization (PPO) has gained popularity in recent years due to its balance in performance, training stability, and computational efficiency. In contrast with their nonlinear counterparts, linear systems are simpler, more predictable, and easier to analyze. Koopman Theory has emerged as a powerful framework for studying nonlinear systems through a globally-linear operator that acts on a higher-dimensional space of measurement functions. Combining these two ideas, Koopman-Inspired Proximal Policy Optimization (KIPPO) extends PPO to learn a simplifying representation of the underlying system's dynamics while retaining essential features for effective policy learning. This is achieved through a Koopman-approximation auxiliary network and carefully designed constraints that enable balancing the complexity of latent dynamics. Results demonstrate improvements over the PPO baseline with 8-60% increased performance while reducing variability by up to 91% when evaluated on diverse continuous control tasks. The study also examines the effects and interactions of key hyperparameters and the impacts of individual loss components through an ablation study, providing a comprehensive analysis of the approach.

Subjects

Reinforcement Learnin...

Policy Optimization

Policy Gradients

Koopman Operator Theo...

Representation Learni...

Dynamical Systems

Disciplines
Artificial Intelligence and Robotics
Theory and Algorithms
Degree
Master of Science
Major
Computer Science
File(s)
Thumbnail Image
Name

acozma_MSthesis_final.pdf

Size

1.54 MB

Format

Adobe PDF

Checksum (MD5)

83ef0aca1397dde31bb1004eed52540e


University Libraries

1015 Volunteer Boulevard
Knoxville, TN 37996
865-974-4351

Map & Directions
Donate to the Libraries
  • About
  • John C. Hodges Society
  • Speaking Volumes magazine
  • Outreach
  • Directory
  • Employment
  • Policies
  • Library Intranet
University of Tennessee power T logo

The University of Tennessee, Knoxville
Knoxville, Tennessee 37996
865-974-1000

Events
A-Z
Apply
Privacy
Map
Directory
Give to UT
Accessibility

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science