Repository logo
Log In(current)
  1. Home
  2. Colleges & Schools
  3. Graduate School
  4. Masters Theses
  5. Bridging HPC Data Gaps: Novel Tools for GPU Resource Monitoring and Scheduler Emulation
Details

Bridging HPC Data Gaps: Novel Tools for GPU Resource Monitoring and Scheduler Emulation

Date Issued
May 1, 2025
Author(s)
Ashworth, Walter Jay
Advisor(s)
Michela Taufer
Additional Advisor(s)
Michela Taufer
Jack Marquez
Catherine Schuman
Permanent URI
https://trace.tennessee.edu/handle/20.500.14382/35448
Abstract

The evolution of High-Performance Computing (HPC) systems is introducing ever greater complexity in resource management and performance monitoring, creating critical gaps in data availability. With their expanding significance in AI and data-intensive applications, GPUs now serve as pivotal components of HPC workloads. Simultaneously, the push toward exascale computing necessitates increasingly efficient and scalable scheduling policies. This thesis presents two novel tools designed to address the data availability challenges introduced by these ongoing shifts in modern HPC environments. The first tool bridges a crucial gap in per-job GPU resource monitoring for SLURM-managed clusters, which lack this support natively. By enabling detailed post-job analysis of GPU utilization metrics, it allows researchers to identify underutilization issues—ranging from configuration errors to algorithmic inefficiencies—while helping administrators accurately assess resource usage in planning future upgrades. The second tool, the Flux Emulator, builds upon a preliminary prototype and extends the capabilities of the Flux Framework, a cutting-edge resource management and scheduling system tailored for exascale HPC. Through the simulation of historical job workloads, the emulator enables the evaluation of various scheduling policies and strategies without the need for physical cluster resources. By illuminating the effects of different scheduler configurations on performance metrics such as job makespan, it empowers system software developers to refine algorithms and policies for more efficient utilization of emerging exascale systems. By providing detailed GPU resource monitoring within established scheduling systems and offering a scalable emulator for evaluating multiple scheduling policies, this thesis lays the groundwork for a more data-driven approach to HPC resource utilization and optimization. Together, these tools directly address the critical gaps in data availability and performance insight that arise as HPC environments continue to grow in complexity.

Subjects

HPC

schedulers

performance analysis

emulation

Disciplines
Other Computer Sciences
Degree
Master of Science
Major
Computer Science
File(s)
Thumbnail Image
Name

Thesis_2025_Jay_Final.pdf

Size

5.56 MB

Format

Adobe PDF

Checksum (MD5)

1b2b04d9d0ec1a29627f11af32a5535d


University Libraries

1015 Volunteer Boulevard
Knoxville, TN 37996
865-974-4351

Map & Directions
Donate to the Libraries
  • About
  • John C. Hodges Society
  • Speaking Volumes magazine
  • Outreach
  • Directory
  • Employment
  • Policies
  • Library Intranet
University of Tennessee power T logo

The University of Tennessee, Knoxville
Knoxville, Tennessee 37996
865-974-1000

Events
A-Z
Apply
Privacy
Map
Directory
Give to UT
Accessibility

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science