Publication Date

2025

Document Type

Dissertation

Committee Members

Ashutosh Shivakumar, Ph.D. (Committee Chair); Sean Banerjee, Ph.D. (Committee Co-Chair); Kendall Goodrich, Ph.D. (Committee Member); Natasha Banerjee, Ph.D. (Committee Member); Thomas Wischgoll, Ph.D. (Committee Member)

Degree Name

Doctor of Philosophy (PhD)

Abstract

As AI-driven workloads accelerate the growth of cloud initiatives and spending, resource waste also increases due to persistent inefficiencies in cloud compute and infrastructure management. Overprovisioned resources and suboptimal configurations often lead to operational inefficiencies and unnecessary financial overhead. These challenges arise from the difficulty of anticipating resource demands in dynamic workloads and selecting suitable virtual machines to ensure optimal performance. Our research proposes a holistic, data-driven framework for managing cloud compute resources that reduces costs without compromising application performance. We integrate a predictive, model-driven, threshold-based autoscaling solution for cloud-native applications with an optimized instance right-sizing approach to select cost-effective instance types that meet performance requirements. Our autoscaling framework employs correlation analysis to eliminate the uncertainty of arbitrary metric selection by identifying relevant resource utilization or workload volume metrics that influence application performance, such as response time or latency. Using these selected metrics along with historical data, our framework evaluates multiple regression models and selects the most accurate and reliable one for predicting the minimum resource requirements needed to maintain application performance. The upper threshold value for autoscaling is derived from model predictions that balance cloud cost and application performance, while the lower threshold values are determined through simulation analysis to minimize inefficient scaling behavior. While autoscaling mitigates the overprovisioning of cloud resources under fluctuating workloads, the selection of appropriate cloud instances is often overlooked. To address this, guided by empirical analysis, the research introduces a cloud instance right-sizing strategy, that optimizes both cost and performance for cloud-hosted applications. Our approach considers business-imposed cost and performance constraints and applies Pareto optimization to identify instance types that achieve an optimal cost-performance balance. A weighted-sum method further refines the selection by prioritizing specific business objectives and determining the most suitable instance from the Pareto-optimal set. Experimental results validate the effectiveness of the proposed autoscaling and right-sizing frameworks in optimizing both application performance and resource costs for low-latency cloud applications.

Page Count

124

Department or Program

Department of Computer Science and Engineering

Year Degree Awarded

2025


Share

COinS