Publication Date
2025
Document Type
Dissertation
Committee Members
Ashutosh Shivakumar, Ph.D. (Committee Chair); Sean Banerjee, Ph.D. (Committee Co-Chair); Kendall Goodrich, Ph.D. (Committee Member); Natasha Banerjee, Ph.D. (Committee Member); Thomas Wischgoll, Ph.D. (Committee Member)
Degree Name
Doctor of Philosophy (PhD)
Abstract
As AI-driven workloads accelerate the growth of cloud initiatives and spending, resource waste also increases due to persistent inefficiencies in cloud compute and infrastructure management. Overprovisioned resources and suboptimal configurations often lead to operational inefficiencies and unnecessary financial overhead. These challenges arise from the difficulty of anticipating resource demands in dynamic workloads and selecting suitable virtual machines to ensure optimal performance. Our research proposes a holistic, data-driven framework for managing cloud compute resources that reduces costs without compromising application performance. We integrate a predictive, model-driven, threshold-based autoscaling solution for cloud-native applications with an optimized instance right-sizing approach to select cost-effective instance types that meet performance requirements. Our autoscaling framework employs correlation analysis to eliminate the uncertainty of arbitrary metric selection by identifying relevant resource utilization or workload volume metrics that influence application performance, such as response time or latency. Using these selected metrics along with historical data, our framework evaluates multiple regression models and selects the most accurate and reliable one for predicting the minimum resource requirements needed to maintain application performance. The upper threshold value for autoscaling is derived from model predictions that balance cloud cost and application performance, while the lower threshold values are determined through simulation analysis to minimize inefficient scaling behavior. While autoscaling mitigates the overprovisioning of cloud resources under fluctuating workloads, the selection of appropriate cloud instances is often overlooked. To address this, guided by empirical analysis, the research introduces a cloud instance right-sizing strategy, that optimizes both cost and performance for cloud-hosted applications. Our approach considers business-imposed cost and performance constraints and applies Pareto optimization to identify instance types that achieve an optimal cost-performance balance. A weighted-sum method further refines the selection by prioritizing specific business objectives and determining the most suitable instance from the Pareto-optimal set. Experimental results validate the effectiveness of the proposed autoscaling and right-sizing frameworks in optimizing both application performance and resource costs for low-latency cloud applications.
Page Count
124
Department or Program
Department of Computer Science and Engineering
Year Degree Awarded
2025
Copyright
Copyright 2025, all rights reserved. My ETD will be available under the "Fair Use" terms of copyright law.
