Publication Date

2024

Document Type

Dissertation

Committee Members

Lingwei Chen, Ph.D. (Advisor); Michael L. Raymer, Ph.D. (Committee Member); Tanvi Banerjee, Ph.D. (Committee Member); Josh Ash, Ph.D. (Committee Member)

Degree Name

Doctor of Philosophy (PhD)

Abstract

The digital landscape is ever-evolving. In recent years the amount of bot traffic, traffic generated by autonomous applications over the internet has increased significantly. Many bots perform useful and needed functions, however, malicious bots are known sources of both common and emerging security threats. Denial-of-Services (DoS), information theft, and credential stuffing have all been conducted by malicious software running on unknowingly infected machines. The dichotomy of useful bots operating in the same networks as malicious bots combined with novel bot attacks and an ever-increasing number of personal devices connecting to the Internet drives the need for continued advancement of malicious bot detection techniques. Graph Neural Networks, a class of neural networks that are enriched with structural and relationship information, have emerged as a possible advancement on existing bot detection methods. Unfortunately, GNNs have known limitations pertaining to extreme data imbalances and heterophily, two common attributes of bot datasets. Moreover, a lack of accurate and completely labeled bot data for training purposes further limits GNN performance. This dissertation outlines multiple novel approaches for addressing data imbalances, heterophily, and scarce data enabling the application of GNNs to malicious bot detection. First, I introduce the HOVER (Homophilic Oversampling Via Edge Removal) technique, an oversampling approach capable of mitigating heterophily in graphs and learning improved data representations. Next, I explore the utility of out-of-distribution detection (ODD) on imbalanced graphs. Two different ODD approaches are developed demonstrating the effectiveness of energy-based and isotropic ODD for bot detection. I supplement shortcomings not addressed in earlier ODD research by enhancing learning-free belief propagation schemes to boost detection performance. Finally, I present a selective graph rewiring methodology derived from the topological concept of curvature. This rewiring method identifies potentially problematic graph structures limiting the scope of graph rewiring process to preserve the original semantics of the graph while still improving GNN utility. I pair this rewiring approach with a label-informed similarity mechanism, tailoring the method to real-world scenarios with minimal training data that may not be entirely trustworthy. Experimental evaluation is conducted on publicly available collection of network traffic containing bot data. The results of these experiments demonstrate my methodologies achieve state-of-the-art performance under extremely challenging graph conditions not typically addressed by other research.

Page Count

122

Department or Program

Department of Computer Science and Engineering

Year Degree Awarded

2024

ORCID ID

0009-0004-5411-4799


Share

COinS