Publication Date

2025

Document Type

Thesis

Committee Members

Lingwei Chen, Ph.D. (Advisor); Krishnaprasad Thirunarayan, Ph.D. (Committee Member); Michael Raymer, Ph.D. (Committee Member)

Degree Name

Master of Science (MS)

Abstract

Modern machine learning (ML) models rely on large amounts of high-quality labeled data to achieve optimal performance. However, in many real-world domains, such as cyber security, acquiring sufficient labeled data is often infeasible due to cost, privacy concerns, and the rapid evolution of underlying phenomena. This challenge underscores the importance of learning under data scarcity. This thesis addresses this challenge by proposing distinct, modality-specific techniques for text and graph domains, which allow models to generalize effectively with minimal data. For text classification task, we incorporate distilled rationales from large language models and adversarial perturbations into the input space to improve data efficiency and reduce the need for extensive fine-tuning. For graph-structured data, we introduce a novel negative distillation framework, where adversarially identified high-confidence nodes are per turbed to generate negative training signals that improve class separation in graph neural networks. Both approaches are implemented within a dual-branch learning architecture, where a "teacher" branch encodes structured reasoning or negative knowledge to guide a more compact "student" model. By conducting experiments on real-world datasets, we show that the proposed reasoning and negative distillation strategies significantly improve performance in data-limited settings, often surpassing state-of-the-art baselines. This thesis highlights the potential of targeted reasoning, adversarial learning, and negative distillation techniques to extend the applicability of machine learning to domains constrained by limited supervision.

Page Count

64

Department or Program

Department of Computer Science and Engineering

Year Degree Awarded

2025


Share

COinS