Graduation Year

2026

Document Type

Thesis

Degree

M.S.C.S.

Degree Name

MS in Computer Science (M.S.C.S.)

Degree Granting Department

Computer Science and Engineering

Major Professor

Sriram D. Chellappan, Ph.D.

Committee Member

Amit Seal Ami, Ph.D.

Committee Member

Anowarul Kabir, Ph.D.

Keywords

HCI, LLM, NLP, QLoRA

Abstract

The detection of sensitive content has become a critical task in natural language processing due to the widespread use of online platforms and the social harm caused by it. Even though current methods have advanced significantly, many of them mainly rely on surface-level text and neglect to take into consideration contextual and cultural elements that are crucial for correctly assessing sensitivity of any content to a particular individual. Due to this, implicit offense, sarcasm, and culturally grounded expressions frequently cause problems for such systems. To enable culturally grounded sensitive content detection, this thesis presents Sensi-Train, a context-rich dataset and modeling framework, that is in some sense, culturally personalized. Three essential elements are annotated on each data instance in Sensi-Train: the text, a target culture, and a contextual description. This study assesses four modeling paradigms to determine the influence of contextual information: PlainBERT, the already existing BERT model, Plain LLM, a zero-shot large language model assessed using both Mistral-7B-Instruct and GPT-3.5-Turbo, SensiBERT, a context-aware transformer refined on Sensi-Train, and SensiLLM, a context-aware large language model refined on Sensi-Train utilizing structured prompts. While context-aware models are trained and tested using the entire feature set, zero-shot models are assessed using text-only input. A held-out test set is used for the experiments. The findings show that context-aware fine-tuned models outperform inference-only and zero-shot models. SensiLLM achieves the best performance, while SensiBERT significantly outperforms baselines. These results offer compelling empirical proof that effective sensitive content detection requires explicit contextual and cultural supervision. To summarize, this thesis outlines the drawbacks of text-only methods and shows how context-aware learning can be used to create NLP systems that are socially conscious and grounded.

Share

COinS