Graduation Year

2025

Document Type

Dissertation

Degree

D.B.A.

Degree Granting Department

Business Administration

Major Professor

Uday Murthy, Ph.D.

Co-Major Professor

Christian Koch, D.B.A.

Committee Member

Dahlia Robinson, Ph.D.

Committee Member

Jennifer Wolgemuth, Ph.D.

Keywords

Artificial Intelligence, Business Intelligence, Data Validation, Elaborated Action Design Research, Large Language Models, Vendor Reconciliation

Abstract

Organizations struggle with data quality problems that cost them millions of dollars annually. Manual data cleanup and standardizing, also known as reconciliation, can take skilled analysts days to complete and can still contain human errors. This dissertation addresses this challenge by developing and validating the Origin Data Management System, which automates processes that traditionally require days of manual work. The system achieved up to 100% accuracy in standardizing vendor names and identifying CEO information across real-world test datasets from Jefferson County Public Schools, Forbes, and Fortune Magazines. Processing time dropped from days to minutes—a 96-fold improvement over manual methods. These results emerged from over 40,000 systematic experiments testing different AI configurations, search strategies, and validation approaches using the Elaborated Action Design Research (eADR) methodology.

The research builds on recent advances in multi-agent systems, where multiple AI systems, such as Large Language Models (LLMs), work together to verify information. While previous studies demonstrated 85-100% success rates in detecting hallucinations in simple narrative text (Kwartler et al., 2024), this work addresses the more complex reality of organizational data, such as vendor names, financial records, and entity relationships, which appear in inconsistent formats across different systems.

Through rigorous testing, the research identified "context poisoning," a phenomenon where one AI agent's feedback inadvertently influences another agent to adopt plausible but incorrect answers rather than maintain independent verification. This discovery challenged prevailing assumptions about AI collaboration, showcasing that agents working independently could outperform agents that discussed their findings with each other. This discovery has significant implications for designing reliable AI systems for structured data validation.

The dissertation makes three primary contributions. First, it delivers a functional system that automates costly reconciliation processes and provides integrated analytics capabilities, directly improving the accuracy, timeliness, and accessibility of organizational data. Second, it introduces the LLM Buddy tool for comprehensive prompt capture and version control in AI-augmented research, ensuring that AI-assisted development processes can be documented, reproduced, and audited—addressing a critical gap in research methodology. Third, it presents the Conversational Forking methodology for non-linear problem-solving with Large Language Models, enabling researchers to explore solution alternatives systematically rather than following linear conversation paths.

Grounded in Wang & Strong's (1996) Data Quality Framework, the Origin system demonstrates measurable improvements across multiple data quality dimensions, including accuracy, believability, and accessibility. These findings provide both theoretical advancement in multi-agent system design and immediate practical value for organizations seeking to enhance their data quality management capabilities.

Share

COinS