European Journal of Computer Science and Information Technology (EJCSIT)

data quality

Data Engineering: The Catalyst for Aviation Industry Transformation (Published)

The aviation industry is experiencing a transformative shift driven by data engineering innovations that optimize operations and enhance passenger experiences. As global air travel expands and consumer expectations evolve, airlines and airports increasingly rely on sophisticated data infrastructure to manage complex operations. Through real-world implementations at major aviation hubs, data engineering has revolutionized critical functions from baggage handling to aircraft maintenance. London Heathrow’s event-driven architecture for baggage management illustrates how real-time data processing eliminates historical pain points, while Lufthansa’s predictive maintenance system demonstrates how properly structured data pipelines enable effective artificial intelligence applications. Singapore Changi Airport’s implementation of graph-based data models for passenger flow optimization showcases the importance of selecting appropriate data modeling paradigms for specific problem domains. These successes contrast with cautionary examples where inadequate data quality undermined otherwise promising initiatives, highlighting data quality as a foundational requirement rather than a technical afterthought. The integration of batch and streaming capabilities, appropriate data model selection, and rigorous quality assurance represent defining characteristics of successful aviation data architectures that deliver measurable operational improvements and enhanced passenger experiences. The economic impact of these implementations extends beyond operational efficiencies to include enhanced revenue opportunities, improved asset utilization, and strengthened competitive positioning in an increasingly digital marketplace. Aviation entities that fail to embrace modern data engineering principles risk falling behind as the gap between data-driven organizations and traditional operators continues to widen. The remarkable improvements in passenger satisfaction metrics and operational key performance indicators demonstrate that data engineering has moved from a supporting technical function to a strategic business capability that directly influences both the bottom line and customer loyalty.

Keywords: Predictive Maintenance, aviation analytics, data engineering, data quality, event-driven architecture, graph databases, passenger experience, real-time processing

Developing an AI-Driven Anomaly Detection System for Cloud Data Pipelines: Minimizing Data Quality Issues by 40% (Published)

This article presents an innovative AI-driven anomaly detection system designed specifically for cloud data pipelines, addressing the critical challenge of ensuring data quality at scale in increasingly complex cloud-native architectures. As organizations transition from monolithic to microservices-based approaches, traditional rule-based monitoring methods have become insufficient for detecting the multitude of potential quality issues that arise across distributed infrastructures. Our system employs a multi-layered architecture that combines statistical profile modeling, deep learning techniques, and semantic anomaly detection to identify subtle pattern deviations across diverse data environments. By leveraging ensemble learning approaches, temporal pattern recognition, and adaptive thresholding, the system demonstrates significant improvements in reducing data quality incidents, minimizing detection latency, and lowering false positive rates. The implementation methodology incorporates specialized transformer-based neural architectures that operate across both streaming analytics and batch-oriented data lake environments. Case studies across multiple industry deployments, particularly in financial services, validate the system’s effectiveness in enhancing operational efficiency, reducing compliance risks, and improving decision-making processes while maintaining adaptability across heterogeneous data infrastructures

Keywords: Cloud data pipelines, anomaly detection, data quality, machine learning, predictive analytics, self-healing systems

Data Quality, Feature Engineering, and Model Reliability in Large-Scale AI Multi-Agentic Systems (Published)

Large-scale artificial intelligence (AI) systems increasingly operate not as a single monolithic model but as a population of interacting, specialized agents separate models or decision-making components responsible for functions such as pricing, fraud detection, ranking, routing, and customer support that share underlying data infrastructure and, in many cases, influence one another’s inputs and outputs. This article synthesizes peer-reviewed literature on data quality assurance, feature engineering, and model reliability to examine how these three concerns interact once an AI system is decomposed into multiple cooperating agents operating at scale. Seventeen primary sources are reviewed, spanning foundational work on technical debt in machine learning systems, multi-dimensional data quality frameworks, scalable and automated data quality verification, data lifecycle management, empirically grounded data-management taxonomies, production-readiness testing rubrics, scalable and automated feature engineering, hyperparameter optimization, organizational workflow studies, a production-scale ML platform, concept drift, large-scale academic surveys of ML testing, systematic reviews of industrial ML challenges, and fault-tolerant cooperative control of multi-agent systems. Drawing on this literature, the article proposes a conceptual framework linking data quality assurance, feature engineering, model training and reliability testing, and deployment to an agent population, closed by a cross-agent monitoring and feedback loop. The review finds that data quality problems and model reliability failures do not remain confined to the agent in which they originate: because agents in a large-scale AI system typically share upstream data sources, feature pipelines, or downstream state, a defect introduced at the data or feature layer of one agent can propagate through the interactions between agents, producing system-level reliability failures that are not visible from the perspective of any single agent’s test suite. Ensuring reliability in such systems therefore requires treating data quality, feature engineering, and testing as cross-cutting, system-wide concerns rather than as properties to be verified independently within each agent.

 

Keywords: ML testing, concept drift, data quality, fault-tolerance, feature engineering, large-scale AI, model reliability, multi-agent systems, technical debt

Scroll to Top

Don't miss any Call For Paper update from EA Journals

Fill up the form below and get notified everytime we call for new submissions for our journals.