European Journal of Computer Science and Information Technology (EJCSIT)

Apache Kafka

Streaming Data Pipelines and AI-Driven Cleansing: A Financial Institution’s Journey to Enhanced Risk Assessment (Published)

Financial institutions face mounting challenges in processing vast transactional datasets while maintaining regulatory compliance and detecting fraudulent activities. This article examines how a global banking enterprise implemented an integrated data architecture utilizing AWS Aurora and Redshift to consolidate disparate transactional systems. The implementation resulted in significant reduction of risk assessment timeframes while enhancing analytical capabilities. Apache Kafka-powered streaming pipelines provided the foundation for real-time fraud detection mechanisms, seamlessly supporting compliance monitoring across multiple jurisdictions. The migration process incorporated AI-driven data cleansing protocols to maintain data integrity and ensure analytical accuracy. Particularly noteworthy was the development of scalable analytical models designed specifically to process volatile market data during periods of financial uncertainty. The architectural solutions described demonstrate how strategic data engineering investments enable financial institutions to navigate complex regulatory landscapes while simultaneously improving operational efficiency. These findings contribute to understanding how modern data infrastructure can transform risk assessment capabilities in the financial services sector.

 

Keywords: AWS aurora, Apache Kafka, Financial data engineering, Fraud Detection, regulatory compliance, risk analytics

Building an End-to-End Reconciliation Platform for Accurate B2B Payments in New-Age Fintech Distributed Ecosystems: A Case Study using Microservices and Kafka (Published)

The evolution of fintech ecosystems toward distributed architectures and microservices has revolutionized financial services by providing unprecedented scalability and flexibility. However, these advancements introduce significant complexities in B2B payment reconciliation processes where precision is critical. This article presents a comprehensive framework for an end-to-end reconciliation platform powered by Apache Kafka for real-time event streaming within microservices-based environments. The solution addresses key challenges including data consistency, transaction integrity, eventual consistency, distributed transactions, error detection, scalability, and timeliness to ensure accurate payment reconciliation during each pay cycle. Through a detailed architectural analysis featuring data collectors, matching engines, exception handlers, and reporting modules, the article explores how event sourcing, CQRS patterns, and idempotent processing can be leveraged to build robust reconciliation systems. Technical implementation considerations spanning horizontal scaling, performance optimization, and security controls provide practical guidance for deploying these systems in production environments. This framework offers valuable insights for fintech practitioners and researchers seeking to implement reliable reconciliation solutions in complex distributed payment ecosystems.

Keywords: Apache Kafka, distributed systems, event-driven architecture, microservices, payment reconciliation

Scalable Real-Time Data Pipelines for AI and Machine Learning–Driven Enterprise Systems (Published)

The growth of enterprise data volumes across the 2000s and 2010s pushed traditional batch-oriented data processing infrastructures past their practical limits, motivating a sustained shift toward distributed, stream-based architectures capable of supporting real-time analytics and machine learning (ML). This article synthesizes foundational and applied research published between 2001 and 2019 on distributed batch processing, distributed structured and key-value storage, early continuous query engines, in-memory cluster computing, micro-batch and internet-scale stream processing, log-based messaging, and unified batch/streaming programming models, to examine how scalable real-time data pipelines can be designed to support artificial intelligence (AI) and ML-driven enterprise systems. Ten distinct systems and thirteen primary sources are reviewed in depth. Drawing on this literature, the article proposes a five-layer architectural framework ingestion, stream processing, batch/model training, durable storage, and analytics/serving and evaluates the quantitative performance data, scalability mechanisms, fault-tolerance strategies, and enterprise implementation challenges reported across these sources. Reported figures include Google’s documented execution of roughly 100,000 MapReduce jobs per day, processing more than twenty petabytes of data daily; Amazon’s Dynamo latency service objective of sub-300-millisecond response at the 99.9th percentile; the 0.5-to-2-second target latency of the D-Streams micro-batch model; and the adoption of Storm by more than sixty production organizations by 2014. The review concludes that horizontally scalable, log-based messaging, combined with fault-tolerant, in-memory and micro-batch computation and durable, replicated storage, constituted the technical foundation that made real-time, ML-driven enterprise analytics feasible within this period, and that this layered architecture continues to underpin modern enterprise AI infrastructure

Keywords: Apache Kafka, Apache Spark, Bigtable, Dataflow Model, Distributed Computing, Dynamo, MapReduce, MillWheel, Real-time data pipelines, enterprise artificial intelligence, fault-tolerance, scalability, stream processing

Scroll to Top

Don't miss any Call For Paper update from EA Journals

Fill up the form below and get notified everytime we call for new submissions for our journals.