European Journal of Computer Science and Information Technology (EJCSIT)

Dynamo

Scalable Real-Time Data Pipelines for AI and Machine Learning–Driven Enterprise Systems (Published)

The growth of enterprise data volumes across the 2000s and 2010s pushed traditional batch-oriented data processing infrastructures past their practical limits, motivating a sustained shift toward distributed, stream-based architectures capable of supporting real-time analytics and machine learning (ML). This article synthesizes foundational and applied research published between 2001 and 2019 on distributed batch processing, distributed structured and key-value storage, early continuous query engines, in-memory cluster computing, micro-batch and internet-scale stream processing, log-based messaging, and unified batch/streaming programming models, to examine how scalable real-time data pipelines can be designed to support artificial intelligence (AI) and ML-driven enterprise systems. Ten distinct systems and thirteen primary sources are reviewed in depth. Drawing on this literature, the article proposes a five-layer architectural framework ingestion, stream processing, batch/model training, durable storage, and analytics/serving and evaluates the quantitative performance data, scalability mechanisms, fault-tolerance strategies, and enterprise implementation challenges reported across these sources. Reported figures include Google’s documented execution of roughly 100,000 MapReduce jobs per day, processing more than twenty petabytes of data daily; Amazon’s Dynamo latency service objective of sub-300-millisecond response at the 99.9th percentile; the 0.5-to-2-second target latency of the D-Streams micro-batch model; and the adoption of Storm by more than sixty production organizations by 2014. The review concludes that horizontally scalable, log-based messaging, combined with fault-tolerant, in-memory and micro-batch computation and durable, replicated storage, constituted the technical foundation that made real-time, ML-driven enterprise analytics feasible within this period, and that this layered architecture continues to underpin modern enterprise AI infrastructure

Keywords: Apache Kafka, Apache Spark, Bigtable, Dataflow Model, Distributed Computing, Dynamo, MapReduce, MillWheel, Real-time data pipelines, enterprise artificial intelligence, fault-tolerance, scalability, stream processing

Scroll to Top

Don't miss any Call For Paper update from EA Journals

Fill up the form below and get notified everytime we call for new submissions for our journals.