article · IEEE Access
Storing massive volumes of data in a single facility is increasingly impractical, leading organisations to distribute information across multiple geographic locations with or without data replication. However, analysing information spread across these distant locations poses significant technical challenges. Two data distribution approaches address this issue using the Random Sample Partition model, which breaks large datasets into smaller random sample blocks. In the non-replicated approach, representative sample blocks are gathered from various facilities and brought to a single location for processing. In the replicated approach, critical sample blocks are pre-distributed across multiple facilities, enabling analysis to occur entirely within any single centre without requiring inter-facility data transfers. Simulation across a local facility and four commercial cloud centres in North America, Asia, and Australia demonstrates the practical performance of both strategies.
Modern digital enterprises generate unprecedented volumes of information and must disperse their storage globally for resilience and capacity. Moving massive datasets across international networks for routine processing causes high latency and network costs. Providing structured methods to analyse statistically representative samples across distributed sites helps organisations extract insights rapidly while avoiding prohibitive network strain and operational overheads.
This work applies directly to multinational enterprises and cloud platform operators managing multi-region storage architectures. By reducing or eliminating heavy cross-centre data transfers during analytics tasks, the methodology enables cost and time savings. Having been evaluated through simulations on global AWS infrastructure across three continents, the underlying concepts appear to be applied research, requiring integration into existing commercial database engines before market deployment.
AI-generated from the published abstract. Always read the original work before citing.
As the volume of data grows rapidly, storing big data in a single data center is no longer feasible. Hence, companies have developed two scenarios to store their big data in multiple data centers. In the first scenario, the company's big data are distributed in multiple data centers without data replication. In the second scenario, data are also stored in multiple data centers but important data are replicated in these data centers to increase data safety and availability. However, in these scenarios, analyzing big data distributed in multiple data centers becomes a challenging task. In this paper, we propose two data distribution strategies to support big data analysis across geo-distributed data centers. In these strategies, we use the recent Random Sample Partition data model to convert big data into sets of random sample data blocks and distribute these data blocks into multiple data centers either without replication or with replication. In analyzing big data in multiple data centers without replication, we randomly select samples of data blocks from multiple data centers and download the sample data blocks to one data center for analysis. In the second strategy with replication of data blocks, we can analyze big data on any data center by randomly selecting a sample of data blocks replicated from other data centers. This strategy avoids data transformation between data centers. We demonstrate the performance of the two strategies in big data analysis by using simulation results produced on one local data center and four AWS data centers in North America, Asia, and Australia.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/access.2020.3027675
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.