MARATTO

article · Transactions in GIS

Leveraging Space Filling Curves for Efficient Storage and Processing of Spatial Data in the Data LakeHouse

Abstract

ABSTRACT The rapid growth of data‐driven decision making has led to an increasing need for efficient and scalable data processing architectures. Data Lakehouse has emerged as a solution for nowadays data workload needs. It's a paradigm that combines the benefits of data lakes and data warehouses and provides a unified platform for a variety of operational workloads, including business intelligence, analytics, data science, and AI. In this paper, we present the concept of space filling curves, particularly the three most known curves, namely: Z‐order, Hilbert, and Gray code curves. We investigate how they improve both data locality and access patterns in the context of spatial big data in Data Lakehouse. This improvement has a direct impact on storage optimization and query performance. As a practical use case, the paper covers the implementation of the Z‐order curve, along with a discussion of the main algorithm's components and how it fits into the current Data Lakehouse system. We show that these optimizations significantly enhance the capabilities of Data Lakehouse architectures, ensuring better resource utilization across various spatial workloads.

Research topics

  • Advanced Database Systems and Queries
  • Data Quality and Management
  • Data Management and Algorithms

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1111/tgis.70137

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.