Efficient Storage and Query Optimization for Large-Scale Data Sets in Distributed Architectures
Main Article Content
Abstract
With the rapid development of the big data industry, data volume across various industries has exploded, and large-scale datasets at PB and EB levels have become mainstream objects for data processing. Relying on core theories of distributed storage and query, this paper constructs an integrated collaborative optimization system covering storage, query and caching. Storage efficiency is improved by designing an adaptive dynamic sharding strategy and intelligent multi-replica placement policy, building a tiered storage architecture for hot and cold data, and optimizing storage encoding and compression mechanisms. Query overhead is reduced by reconstructing query execution plans based on cost models, pushing down operators, and optimizing cross-node transmission. A collaborative global-local indexing framework and a multi-level caching linkage mechanism are established to further boost query response performance. A standardized distributed experimental cluster is deployed, and multi-dimensional comparative experiments are conducted to verify the performance of the proposed optimization scheme. Experimental results demonstrate that the proposed optimization solution can effectively cut storage redundancy overhead, improve cluster resource utilization, drastically reduce query latency for large-scale data, and raise concurrent throughput. It can provide theoretical support and engineering practice references for distributed system optimization in industrial big data, internet massive data processing and other scenarios.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
S. Patil and S. Sangam, “An Exhaustive Survey of Big Data Storage Reduction Techniques,” Cureus J. Comput. Sci., vol. 2, no. 1, p. 3518, 2025, doi: 10.7759/S44389-025-03518-3.
S. Guan, C. Zhang, Y. Wang, et al., “Hadoop-based secure storage solution for big data in cloud computing environment,” Digit. Commun. Netw., vol. 10, no. 1, pp. 227–236, 2024, doi: 10.1016/J.DCAN.2023.01.014.
G. Jifu, H. Chunlin, H. Jinliang, “A Scalable Computing Resources System for Remote Sensing Big Data Processing Using GeoPySpark Based on Spark on K8s,” Remote Sens., vol. 14, no. 3, p. 521, 2022, doi: 10.3390/RS14030521.
B. J. Kim, “Verification of scalability and compatibility of TiDB for DBMS migration,” J. Digit. Contents Soc., 2022, doi: 10.9728/dcs.20 22.23.11.2299.
S. Huijun and R. Ruonan, “Scalable distributed RDFS reasoning using MapReduce and Bigtable,” Shanghai Jiao Tong Univ., China, 2013, doi: 10.1117/12.2010731.
A. Sale and N. Enas, “Survey on a Google File System (GFS),” Int. J. Comput. Appl., vol. 181, no. 28, pp. 9–16, 2018, doi: 10.5120/ijca201891 8084.
J. C. C, H. Peter, H. Wilson, et al., “Spanner: Google’s Globally Distributed Database,” ACM Trans. Comput. Syst., vol. 31, no. 3, pp. 1–22, 2013, doi: 10.1145/2518037.2491245.
H. Fei, Y. Chaowei, J. Yongyao, et al., “A hierarchical indexing strategy for optimizing Apache Spark with HDFS to efficiently query big geospatial raster data,” Int. J. Digit. Earth, vol. 13, no. 3, pp. 410–428, 2020, doi: 10.1080/17538947.2018.1523957.
H. Su, J. Li, L. Guo, et al., “Massive Data HBase Storage Method for Electronic Archive Management,” Int. J. Netw. Manage., vol. 35, no. 1, p. e2308, 2024, doi: 10.1002/NEM.2308.
K. Yan, J. Jia, L. Lv, et al., “OceanBase Database Health Status Prediction System Based on the HPSA-LSTM Model,” Acad. J. Comput. Inf. Sci., vol. 8, no. 8, 2025, doi: 10.25236/AJCIS.2025.080809.
C. Yang, D. Qingyang, W. Dan, et al., “TIDB: a comprehensive database of trained immunity,” Database, 2021, doi: 10.1093/DATABASE/BAAB041.
R. M. Kaseb, Haytamy, et al., “Distributed query optimization strategies for cloud environment,” J. Data Inf. Manage., pp. 1–9, 2021, doi: 10.1007/S42488-021-00057-Z.
L. Jongtao, K. Byounghoon, L. Hyeonbyeong, et al., “An Efficient Distributed SPARQL Query Processing Scheme Considering Communication Costs in Spark Environments,” Appl. Sci., vol. 12, no. 1, p. 122, 2021, doi: 10.3390/APP12010122.
Y. Du, Z. Ding, Z. Cai, et al., “Join query optimization in distributed database based on multi-source mating selection evolutionary algorithm,” Cluster Comput., vol. 28, no. 5, p. 281, 2025, doi: 10.1007/S10586-024-04905-6.
K. Fan, “Research on distributed query optimization technology based on index,” M.S. thesis, Univ. Electron. Sci. Technol. China, Chengdu, China, 2024, doi: 10.27005/d.cnki.gdzku.2024.005185.
Z. Aoujil, M. Hanine, Z. Baybih, et al., “NoSQL Vis: A unified visualization and exploration platform for heterogeneous NoSQL datastores,” SoftwareX, vol. 35, p. 102778, 2026, doi: 10.1016/J.SOFTX.2026.102778.
R. R. E. F. Ferreira and N. D. R. Fidalgo, “A Performance Analysis of Hybrid and Columnar Cloud Databases for Efficient Schema Design in Distributed Data Warehouse as a Service,” Data, vol. 9, no. 8, p. 99, 2024, doi: 10.3390/DATA9080099.