review · Machine Learning with Applications
Serverless computing has changed the shape of cloud architecture by removing the load of infrastructure management and enabling event-driven, stateless function execution with dynamic resource allocation. This model delivers notable benefits in scalability, cost efficiency, and energy savings, yet faces challenges such as cold start latency, workload diversity, and complex resource scheduling. This study examines core principles of serverless computing, contrasting it with traditional cloud models, and evaluates key performance metrics, including latency, execution time, throughput, resource utilization, energy consumption, and cost-performance trade-offs. Performance metrics are analyzed through both infrastructure and application perspectives by considering platform heterogeneity, virtualization overheads, hardware diversity, deployment package size, workflow complexity, and data access patterns. Following PRISMA guidelines, this systematic literature review (SLR) analyzed 1384 initial records across six databases, ultimately selecting 80 high-quality studies published between 2020 and 2025. The analysis identifies seven primary performance metrics, with Latency accounting for 31% of 149 metric mentions, followed by Cost (24%) and Execution Time (15%). A 17-fold growth in serverless performance publications was observed from 2020 to 2024. Optimization strategies are categorized into dynamic scaling, heuristic and greedy algorithms, cost-aware provisioning, and AI-driven approaches, including machine learning, deep learning, and deep reinforcement learning. Deep reinforcement learning emerges as the leading ML paradigm (10% of studies), indicating a shift toward adaptive, intelligent optimization. The role of service-level agreements and user requirements is highlighted in the context of heterogeneous, rapidly changing workloads. The objective of this review is to equip researchers with valuable insights into the methodologies and technical landscape of serverless cloud computing, making it a key reference in the field.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.mlwa.2026.100994
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.