{STRATA2.0}: A Serverless Middleware for Machine Learning Training
Serverless computing has gained increasing interest in recent years for enabling large-scale machine learning tasks. However, training a machine learning model in a serverless setting is a complex task and several challenges need to be addressed particularly in data distribution, result aggregation, resource heterogeneity, failures, container ephemerality, network and execution cost. These difficulties stem from the inherent complexity of distributed computation and the coordination demands of the machine learning algorithms. We propose STRATA2.0, a serverless middleware for Machine Learning training in serverless environments. STRATA2.0 provides a comprehensive suite of mechanisms designed to support efficient training of machine learning models on serverless infrastructures and address key challenges related to efficient data communication, coordination and synchronization, scalable and time efficient training of ML models using heterogeneous containers. Our extensive experimental results demonstrate that STRATA2.0 achieves the same level of accuracy with fewer data points, is on average three times faster in training time compared to centralized approaches, reduces energy consumption by up to 50\%, and remains resilient when up to 60\% of training instances fail.
- Published in:
IEEE Transactions on Parallel and Distributed Systems - Type:
Article - Authors:
- Year:
2026 - Source:
https://doi.org/10.1109/TPDS.2026.3702338
Citation information
: {STRATA2.0}: A Serverless Middleware for Machine Learning Training, IEEE Transactions on Parallel and Distributed Systems, 2026, 1--12, IEEE, https://doi.org/10.1109/TPDS.2026.3702338, Tomaras.etal.2026a,
@Article{Tomaras.etal.2026a,
author={Tomaras, Dimitrios; Buschjäger, Sebastian; Kalogeraki, Vana; Morik, Katharina; Gunopulos, Dimitrios},
title={{STRATA2.0}: A Serverless Middleware for Machine Learning Training},
journal={IEEE Transactions on Parallel and Distributed Systems},
pages={1--12},
publisher={IEEE},
url={https://doi.org/10.1109/TPDS.2026.3702338},
year={2026},
abstract={Serverless computing has gained increasing interest in recent years for enabling large-scale machine learning tasks. However, training a machine learning model in a serverless setting is a complex task and several challenges need to be addressed particularly in data distribution, result aggregation, resource heterogeneity, failures, container ephemerality, network and execution cost. These...}}