Predictive GPU Memory Modeling in Containerized Cloud Environments Using Self-Supervised Multi-Scale Learning

Main Article Content

Xiayuan Liu
Kai Gong

Abstract

This paper addresses the complexity and challenges of GPU memory usage prediction in containerized cloud environments and proposes a framework based on self-supervised learning and multi-scale time series modeling. The study first analyzes the dynamic characteristics of GPU memory utilization under multi-tenant and multi-task concurrency and points out that its usage patterns show significant non- stationarity and multi-scale dependencies. To address this, a multi-scale embedding mechanism is designed to unify sequence information at different time granularities, and a self-attention mechanism is applied to capture global dependencies, enhancing the representation of both short-term fluctuations and long-term trends. At the same time, a self-supervised pretraining module is introduced, which extracts latent temporal patterns from large-scale unlabeled data through a masking reconstruction task, reducing the reliance on labeled data and improving model generalization. In the prediction stage, the model uses fused feature representations to perform regression learning of GPU memory usage and adopts a joint optimization objective to improve accuracy while maintaining stability. Comparative and sensitivity experiments are conducted, covering mainstream methods such as LSTM, MLP, 1DCNN, and TimeMixer. The results show that the proposed method outperforms existing models in R², MAPE, MAE, and RMSE, and demonstrates robustness across multiple conditions, including window size, missing rate, and workload variations. These findings confirm the practical value and wide applicability of the framework in containerized cloud environments.

Article Details

Section

Articles