The data center industry first took notice of the erratic power profile of AI training systems in 2023, at the time when the first truly large compute clusters came online to develop AI models even larger and better than OpenAI's ChatGPT-3. Now, the phenomenon is widely documented and heavily researched. In short, training generative pre-trained transformers (GPTs, the dominant crop of popular AI models today) on massively parallel compute clusters will create various types of repeated power fluctuations on a facility scale. AI inference does not create similar electrical patterns.
There are several contributing factors to these power fluctuations, but the dominant one is the training pipeline. Computationally intense stages of the training run across thousands of processing units are followed by memory operations moving large volumes of data. Every few seconds, this creates large-amplitude step changes in the load. Large clusters also rely on regular storage checkpoints to minimize any loss of work, during which power levels drop steeply before shooting up tens of seconds later.
Apply for a four-week evaluation of Uptime Intelligence; the leading source of research, insight and data-driven analysis focused on digital infrastructure.
Already have access? Log in here