UII UPDATE 515 | JULY 2026
The data center industry first took notice of the erratic power profile of AI training systems in 2023, at the time when the first truly large compute clusters came online to develop AI models even larger and better than OpenAI's ChatGPT-3. Now, the phenomenon is widely documented and heavily researched. In short, training generative pre-trained transformers (GPTs, the dominant crop of popular AI models today) on massively parallel compute clusters will create various types of repeated power fluctuations on a facility scale. AI inference does not create similar electrical patterns.
There are several contributing factors to these power fluctuations, but the dominant one is the training pipeline. Computationally intense stages of the training run across thousands of processing units are followed by memory operations moving large volumes of data. Every few seconds, this creates large-amplitude step changes in the load. Large clusters also rely on regular storage checkpoints to minimize any loss of work, during which power levels drop steeply before shooting up tens of seconds later.
When a training cluster uses several megawatts at full load, such swings in power can become an issue. To make matters worse, modern silicon exacerbates the issue by producing power excursion events above nominal design power ratings — creating stronger power swings and potentially overload conditions if capacity sizing of equipment is not appropriate, also considering inrush currents and power factor. Repeated overloads may cause breakers to trip, make UPS systems to use energy storage to match the load, or even force UPS systems to run in bypass mode (leaving the IT load unprotected against grid disturbances). Figure 1 shows the power profile issues of AI training (see Electrical considerations with large AI compute for a more detailed discussion).
Figure 1 Power profile of simulated GPU-based training cluster

Discussions with industry participants highlighted four major areas of concern:
The challenge of managing these AI training related power swings is still relatively new, and no single solution has yet emerged to address the issues described above for all scenarios. The problem is also becoming more prevalent, as AI training systems grow bigger (surpassing 10 MW) and become increasingly prevalent worldwide — establishing a new class of computer system. For comparison, as of July 2026, a mere handful of research supercomputers worldwide require more than 10 MW of power, and only a few dozen need more than 5 MW.
For operators with pre-existing large loads and spare capacity to host an AI training cluster, load diversity is their first best option to absorb most of the load swings associated with AI training. Load diversity simply means that the electrical effects of these sudden, synchronized changes in power demand are diluted by other workloads, making them less substantial for the data center power system. Some power delivery equipment, such as power distribution units and breakers directly serving the AI training system, will remain exposed to these load swings, but the additional costs (oversizing) are limited and the risks are more contained overall.
Addressing the root of the issue — the IT hardware — can be highly effective, both technically and economically, for new builds. This is especially true in existing data center capacity, where retrofitting electrical systems (such as changing UPS systems and batteries or adding flywheels) can introduce operational risks and may not be financially viable. Below is a summary of major IT options, including a combination of them:
In the foreseeable future, AI training clusters will become denser and larger in power demand, surpassing most research supercomputers. Next-generation AI training systems developing cutting-edge models could exceed 40 MW, with compute rack densities expected to exceed 400 kW before the end of the decade (although using wider, deeper cabinets). Current AI hardware technology roadmaps at Nvidia and AMD call for close coupling of increasing amounts of compute and memory resource to enable faster synchronization. This trend will exacerbate the intrinsic step load behavior, making for mitigation measures an important consideration in the infrastructure planning stage.
Cooperation between IT infrastructure teams (internal tenants or external customers) can avoid or reduce the need for more expensive mitigation layers in the facility power infrastructure. Manufacturers of UPS systems are honing their products to be able to almost fully dampen the sub-second and second-level oscillations of large AI training systems, even without load diversity. Addressing larger and longer-duration load steps to shield engine generators and the grid requires energy storage systems that are both sufficiently large and responsive. These systems can absorb high-energy events, such as checkpoints or starts and ends to training jobs. Additional options include battery energy storage systems and, alternatively, flywheel energy storage systems (where rotating mass stores kinetic energy and converts it into electrical energy when needed and vice versa).
The importance of large step loads extends beyond data center electrical systems. Although owners and operators already face scrutiny over power and water use, a relatively recent concern is how a high concentration of data centers is affecting — or will affect — the grid in some regions. This includes the need for network upgrades and the potential impact on grid stability. Several regulators and power grid operators worldwide are considering revising the rules of grid connection to address data center-specific load issues, which will include fast load changes (see Draft grid rules position data centers as active grid participants).
To date, much of the industry's attention has been duly directed toward solving the rack power density problem as training leading-edge generative AI models increasingly resembles supercomputing. However, large and rapid power swings place significant strain on electrical equipment and may affect grid stability. Addressing the issues at their root cause — the IT system software and hardware — can be both highly effective and less costly and complex than mounting layers of defence in the facility infrastructure.