UII UPDATE 524 | AUGUST 2026
Multi-gigawatt "AI factory" facilities are often announced with great optimism, backed by developers' positive outlooks and strong investor confidence. Yet many of these projects have often struggled or encountered practical constraints, resulting in significant cost overruns. Similar issues have affected smaller projects, data center upgrades and retrofits.
New projects continue to be announced, with the US accounting for 80% of the large data center projects unveiled in 2025 (see Giant data center power plans reach extreme levels). Uptime Intelligence research has found that developers frequently struggle to deliver the capacity they have promised, resulting in cancellations, delays and downgrades (see Data center cancellations on the rise as public opposition grows and Many giant data center projects advance, despite risks).
This trend was recently reflected in the US project pipeline. Despite a narrative of accelerating growth, less US data center capacity was announced in the last quarter of 2025 than in the previous quarter: 25 GW of new project proposals in Q4 versus 49 GW in Q3 (see US capacity growth stumbled in 2025: what happened?).
Many factors can hinder a project, but rigorous methodology and careful planning can significantly improve the chances of success.
Like all data center projects, AI infrastructure needs to meet the demands for resiliency, efficiency, performance and scalability, while matching business needs. High-density AI projects face many common constraints, including:
The novel nature of AI data center infrastructure projects contributes to these constraints. AI facilities are often much larger than most data centers announced before 2020 and need to support GPUs operating at much higher rack power densities, alongside demanding latency and interconnectivity requirements. In addition to new cooling and power technologies, AI facilities require more total grid power, and their power consumption profiles have a greater impact on the grid. Increasing demand for equipment exacerbates supply chain issues, and concerns over the size and environmental demands of these facilities have provoked an unprecedented level of antipathy and public opposition.
Many of these factors are external to the project and may appear beyond the direct control of developers. Table 1 lists some practical measures that a methodical and flexible project plan can take to reduce the risks arising from external factors.
Table 1 AI facility development: issues and solutions

Because AI facilities are a new category of infrastructure, a methodical development plan will require fresh thinking and careful planning. Developers should seek guidance to address new technical and operational challenges. One source of advice is the AI Infrastructure Advisory service from Uptime Institute Professional Services (UIPS), which is part of Uptime Institute. This report will examine the content and outline the approach suggested by this advisory service.
The UIPS AI Infrastructure Advisory service consists of five pre-packaged consulting and support services that guide owners and developers through each stage of the process to meet the novel requirements of AI data center projects. These services progress through design, technology and vendor evaluation, construction, commissioning, and operations and management. An expanded description is available in five short guidance papers, which are summarized in this report and are available on the Uptime Institute website.
AI data center design requires a holistic approach that goes beyond traditional resiliency and availability. Power, cooling, space and workload requirements should be established as a design foundation before thermal and electrical strategies are validated against current and future demands.
Nvidia GPUs, which are ubiquitous in this sector, can have a thermal design power (TDP) of 700 W or more per chip, and are packaged in multi-GPU nodes. A single rack can have a total power demand of 100 kW or more (for the GPUs alone), compared with an average rack density in conventional facilities of less than 20 kW.
The UIPS AI Infrastructure Advisory service presents modular, scalable scenario designs for high-density data centers housing 130 kW racks combining air cooling and direct liquid cooling (DLC), and a modular power infrastructure. UIPS provides two scenarios, designed to deliver the performance specified in Uptime Institute's Tier Classification System at two levels: Tier III (concurrently maintainable) and Tier IV (fault-tolerant)
The concurrently maintainable scenario proposes construction in 15 MW phases, with each phase comprising three 5 MW data halls with independent power and cooling. The architecture uses a 5/4N distributed redundant power architecture and two independent bi-directional chilled water loops that supply both CDUs for liquid cooling and fan wall units for air cooling.
AI facilities are likely to combine air and direct liquid cooling (DLC); DLC availability is critical, with a short interval between cooling loss and shutdown. UIPS recommends that facilities using direct DLC adopt continuous cooling and provide the cooling equipment with UPS support.
The building design needs to accommodate AI racks, which can be up to 48U high and weigh 2,000 kg — substantially larger and heavier than conventional racks. The completed design should be evaluated for feasibility — including whether it can be constructed on schedule at the intended location — and that it meets all resiliency and efficiency requirements before finalization and approval.
With the design established, the AI Infrastructure Advisory service recommends the developer follow a carefully defined process to identify and procure infrastructure and equipment. Developers need a strategy for working with vendors to turn design concepts into actionable, contractual and technical documents. This guides the selection and management of suppliers and contractors responsible for delivering an integrated, functional facility that meets the owner's requirements.
Facility systems are specified, including generators, UPS, power distribution and switchgear, racking and cabling systems, and cooling systems. These are integrated into subsystems and, in turn, the overall design.
High-density computing may require certain designs and technologies that are relatively immature, such as DLC and higher-density power distribution. DLC does not replace air cooling, so a 130-kW rack will require more than 100 kW of DLC and more than 30 kW of air cooling. At minimum, 1-30% of the IT load will likely need air (see Guiding questions for liquid-cooled colocation planning). To deliver high power densities, it may be necessary to distribute power at higher voltages: roadmaps from GPU makers suggest 800 volts DC will become widely used, which may be expensive and harder to source.
Developers should produce requests for proposal (RFPs) or bid tender specifications that contain sufficient information to evaluate vendor bids. As AI infrastructure evolves, bid documents should ensure future-proofing by specifying requirements for easy upgrades and redeployment.
To mitigate supply chain risk, the bid documents should set required delivery timelines and outline the financial consequences of failing to meet these deadlines. The AI Infrastructure Advisory service recommends evaluating three vendors for each type of equipment (including UPS, engine generators, cooling units), a process which can be streamlined by pre-qualifying vendors using factors such as supported technology types and location.
During submission, vendors can ask questions to clarify project requirements and obtain product information not fully captured in the design. Bids will be evaluated using a template that aligns with the business standards used to develop contracts for delivery and integration. This process, broadly aligned with the Royal Institute of British Architects (RIBA) Plan of Work Stages 3 and 4, will deliver a detailed design to hand over to the construction phase.
Facilities for AI workloads often need to be built quickly — but not at the expense of quality or overall business objectives. They also need to meet a requirement for radical flexibility because AI is a developing field in which business models, IT loads and facility technologies are all subject to change.
Even with fast builds , drastic design changes may be necessary during construction. As AI infrastructure complexity increases, so does the risk that installed systems deviate from design intent — compromising performance, resilience or future flexibility. This is when design-to-build drift becomes a critical risk.
The AI Infrastructure Advisory service recommends:
AI training clusters use low-latency network connections between large numbers of GPU servers. The largest AI training facilities are moving to multiple stories, as systems on an upper floor will be closer to IT and cooling systems on the floor below than they would be in an adjacent building. The upper floors of these buildings must be strengthened appropriately, with prefabricated openings for connections between racks at different levels, with lifts and access corridors sized for AI racks.
While heavy liquid-cooled racks might warrant solid concrete slab floors, DLC introduces other considerations. Operators generally prefer to locate liquid circulation below the rack so any leaks can be contained and kept away from IT equipment. Strengthened raised floors may be required, with careful consideration of coolant-handling requirements and the possibility of future modifications.
In AI data centers, power and cooling equipment in the technical space is likely to be heavy and bulky. Cooling distribution units (CDUs) installed for DLC include large, heavy heat exchangers. Equipment in the gray space may therefore be larger and heavier than the racks it serves; the ratio of gray to white space will differ from that in conventional facilities, and gray space must be designed to support this equipment and enable its physical delivery and installation. If possible, the construction design should allow gray and white space allocation to be modified with relatively little effort and expense.
In conventional data centers, large items such as chillers or UPS systems are often installed early in the construction phase, with the building completed around them. In flexible AI designs, this may need to be reconsidered.
The AI Infrastructure Advisory service recommends phased construction, using a monitoring and testing regime developed from conventional builds to support the more exacting requirements of AI infrastructure.
Despite the urgent business need for the facility, the construction process is likely to take longer than a conventional build. This is partly because of anticipated changes, but also because of simple physical processes: larger concrete slabs require more time to pour and cure.
Periodic inspections are required:
Commissioning a data center is a systematic quality assurance process that ensures all technical systems function in accordance with owner and/or operator requirements prior to handover. It compares the delivered function with the design intent, assesses the facility's operational efficiency, resiliency and safety, verifies compliance with relevant regulations, and includes load testing to demonstrate that building systems can deliver sufficient power and cooling at the required levels of efficiency and resiliency.
High-density AI facilities have greater demands, with their power densities and use of emerging technologies such as DLC and medium-voltage power distribution. The commissioning process will involve closer co-operation with grid power operators, as the site may use temporary or permanent on-site power, and the dynamic power demands of AI training equipment have been shown to disrupt electric power grids. In other cases, the site may not interact with the grid at all.
Commissioning includes verification of compliance with relevant local regulations and applicable building standards. Verification should demonstrate efficiency levels, safe handling of materials such as cooling fluid, and satisfactory noise and emissions levels. This conformance will be shown by in-factory testing of equipment such as diesel gensets, as well as on-site testing, particularly where novel power sources are deployed.
The AI Infrastructure Advisory service recommends that facility commissioning use dummy loads ("load banks") in the white space, specifically designed to simulate AI workloads running on GPUs. The facility's systems will power and cool these loads, including DLC and air-cooling infrastructure.
The performance of the entire DLC system and its components needs to be tested under simulated regular operation as well as during planned and unplanned outages. Any novel equipment or modified designs will require new, tailored commissioning tests.
Fluids in the DLC testing process should be clean, as impurities can clog the narrow cooling circuits within the cold plates used to remove heat from the electronics. Accordingly, some operators may choose to retain in-house control of key parts of the load bank systems, such as cooling manifolds, or even entire load banks.
Where equipment comes from multiple suppliers or is owned by different organizations, a controlled staging area is essential so equipment can be tested in a known environment before installation. IT systems testing should be included in an integrated testing process that ultimately establishes that the facility is ready for operation. At this stage, documentation needs to be completed, checked and handed over to the operations staff.
The design, construction and commissioning of a facility have more-or-less defined start and endpoints, but its operation is open-ended. The operator's original business goals and requirements form the basis of a site operations program, which builds an awareness of day-to-day operational demands and should be integrated into every stage of the project from the outset.
Operational staff should receive all necessary documentation and be fully trained in operational procedures. The AI Infrastructure Advisory service recommends that these procedures are developed alongside the delivery process, so they are ready for handover and training well before project completion. Design and technology choices influence data center operations and should be made with a clear understanding of the eventual demands on operational staff.
AI data centers are likely to be far more automated than conventional data centers and will rely heavily on predictive control systems. Operating GPUs at high power with liquid cooling leaves little margin for error because a GPU can burn out in 30 seconds if coolant flow is interrupted.
This level of criticality means experienced staff will still be required to supervise and manage the facility. To ensure sufficient qualified staff are available to support business goals, the operator needs to establish management structures for hiring, training and (amid intense industry competition) retaining staff (see Survey highlights industry staffing crisis). An organizational framework will include job descriptions, clearly defined roles and responsibilities, the qualifications required for each role, and strategies for securing, retaining, developing and replacing staff.
In this new field, many staff members will need to be recruited and trained from scratch and many tasks will require external contractors that should be vetted and managed through careful procedures and detailed contractual agreements.
Operating procedures should ensure the safety and efficiency of new technologies. AI architectures may require rethinking traditional boundaries between facilities (OT) and IT functions, as OT-provided coolant circulates directly within IT racks. Operational procedures will specify maintenance schedules and upgrade procedures, which may occur at an accelerated rate compared with conventional facilities.