Energy and Cooling Define the Future of AI Data Centers: An AI Rack Is More Than a Collection of GPUs

Energy and Cooling Define the Future of AI Data Centers: An AI Rack Is More Than a Collection of GPUs
Dražen Tomić / Tomich Productions

Converting a conventional data center into AI-ready infrastructure requires several times more electrical power, liquid cooling, and substantially greater structural load capacity. The decisive factor is the availability of power, which increasingly determines whether a project can be realised at all. A1 Croatia has already upgraded its data center with GPU infrastructure based on NVIDIA Blackwell technology. At the same time, it is increasing the share of energy from renewable sources and expanding capacity to meet growing demand for AI services.

Artificial intelligence is reshaping the fundamental architecture of data centers and placing demands on operators that conventional infrastructure often cannot meet. In Croatia and the wider European market, the single biggest constraint is the availability of electricity from the local grid, according to Bojan Klasan, Data Center Manager at A1 Croatia, speaking to ICTbusiness Media - ICTbusiness.info. Converting an existing facility first and foremost requires a several-fold increase in electrical power per rack. Instead of the current 5 to 15 kW, new AI racks may require 40 to 120 kW or more. Higher power also brings liquid cooling, larger pipes, new cable routes, and substantially greater floor loads. If the building lacks the required height, load-bearing capacity, or the ability to connect to the power grid, construction of a new facility becomes unavoidable. Retrofitting generally ceases to be cost-effective once its cost exceeds 60 to 70 percent of the price of a new data center. The transformation can nevertheless be carried out in phases by building isolated AI zones with their own power supply and cooling. This approach allows GPU infrastructure to be introduced without interrupting existing IT systems. Long-term sustainability will be determined not only by processors, but equally by energy, cooling, modularity and the efficiency of the entire system.

How can an existing data center be converted into an AI data center?

At its core, converting a conventional data center into an AI data center requires a dramatic leap from 5 to 15 kW to 40 to 120 kW or more per rack. In existing facilities, it is usually possible to upgrade the networking equipment and internal power distribution, and install piping for liquid cooling, or completely replace the cooling system if necessary. However, construction of a new facility becomes unavoidable if the building lacks sufficient ceiling clearance for large pipes and cable routes, or if the floors cannot bear the load of new AI racks weighing more than 2.5 tonnes.

In Croatia and the wider European market, the single biggest constraint is the electricity available from the local grid. If the local substation cannot deliver the required power, the project stops before it begins. Immediately after energy comes the physical limitation of the cooling system, because conventional air-cooling systems can no longer dissipate such high heat loads. Let us assume that sufficient grid capacity is available, but it is unclear whether a brownfield retrofit is worthwhile. Its viability is assessed by comparing the cost of adaptation with the cost of building a new facility; anything above 60 to 70 percent of the cost of a new build pushes the decision towards a new facility. The target PUE is equally important, because an inefficient retrofitted system will quickly consume the initial investment savings through high electricity bills, while creating a constant risk to operational stability.

The transformation can nevertheless be carried out in phases and without interrupting existing systems. This is achieved by introducing hybrid zones or modular units within the data center. Smaller isolated "pods" are created, each with its own liquid-cooling loop and dedicated power supply, allowing new AI capacity to come online step by step while conventional IT systems continue to operate without disruption. A considerable number of hyperscalers use this approach to retrofit their ageing data centers.

What does an optimally designed NVIDIA AI rack look like?

An optimally designed NVIDIA AI rack is not simply a collection of expensive graphics processors, but a complete, integrated supercomputer in miniature. To avoid a scenario in which costly accelerators sit underutilised while waiting for data, the entire architecture must be precisely aligned, orchestrated, and free of bottlenecks. The choice of architecture, such as NVIDIA HGX or NVL platforms, depends on whether the system will be used for large-scale training of complex models or for their subsequent real-time deployment.

Aligning the hardware requires complete synergy among all resources. InfiniBand or a dedicated ultra-low-latency Ethernet network is essential for training AI models. Within the node itself, NVLink enables direct communication between GPUs, while NVMe storage with GPUDirect Storage sends data straight to GPU memory, bypassing the CPU. At the same time, DPUs take over network traffic and security, freeing the graphics accelerators for pure AI computation. In this equation, the cooling system reaches a critical threshold above 25 to 30 kW per rack, when air cooling can no longer physically keep pace with the load. Rear-door heat exchangers provide a good transitional solution for relatively modest loads of up to 60 kW, but direct-to-chip liquid cooling has become the new industry standard for AI, removing most of the heat and allowing GPUs to operate at maximum performance. Immersion cooling offers exceptional efficiency, but its complex maintenance requirements and specialised fluids mean that it remains a solution for specific niches.

The ultimate economics depend heavily on software orchestration and dynamic workload scheduling. A1 Croatia applies this principle of an integrated and well-orchestrated system in practice. We have upgraded our data center by introducing next-generation GPU infrastructure based on NVIDIA RTX 6000 Pro Blackwell GPUs, in partnership with Exoscale. The technology provides high-performance GPUs for training large neural networks, deep learning, AI research models, faster data processing, and real-time analytics. Exoscale and A1 also ensure full data sovereignty and transparency regarding data location, as well as infrastructure compliant with the GDPR and ISO standards, supported by a team based in Croatia and the EU.

Can energy infrastructure keep pace with the growth of AI?

The availability of electricity and a grid connection has become the primary criterion for selecting sites for new AI data centers, pushing the previously decisive proximity to financial centers or fibre-optic hubs into second place. Global energy consumption in the sector could double by the end of the decade, while demand for AI itself is growing by more than 25 percent annually. Amid this rapid expansion, renewable energy sources are no longer merely an environmental choice; they are becoming a central pillar of the industry's sustainability, competitiveness and long-term survival.

Without the large-scale integration of green energy, it will be difficult to justify such enormous computing capacity. Long-term stability and cost control are therefore increasingly based on corporate power purchase agreements (PPAs) for the direct purchase of energy from wind farms, nuclear power plants and solar power plants, significantly reducing exposure to market shocks and fluctuations in fossil-fuel prices.

Because solar and wind generation are inherently variable, the key challenge is to align their output with the uninterrupted operation of a data center. Industrial battery energy storage systems (BESS), combined with on-site microgrids, take on this role by stabilising the power supply, shaving peak loads and reducing operating energy costs by 15 to 20 percent. To provide reliable baseload power when there is no wind or sun, modular gas plants designed for conversion to green hydrogen are being developed in parallel, along with local waste-to-energy systems.

The exponential growth of AI workloads also requires clear support from regulators through accelerated investment in the transmission grid, stronger incentives for renewable energy and the use of AI solutions for smart energy management, so that the digital and green transitions can work in tandem.

In this context, the A1 data center already provides a practical example of how theory can be put into practice. Our infrastructure uses advanced consumption metering, optimised cooling systems and precise monitoring of energy indicators. In addition, 93 percent of our total energy already comes from verified renewable sources, while some of our needs are met directly by our own solar power plants. With a clear strategic goal of reaching 100 percent renewable energy by 2030, we are ensuring that our infrastructure not only meets the strictest environmental standards, but also provides users with a highly efficient and sustainable long-term computing environment.

Where is the business value of an AI data center created today?

The business value of an AI data center is no longer measured in square metres or megawatts, but by how efficiently expensive hardware is converted into billable computing hours. The most successful models are GPU-as-a-Service providers and advanced colocation operators with liquid-cooling-ready infrastructure. Demand for sovereign AI clouds is also rising sharply in Europe because of strict data-protection regulations.

The financial model is extremely aggressive because processors lose their competitive edge after just three years. Traditional multi-year depreciation schedules no longer apply. The investment must pay for itself within 36 months, requiring operators to maintain average capacity utilisation above 70 to 80 percent. The risk of overbuilding is therefore real when capacity is built speculatively, without clear demand behind it. Sustainable long-term demand will not come from a short-lived wave of model training, but from large-scale industrial deployment - inference in everyday business processes.

How do you design a data center that will not become obsolete within five years?

Building a facility with a 25-year lifespan for hardware that becomes obsolete in three years requires a complete departure from rigid construction and energy designs. In this race, modular data centers are the future. Pre-engineered, factory-assembled modules allow capacity to be expanded gradually and at precisely the pace at which demand grows, avoiding the risks of overbuilding and obsolescence.

I believe the market will divide into two key layers over the next five years: massive mega-campuses for training large models close to sources of inexpensive energy, and agile, distributed edge AI centers with capacities ranging from 1 to 10 MW. At this distributed edge, the modular concept demonstrates its full potential by combining AI inference with 5G standalone networks and massive IoT deployments. Autonomous systems and industrial robotics require data processing with latency below five milliseconds, which modular edge centers can provide close to the data source while saving network bandwidth and transmission costs.

The modularity of the internal architecture is equally critical to longevity. This means ceiling heights of more than five metres, floor load capacities of at least two tonnes per square metre and oversized cooling mains. Five years from now, the industry will be shaped not only by the pace of chip development, but also by power-grid constraints, strict European regulation and agility at the network edge.