When one AI node is no longer enough
In earlier guides we built a 24/7 AI agent server with ClawBrain (IWILL N1522) and ran several assistants on one mini PC with Proxmox.
All of those scenarios share one dependency: the model itself runs somewhere else. The agent is yours and its memory is yours, but the reasoning happens behind someone else’s API key and is billed per token.
At some point that stops working – because data must not leave the company, because the monthly token bill has overtaken the cost of hardware, or because you want to fine-tune a model on your own documents. Then the conversation moves from a mini PC to a GPU server in your own server room.
This guide covers the practical side of that move:
- when an on-premises GPU server pays off against cloud GPU hours
- how to size your needs in VRAM rather than number of cards
- what really limits a chassis — slots, card length, power and airflow
- the difference between 4U, 5U, 7U and 9U GPU servers
- when a barebone chassis with a BMC beats an empty chassis
- when air cooling gives way to liquid cooling
Four reasons to bring AI workloads in-house
Data stays inside
Contracts, medical records, drawings, customer correspondence. When processing needs a clear lawful basis under UK GDPR, the cleanest answer is to run the model on the same network as the data – with no outbound traffic to an external provider.
Predictable cost under constant load
Tokens are billed per use. If the workload is constant – indexing, classification, an internal assistant for the whole team – the bill grows every month, while hardware is paid for once and then costs electricity.
Fine-tuning on your own data
Training a model on your documentation and terminology is the most direct way to make it useful for your business. It needs a lot of GPU time – exactly where owned hardware pays back fastest.
Independence from providers
Models are retired, prices change, terms are updated. A locally deployed model behaves the same way a year later – important once AI is built into a production process.
The honest downside: your own GPU server also means your own responsibility – power, cooling, spares and updates. For sporadic workloads, cloud GPU hours are almost always cheaper. The move makes sense when the load is constant or when data cannot leave the network.
First decision: how much video memory do you need?
The most common planning mistake is to start from the number of cards. The right starting point is total video memory (VRAM), because it decides which models can be loaded at all. Speed comes second.
Inference or training?
Training and fine-tuning need several times more memory than inference for the same model, because gradients and optimiser state are stored alongside the weights. A server designed only for inference is usually not enough once you decide to train.
How large are your models?
Parameter count and quantisation decide how much VRAM a model occupies. Quantisation cuts the requirement substantially for a small loss in quality – the same model may fit on one card or need several.
How many concurrent requests?
An internal assistant for five people and an inference endpoint for the whole company are different machines. Parallel requests take extra memory for each session’s context.
Rule of thumb: choose a chassis with at least twice today’s capacity. Swapping graphics cards takes an afternoon; swapping the chassis means a new server. That is why choosing between a 5-card and a 10-card chassis matters more than choosing a specific card.
What really limits the number of cards
“Supports 8 GPUs” is not one number but the result of four independent constraints. Miss one and the cards either do not fit, or fit and throttle.
01 · Slot width, not slot count
Modern cards occupy 2, 2.5, 3 or 3.5 slots because of their coolers. A chassis with 11 PCIe slots does not take 11 cards. Read specifications literally: the G4830-P4 is rated for 10 dual-width cards, the G9780-P4 for 10 cards 3.5 slots wide. Different chassis for different cards.
02 · Card length
A hard physical limit. The G4830-P4 accepts cards up to 270 mm; the G4835-P4 and G9780-P4 up to 350 mm. Those 80 mm decide whether a given card fits at all. Measure the actual card, not its catalogue length.
03 · Power and redundancy
Eight cards under load plus CPUs and drives need serious, stable power. GPU chassis use CRPS modules: the G4830-P4 and G9780-P4 run 4+1 redundancy, the F-G587-H12-A8 up to four modules in 3+1 with a 3200 W option, and the F-G791-H12-A8 up to eight modules in 4+4, feeding the motherboard and cards separately.
04 · Airflow through the cards
Tightly packed cards starve each other of air. The answer is high static pressure. The G4830-P4 combines three 120 mm positions at the front with three 80 mm and two 60 mm at the rear; the G9780-P4 has twelve 120 mm fan positions; the barebone models use eight high-performance N080 fans.
4U, 5U, 7U or 9U: what each extra unit buys you
Rack height is a trade-off between density and cooling. If you pay for colocation space, every U has a price. If the server lives in your own room, a taller chassis is usually the calmer choice.
| Format | Model | GPUs | Strength |
|---|---|---|---|
| 4U | G4835-P4 | 5 × 3.5-slot, up to 350 mm | Entry configuration with room to grow |
| 4U | G4830-P4 | 10 × dual-width, up to 270 mm | Maximum density in 4U |
| 4U | G465-12 | GPU + 12 × 3.5" hot-swap | When the data lives with the model |
| 5U | F-G587-H12-A8 | 8 GPUs, barebone | Ready platform with BMC and 32 DDR5 slots |
| 7U | F-G791-H12-A8 | 8 × full-length | Long cards and 4+4 power |
| 9U | G9780-P4 | 10 × 3.5-slot, up to 350 mm | Inference with maximum air volume |
Note the G465-12: it is the only row that solves a different problem. It combines GPU capacity with twelve 3.5″ drive bays and a 12 Gb/s Mini-SAS backplane – for cases where the training dataset or vector database must live in the same machine rather than be pulled across the network on every load.
Barebone platform or empty chassis?
Barebone platform · F-G587-H12-A8, F-G791-H12-A8
The chassis ships with motherboard, power supplies, cooling and BMC as a matched platform. You add CPUs, memory, drives and cards.
- Up to 32 DDR5 DIMM slots (4400/4800/5600 MHz)
- Built-in AST2600 BMC with IPMI 2.0 and Redfish
- 21 MCIO 5.0 interfaces for 10G/25G/40G cards
- OCP 3.0 slot and dedicated management port
- Manufacturer-validated compatibility
Empty chassis · G4835-P4, G4830-P4, G465 series
You choose the motherboard, power and cooling. More freedom and usually a lower entry cost, in exchange for more integration work.
- Motherboards up to 12" × 13" (E-ATX/ATX)
- Needs 5 × PCIe 4.0 x16 or 10 × x8 slots
- Free choice of CRPS or ATX power
- BMC depends on the chosen motherboard
- Component matching is your job
Which when: if the server will run in production with nobody on site every day, a barebone platform with a BMC pays for itself at the first remote reboot. If you have a team that builds and maintains hardware, an empty chassis gives you more machine for the same budget.
Cooling: where air stops being enough
A GPU server is above all a thermal problem. Cards run at full load for hours or days, not in few-second bursts. If airflow is insufficient, the cards throttle and the training you paid for simply takes longer.
Air cooling covers most scenarios – provided the chassis is designed for high static pressure and the room can actually remove the heat. Liquid cooling makes sense in three cases:
- High density – the maximum number of cards working constantly at the same time.
- Noise – the server is not in an isolated room, and high-static-pressure fans are loud.
- Room limits – the air conditioning cannot absorb the extra heat load.
From the liquid-cooled range, the Y-G587-H12-A8 handles up to eight cards in the RTX 4090/5090, A100, H100 and H200 class, with an external system separating cold and hot loops on 1″ or 0.75″ pipes. It takes up to twelve 2.5″/3.5″ drives, 32 DDR5 slots and CRPS power in 3+1 with 2000 W and 3700 W options.
Before you order: check three numbers for the room: the available power on the supply circuit, the air-conditioning capacity in kW and the rack depth. Chassis such as the G4830-P4 and G9780-P4 are over 630 mm long and will not fit a shallow rack.
A GPU server is also a network asset
The machine that holds your models and training data is among the most sensitive assets on the network. It has a BMC with full hardware control, often an exposed inference endpoint and access to internal data stores. The rules from ClawFirewall for self-hosted and hybrid AI apply twice over:
- The BMC port never faces the internet – a separate management VLAN, reachable only over VPN.
- Egress filtering – a server holding your data has no reason to start arbitrary outbound connections.
- Segmentation – publish the inference endpoint only to the networks that actually use it.
For organisations in scope of the UK NIS Regulations, or of the EU’s NIS2 Directive through their EU operations, this is part of risk management rather than a nice-to-have: local processing solves one problem but creates a new perimeter that must be documented and protected.
Three starting points
First server for a team · G4835-P4
4U, five 3.5-slot cards up to 350 mm. Start with one or two GPUs and grow without changing chassis.
Dense, constant inference · G4830-P4
4U, ten dual-width cards up to 270 mm and 4+1 CRPS power – maximum VRAM per rack unit.
Production platform · F-G587-H12-A8
5U barebone with BMC, 32 DDR5 slots and up to eight GPUs – manageable remotely from day one.
Sizing a GPU server for your models
IWILL GPU and AI server chassis are supplied to UK projects on enquiry and are not listed in the online catalogue. Tell us which models you want to run, whether you plan to fine-tune, how many concurrent users you expect and the power and rack depth available – we will propose a chassis and configuration.
