Local GPU Infrastructure Docs Local GPU Infrastructure

Local GPU infrastructure overview#

Status: Current onboarding summary Last consolidated: 2026-09-23 Scope: The benchmark evidence and purchasing decision for a second Malaysian RTX 5090 ComfyUI service node Current decision: Retain the selected hardware design, quote the Corsair/Seasonic/Lian Li PSU gate, re-quote the procurement route, and accept the node only after topology, storage, thermal, and end-to-end validation Retailer-package ceiling: RM38,647 with the Seasonic PSU and separate 1 TB system drive

One-minute summary#

This directory records why the next ComfyUI worker is designed around an RTX 5090, two Samsung 9100 Pro SSDs in RAID 0, and a separate system drive. The workload changes a roughly 42.47 GB model set. Warm runs are mainly GPU-bound, but cold model changes are sensitive to the complete storage path through Linux, containers, the filesystem, and ComfyUI.

The benchmark evidence shows that every measured host sustaining approximately 10 GiB/s on the real model set used two NVMe drives in parallel. A single Samsung 9100 Pro reached only 5.93 GiB/s in the observed application path, despite much stronger advertised and synthetic speeds. The selected machine therefore uses two 2 TB Samsung 9100 Pros as a 4 TB disposable model cache, with both drives on independent CPU-connected PCIe 5.0 x4 links. Linux and all durable data live on a separate 1 TB Kioxia drive behind the chipset.

The MSI MAG X870E Tomahawk Max is retained because it can provide those two CPU-connected M.2 paths while keeping the RTX 5090 at PCIe 5.0 x16. AM5 cannot also provide a third CPU-direct x4 system-drive path without taking lanes from the GPU. Moving to Threadripper solely for the system drive costs at least RM5,800 more for the CPU and motherboard before new memory, cooling, and storage, with no measured workload need. It is rejected.

The component design is settled, but the purchase route is not. IdealTech's RM38,647 package is the ceiling. Obtain current written quotes from IdealTech and authorised Malaysian channels, then split the order only when the landed saving remains meaningful after delivery, tax, warranty, and assembly costs.

What is being built#

This is the second local ComfyUI service worker. The first worker already uses a Ryzen 9 9950X, RTX 5090, 64 GB of RAM, and two Samsung 9100 Pro 2 TB drives. Keeping the second worker close to that platform makes performance, power, thermal, and failure comparisons easier.

Component Current selection Why it stays
CPU AMD Ryzen 9 9950X The present workload does not require it, but the RM700 promotional upgrade adds preprocessing and concurrent-service headroom while matching the first worker
GPU Zotac RTX 5090 Solid OC 32 GB The non-negotiable throughput and VRAM component; Zotac is requested for fleet consistency
Motherboard MSI MAG X870E Tomahawk Max WiFi Supplies two independent CPU-connected Gen5 x4 M.2 paths without reducing the GPU below x16
Memory 64 GB DDR5-6000 Proven sufficient for the benchmark; 128 GB is deferred until production telemetry shows pressure
Model cache 2x Samsung 9100 Pro 2 TB Configured as Linux mdadm RAID 0 to target at least 10 GiB/s on real model reads
System drive Kioxia Exceria Basic 1 TB Keeps Linux, containers, logs, downloads, and outputs off the model-cache array
PSU 1200 W ATX 3.1 with its own supplied 600 W 12V-2x6 GPU cable Adds continuous-service margin over the 1000 W minimum without paying for unused 1500-2200 W capacity
Cooler Lian Li HydroShift LCD 360 ARGB Retained as part of the discounted package
Case Lian Li O11 Vision with eight Antec Vision fans Retained as part of the discounted package; airflow configuration matters
OS Linux Required for the operational stack and software RAID; the bundled unactivated Windows licence has no decision value

The target physical topology is:

Ryzen 9 9950X
├── PCIe 5.0 x16 ── Zotac RTX 5090
│
├── M2_1, CPU Gen5 x4 ── Samsung 9100 Pro 2 TB ─┐
│                                                ├─ mdadm RAID 0
├── M2_2, CPU Gen5 x4 ── Samsung 9100 Pro 2 TB ─┘  512 KiB chunks
│                                                   /srv/model-cache
│
├── M2_3, chipset Gen4 x4 ── Kioxia 1 TB
│                             Linux, containers, application, downloads and outputs
└── M2_4, chipset Gen4 x4 ── available for future scratch or output storage

M2_2 shares lanes with the rear USB4 controller. Configure M2_2 for its full Gen5 x4 link and accept that the rear 40 Gbps USB-C ports will be disabled. This does not reduce the GPU's x16 link.

Why storage is central to the decision#

The persisted input and model set is 42,471,968,783 bytes: 42.47 GB decimal or 39.56 GiB. The benchmark does not treat an advertised SSD number or a single fio result as proof. It measures the normal buffered read path used by the application after explicitly evicting files from cache.

The storage probe:

Every accepted cold sample must show physical model reads. Every warm sample must show zero physical model reads. This separates real filesystem-cold behavior from Linux page-cache effects.

Evidence that drives the design#

Observed storage layout Real model read What it establishes
2x Crucial T710 Gen5 RAID 0 15.59 GiB/s Fastest confirmed path, on a Threadripper platform
2x WD SN850X Gen4 RAID 0 on X870/9950X-class hosts 11.11-11.54 GiB/s Repeatable proof that two CPU-root-port drives on a comparable AM5 platform cross the target
2x Samsung 9100 Pro Gen5, effectively striped 11.54 GiB/s Direct support for the selected drive pair
1x Samsung 9100 Pro Gen5 5.93 GiB/s A premium single drive did not meet the application target and added a 21.31-second cold penalty
4x visible SN850X RAID 0 3.49 GiB/s Drive count and advertised topology do not prove that the application is mounted on the fast path

For the confirmed striped hosts, physical-device counters rose in an approximately 50/50 split during the probe. This is evidence that both members served the data rather than a RAM disk or unrelated device.

The conclusion is deliberately narrow: two correctly placed drives are the best-supported route to the 10 GiB/s target for this workload. It does not claim that every two-drive array will be fast or that a single 9100 Pro can never cross the target in a different local configuration.

Why the platform choices stay#

GPU and PCIe width#

The RTX 5090's 32 GB VRAM and measured ComfyUI throughput are the reason to build the node. An RTX 5080 is not an equivalent substitute. The GPU must negotiate PCIe 5.0 x16 under load.

PCIe 5.0 x8 does not halve the GPU's compute capability or VRAM, but it halves theoretical host-link bandwidth from roughly 63.0 GB/s to 31.5 GB/s in each direction. The impact varies by workload, but model upload, CPU offload, memory spill, and streamed transfers all use that link. Sacrificing x16 to add another M.2 path is rejected for a service explicitly optimizing model-loading behavior.

Motherboard and lane budget#

The 9950X exposes 24 usable device lanes: x16 for the GPU and two x4 paths. The Tomahawk uses those lanes for the RTX 5090 plus M2_1 and M2_2. A hypothetical x16 GPU plus three CPU-direct x4 SSDs needs 28 device lanes, which AM5 cannot supply.

Several AM5 boards expose three physical CPU-connected Gen5 M.2 sockets by taking lanes from the graphics slot. Populating those paths reduces the GPU to x8. More expensive X870E boards and another AM5 CPU cannot change the processor's lane budget, so they do not solve the problem.

Threadripper 9960X with a TRX50 AERO D can provide an x16 GPU and three CPU-direct SSD paths. It is technically clean but economically unjustified for the current workload: CPU and motherboard alone cost RM5,800 more than the 9950X/Tomahawk pair, and the migration also requires RDIMM memory, different cooling, and a replacement Gen5 system SSD. Current cold-load traces averaged fewer than two CPU cores, so this would buy lanes for an unmeasured system-drive bottleneck.

CPU and memory#

The included Ryzen 7 9850X3D is already sufficient for the measured workload. The 9950X is retained because the promotional upgrade is only RM700, doubles the core count, matches the first worker, and gives future preprocessing and concurrency headroom. It is an operational-margin choice, not a warm-inference requirement.

Keep 64 GB of RAM initially. A 64 GB host completed the workload, while the quoted 128 GB option cost about RM4,000 more. Monitor RSS, swap, major faults, and cache pressure. If production requires more memory, replace the two-DIMM kit with a validated 2x64 GB kit rather than assuming four DIMMs will remain stable at DDR5-6000.

Power supply, site power, and UPS#

Use a 1200 W ATX 3.1/PCIe 5.1 PSU with its own supplied 600 W 12V-2x6 GPU cable. The GPU is rated at 575 W and the manufacturer's system recommendation is 1000 W. A quality 1000 W unit is adequate, but 1200 W adds useful thermal and aging margin for continuous service. Larger units add no throughput or meaningful protection.

For a unit already owned, keep the Seasonic Vertex PX-1200. For the unbuilt machine, the regular Corsair HX1200i is the balanced value default if its written installed-system price saves at least RM250 against the Seasonic and confirms the exact ATX 3.1 revision, supplied 12V-2x6 cable, and local ten-year warranty. The Lian Li SX1200P remains the cost-first option if it saves at least RM500 and its exact black SKU and warranty are confirmed. Otherwise retain the Seasonic. The detailed evidence and decision tree are in the power-supply and site-power decision.

A premium PSU does not condition the building supply or bridge an outage. The immediate site priorities are one dedicated protected circuit per worker, verified earthing and RCD/RCBO operation, a correctly selected switchboard surge-protection device, and branch power-quality logging installed by a Malaysian licensed electrician.

Do not buy two large UPS units merely to “regulate voltage.” Record at least 30 days of voltage, sag, swell, and outage data first. If interruption cost or customer commitments justify UPS capacity, start around 2200 VA with at least 1800 W continuous pure-sine output per worker, protect the network path too, and test automatic job draining and shutdown.

Procurement position and price gates#

The RM30,999 IdealTech promotion carries roughly RM4,000 or more of package discount compared with the reconstructed standalone component prices. Most requested upgrades are then charged at ordinary retail-price differences. Entering through the promotion remains attractive even though the final design replaces several base components.

The current IdealTech ceiling is:

Line item Amount
Promotional base system RM30,999
Ryzen 9 9950X upgrade +RM700
X870E Tomahawk Max upgrade +RM450
First Samsung 9100 Pro 2 TB upgrade +RM1,950
Seasonic Vertex PX-1200 upgrade +RM900
Second Samsung 9100 Pro 2 TB +RM2,749
Quoted upgraded system RM37,748
Add the included Kioxia 1 TB back as the system drive +RM899
Three-drive ceiling RM38,647

The component decision remains valid, but do not assume every component must be bought from IdealTech. Request exact, current, written, landed quotes.

Decision gates:

  1. Ask IdealTech for the promotional base plus CPU and motherboard substitutions, the Kioxia retained, and side-by-side installed prices for the regular Corsair HX1200i, Seasonic PX-1200, and Lian Li SX1200P.
  2. Split the Samsung purchase only when the final saving survives freight, tax, and separate warranty handling. At RM2,299 per drive or less, split purchasing is compelling; from RM2,300 to RM2,799, judge the modest saving; at RM2,800 or above, use IdealTech.
  3. Ask Sun Cycle and authorised Zotac retailers for written, GPU-only Palit and Zotac 5090 quotes with exact SKU, stock, attachment requirements, DOA terms, and local warranty.
  4. Price a fully independent build only if the authorised RTX 5090 price approaches the promotion's implied allocation of roughly RM18.8k, or savings elsewhere clearly offset the lost package discount.
  5. Compare landed totals and warranty ownership, not web headlines. Historical prices are not purchase assumptions: the Samsung 9100 Pro 2 TB, for example, moved from RM1,630 in March to RM2,592 in August 2026.

A cost-reduced but technically sound version keeps the included 9850X3D and 1000 W PSU while retaining the Tomahawk and two Samsung drives. It totals RM36,947. It remains the fallback if minimizing capital matters more than CPU headroom and PSU margin.

Required build configuration#

Give the builder these non-negotiable instructions:

  1. Supply and invoice the exact RTX 5090 SKU, preferring the Zotac Solid OC used by the first worker.
  2. Retain or add back the Kioxia 1 TB as a third, dedicated system drive.
  3. Install the Samsung drives in M2_1 and M2_2, not in chipset-connected slots.
  4. Install the Kioxia in M2_3.
  5. Configure M2_2 for full Gen5 x4; losing rear USB4 is intentional.
  6. Confirm the RTX 5090 retains its intended x16 link.
  7. Record the exact PSU model and revision. Use only its supplied direct 12V-2x6 GPU cable, fully seated, and never reuse a modular cable from another PSU.
  8. Fit both M.2 heat spreaders and direct case airflow across the Gen5 drives.
  9. Configure bottom and side intake with rear exhaust; use the supported side-radiator orientation.
  10. Put no operating system, irreplaceable output, or other authoritative data on the RAID 0 cache.

Recommended array and filesystem starting points are Linux mdadm RAID 0, 512 KiB chunks, 2 MiB read-ahead, 1 MiB partition alignment, XFS or ext4 with noatime, and periodic fstrim rather than continuous discard on the hot path. Authoritative models must remain on HDD, NAS, or object storage so the cache can be rebuilt.

Production acceptance criteria#

Do not accept the node from its parts list or one synthetic number. Validate in this order:

  1. PCIe topology
    • Both Samsung drives negotiate 32.0 GT/s at x4.
    • The RTX 5090 negotiates x16 under load.
    • Both model drives enumerate on CPU paths; the Kioxia uses the chipset path.
  2. Array correctness
    • mdadm reports two healthy RAID 0 members, 512 KiB chunks, and approximately 4 TB usable capacity.
    • Physical read counters rise approximately equally on both members during a large read.
  3. Synthetic floor
    • Direct 1 MiB QD32 sequential reads comfortably exceed 20 GB/s.
    • Cache-evicted buffered single-stream reads exceed approximately 12 GB/s.
  4. Application storage target
    • The normal buffered model-file reader sustains at least 10 GiB/s across repeated samples after eviction.
  5. End-to-end target
    • Complete five cold/warm pairs.
    • Every cold sample passes physical-read verification.
    • Every warm sample records zero physical model reads.
    • Target a 5-8 second median cold-to-warm penalty for the current workload.
  6. Thermal stability
    • Read at least 200 GB continuously while recording both SSD temperatures and link states.
    • Correct any configuration that passes briefly and then throttles below the target.
  7. Electrical stability
    • Record idle, ComfyUI, and combined CPU/GPU stress input power at the branch meter.
    • Check for voltage sag, nuisance trips, and connector heating.
    • After initial thermal cycling and safe mains isolation, visually inspect the 12V-2x6 connection.

Apply the same acceptance test to the already-built first worker. If it misses the target, investigate M.2 placement, the M2_2/USB4 BIOS setting, actual filesystem/container placement, array configuration, and cooling before changing the second worker's selected components.

Assumptions and boundaries#

Immediate next actions#

Where to go deeper#