System Architecture
Resource allocation on Ferranti is managed using Slurm. Global storage is provided by a Weka file system. Inter-node communication uses a non-blocking InfiniBand fabric, with NDR400 (400 Gb/s) HCAs on the compute nodes. The cluster is housed in 10 air-cooled racks.
Login Nodes
| Hostname | CPU | RAM | Root disk |
|---|---|---|---|
ferranti-login.mlcloud.uni-tuebingen.de |
Intel(R) Xeon(R) Gold 6430 ×2 (64 cores / 128 threads) | 1.0Ti | 851G |
ferranti-login2.mlcloud.uni-tuebingen.de |
Intel(R) Xeon(R) Gold 6430 ×2 (64 cores / 128 threads) | 1.0Ti | 851G |
RAM type: DDR5-4800 ECC.
CPU Compute Nodes
| Nodes | CPU layout | RAM |
|---|---|---|
| mlcbm101, mlcbm102 | 2 sockets, 192 cores (384 threads) | 2.21 TiB |
- CPUs: 2× AMD EPYC 9654 (96 cores, 2.4 GHz, 384 MB L3)
- RAM type: DDR5-4800
- Local storage: 50 TB
GPU Compute Nodes
Compute nodes grouped by GPU model and hardware profile (live cluster state):
| GPU model | GPUs/node | Nodes | CPU layout | RAM/node | Node list |
|---|---|---|---|---|---|
h100 |
8 | 5 | 2 sockets, 96 cores (192 threads) | 1.97 TiB | mlcbm001, mlcbm002, mlcbm003, mlcbm004, mlcbm005 |
h100 |
8 | 10 | 2 sockets, 192 cores (384 threads) | 2.21 TiB | mlcbm006, mlcbm007, mlcbm008, mlcbm009, mlcbm010, mlcbm011, mlcbm012, mlcbm013, mlcbm014, mlcbm015 |
Per-node hardware reference
Per-card and per-node specs that do not change with cluster state:
Two CPU SKUs are deployed across the H100 fleet: Intel Xeon 8468 on mlcbm001–mlcbm005, AMD EPYC 9654 on mlcbm006–mlcbm015. Per-node CPU details are in the Compute-node deployment table.
Architectural per-card facts (same for any H100 SXM5 — see the full H100 spec sheet):
| Feature | H100 SXM5 |
|---|---|
| Accelerator connect | SXM5 |
| GPU memory / card | 80 GB HBM3 (3.35 TB/s) |
| NVIDIA Tensor Cores / card | 528 |
| FP32 cores / card | 16 896 |
| FP64 cores / card | 8 448 |
| Theoretical peak per node (8 GPUs) | FP16 Tensor 15 832 / TF32 Tensor 7 912 (sparse) / FP32 536 / FP64 272 TFLOPS |
Host RAM type: DDR5-4800.
CPU model, NVIDIA driver, CUDA paths, HCA, and /scratch_local per probed profile are in the Compute-node deployment table.
Partitions
| Partition | Nodes | GPUs (total) | Default time | Max time | Max nodes/job |
|---|---|---|---|---|---|
cpu-ferranti |
2 | — | — | 10-00:00:00 | 1 |
h100-ferranti |
15 | 120× h100 | — | 3-00:00:00 | ∞ |
h100-preemptable-ferranti |
13 | 104× h100 | — | 3-00:00:00 | ∞ |
Ferranti Interconnect
Ferranti has a fat tree InfiniBand interconnect topology with the following composition:
| Feature | Specifications |
|---|---|
| Number of core switches: | 4 |
| Number of edge switches: | 7 |
| Interconnect topology and type: | NDR InfiniBand Fat Tree, non-blocking |
| Blocking factor: | 1:1 |
Created: June 21, 2024