HPC & AI infrastructure

Your workload is the benchmark.

Find the constraint, measure it properly, and improve the system that researchers actually use.

Cluster troubleshooting

Failed or stalled jobs, node inconsistency, driver and runtime mismatches, mounting issues, and application environments. Diagnose one agreed failure, document the cause, and plan a safe repair.

Slurm & workload scheduling

Queues, fair-share, GPU allocation, job sizing, resource requests, and usage accounting. Align policy and placement with the needs of a shared research facility.

GPU & AI performance

CPU/GPU overlap, memory pressure, distributed execution, inference batching, concurrency, quantization, and KV-cache use. Compare latency, throughput, quality, and cost.

Storage & network fabric

IBM Storage Scale/GPFS, Lustre, NFS, metadata pressure, data staging, and InfiniBand/RoCE. Separate application, filesystem, network, and hardware limits before prescribing a change.

Reproducible environments

Linux, modules, Apptainer/Singularity, containers, MPI, CUDA, Python, and R environments. Make dependencies, versions, configuration, and run instructions visible.

Architecture & capacity

Design a new cluster or expand an existing one around workload, support, storage, fabric, power, and operating requirements. Procurement and hardware are quoted separately from engineering.

Performance with evidence

Busy GPUs are a signal.
Useful output is the goal.

We look at jobs completed, time to result, inference throughput, tail latency, and resource cost together. Utilization alone cannot tell you whether the system is serving the research well.

What a bounded assessment produces

  1. A baseline: agreed workloads, versions, settings, and measurements.
  2. A diagnosis: evidence linking the bottleneck to the observed behavior.
  3. A ranked plan: expected value, effort, dependencies, and operational risk.
  4. A repeatable check: how to tell whether the next change actually helped.

Production changes require an agreed maintenance window, authorization, and rollback plan. A diagnostic is not a promise that every issue can be fixed within its time cap.

Book an HPC assessment →