Cluster troubleshooting
Failed or stalled jobs, node inconsistency, driver and runtime mismatches, mounting issues, and application environments. Diagnose one agreed failure, document the cause, and plan a safe repair.
HPC & AI infrastructure
Find the constraint, measure it properly, and improve the system that researchers actually use.
Failed or stalled jobs, node inconsistency, driver and runtime mismatches, mounting issues, and application environments. Diagnose one agreed failure, document the cause, and plan a safe repair.
Queues, fair-share, GPU allocation, job sizing, resource requests, and usage accounting. Align policy and placement with the needs of a shared research facility.
CPU/GPU overlap, memory pressure, distributed execution, inference batching, concurrency, quantization, and KV-cache use. Compare latency, throughput, quality, and cost.
IBM Storage Scale/GPFS, Lustre, NFS, metadata pressure, data staging, and InfiniBand/RoCE. Separate application, filesystem, network, and hardware limits before prescribing a change.
Linux, modules, Apptainer/Singularity, containers, MPI, CUDA, Python, and R environments. Make dependencies, versions, configuration, and run instructions visible.
Design a new cluster or expand an existing one around workload, support, storage, fabric, power, and operating requirements. Procurement and hardware are quoted separately from engineering.
Performance with evidence
We look at jobs completed, time to result, inference throughput, tail latency, and resource cost together. Utilization alone cannot tell you whether the system is serving the research well.
Production changes require an agreed maintenance window, authorization, and rollback plan. A diagnostic is not a promise that every issue can be fixed within its time cap.
Book an HPC assessment →