We are excited to announce a major upgrade of the Compute2 Slurm Cluster, bringing significant improvements in security, performance, reliability, and platform capabilities.
💡 What’s new:
- Upgrading to Slurm 25.05.5, featuring the latest scheduler enhancements and fixes.
- Cluster-wide OS and NVIDIA platform driver upgrades for improved security, performance, and compatibility.
- Enhanced High Availability (HA) through network and controller infrastructure improvements.
- Upgrading to Spack v1 provides a more reliable and controllable experience, such as precise module versioning, and makes it easier to install a wider variety of software.
❗️Impact:
Maintenance window: 9:00 AM – 5:00 PM CDT
- All running jobs will be terminated at 9:00 AM CDT on Wednesday, July 15th.
- Compute2 Open OnDemand will be inaccessible during this time.
Please save your work and plan job submissions accordingly. All other services will be unimpacted.
📢 Actions required:
After the upgrade, users will be required to:
- With the upgrades, users are required to:
- Specify a Slurm account when submitting a job ‘-A’ parameter.
- Failure to do so will return, “error: slurm_job_submit: You must specify an account when submitting a job with the ‘-A’ parameter”
- Specify versions when loading/unloading modules.
- Failure to do so will return, “Lmod has detected the following error: These module(s) or extension(s) exist but cannot be loaded as requested: “MODULENAME” Try: “module spider MODULENAME” to see how to load the module(s).”
- Path to software must be exposed for module spider traversal.
- Failure to do so will result in the expected module to not be listed.
- Users need to update/remove the ssh known hosts from their local machine .
- Failure to do so will result in the inability to connect to the compute2 cluster.
- Rebuild user-installed Spack modules for v1 compatibility.
- Resubmit jobs that were terminated.
Have questions? Contact us at the RIS Service Desk.