You can use our service desk portal for getting RIS support. RIS also offers 15 min. virtual office hours session Mon-Thru..
Compute2 Cluster
General Guidelines and Introduction:
Please note ‘Login Nodes’ are not compute nodes, they are intended to only for accessing cluster and doing house keeping tasks such as editing files, moving data etc. Usage on login nodes is restricted to 6GB per user per login node. Please read Compute2 guidelines and policies carefully that describes usage policy and best practices.
For doing computation you will have to submit batch jobs or interactive jobs. Or use our Open-OnDemand portal.
Connect via SSH:
You can connect via ssh to any of our three login nodes:
c2-login-001, c2-login-002 and c2-login-003
by running the below command:
ssh washukey@c2-login-00N.ris.wustl.edurequires a user's washukey and password, as well as the Duo 2FA set up. This means that a push will be sent to a user’s registered device upon login.
where N can be 1,2 or 3.
An example command would be:
ssh user@c2-login-001.ris.wustl.eduUsing Slurm
Load RIS module via ml
See what is available via
mlwithavailcommand
$ ml avail
--------------------------------------------------------- /cm/local/modulefiles ----------------------------------------------------------
boost/1.81.0 cmd gcc/13.1.0 (L) mariadb-libs openldap slurm/compute2/23.02.5 (L)
cluster-tools-dell/10.0 cmjob ipmitool/1.8.19 module-git python3
cluster-tools/10.0 dot lua/5.4.6 module-info python39
cm-bios-tools freeipmi/1.6.10 luajit null shared
--------------------------------------------------------- /cm/shared/modulefiles ---------------------------------------------------------
blacs/openmpi/gcc/64/1.1patch03 intel/compiler/latest intel/mkl/latest
blas/gcc/64/3.11.0 intel/compiler/2024.0.2 intel/mkl/2024.2 (D)
bonnie++/2.00a intel/compiler/2024.2.1 (D) intel/mpi/latest
cm-pmix3/3.1.7 intel/debugger/latest intel/mpi/2021.11
cm-pmix4/4.1.3 intel/debugger/2024.0.0 intel/mpi/2021.13 (D)
cuda11.7/blas/11.7.1 intel/debugger/2024.2.1 (D) intel/oclfpga/latest
cuda11.7/fft/11.7.1 intel/dnnl/latest intel/oclfpga/2024.0.0
cuda11.7/toolkit/11.7.1 intel/dnnl/3.5.0 (D) intel/oclfpga/2024.2.1 (D)
default-environment intel/dpct/latest intel/tbb/latest
fftw3/openmpi/gcc/64/3.3.10 intel/dpct/2024.2.0 (D) intel/tbb/2021.11
gcc11/11.5.0 intel/dpl/latest intel/tbb/2021.13 (D)
gdb/13.1 intel/dpl/2022.3 intel/vtune/latest
globalarrays/openmpi/gcc/64/5.8 intel/dpl/2022.6 (D) intel/vtune/2024.2 (D)
hdf5/1.14.0 intel/ifort/latest iozone/3.494
hdf5_18/1.8.21 intel/ifort/2024.0.2 lapack/gcc/64/3.11.0
hwloc/1.11.13 intel/ifort/2024.2.1 (D) mpich/ge/gcc/64/4.1.1
hwloc2/2.8.0 intel/inspector/latest mvapich2/gcc/64/2.3.7
intel/advisor/latest intel/inspector/2024.0 (D) netcdf/gcc/64/gcc/64/4.9.2
intel/advisor/2024.2 (D) intel/intel_ipp_intel64/latest netperf/2.7.0
intel/ccl/latest intel/intel_ipp_intel64/2021.12 (D) openblas/dynamic/0.3.18
intel/ccl/2021.13.1 (D) intel/intel_ippcp_intel64/latest openmpi/gcc/64/4.1.5
intel/compiler-rt/latest intel/intel_ippcp_intel64/2021.12 (D) openmpi4/gcc/4.1.5
intel/compiler-rt/2024.0.2 intel/itac/latest ris
intel/compiler-rt/2024.2.1 (D) intel/itac/2022.0 (D) ucx/1.10.1Load RIS module
ml load risJob Options
The help option provides a full list of what's available
The basic options are listed here
Account: required as there is no longer a default set for users.
A user can find their accounts by following the information in our Slurm Accounts documentation.
-A, --account=name charge job to specified account
CPU: Use
--ntasksfor increase of CPUs used with--cpus-per-task=1
-c, --cpus-per-task=ncpus number of cpus required per task -n, --ntasks=ntasks number of tasks to runJob name
-J, --job-name=jobname name of jobRam/Memory
--mem=MB minimum amount of real memory --mem-per-cpu=MB maximum amount of real memory per allocated cpu required by the job. --mem >= --mem-per-cpu if --mem is specified.Host/Node/Server
-N, --nodes=N number of nodes on which to run (N = min[-max])Partition/Queue
-p, --partition=partition partition requestedStandard Out (
stdout)
-o, --output=out location of stdout redirectionJob Run Time
-t, --time=minutes time limitUsing a container (Docker image)
--container-image=[USER@][REGISTRY#]IMAGE[:TAG]|PATH [pyxis] the image to use for the container filesystem. Can be either a docker image given as an enroot URI, or a path to a squashfs file on the remote host filesystem.Using a GPU
-G, --gpus=n count of GPUs required for the job --gpus-per-node=n number of GPUs required per allocated node --mem-per-gpu=n real memory required per allocated GPU --cpus-per-gpu=n number of CPUs required per allocated GPU
Using srun
Submits a job that runs in real time
Example using python
$ srun -A compute2-account -p general-interactive python helloworld.py
Hello World!You will need a python script called helloworld.py in order for the command to run. The example here has the following.
print("Hello World!")Storage allocations are mounted and you do not need to mount storage
Storage allocations do need to be mounted if using a container. See the Using Containers section.
You can get a shell connection to a job by using the
--ptyoption and/bin/bashas the command.
srun -A compute2-account -p general-interactive --pty /bin/bashThis works for jobs that use installed applications or containers.
Using sbatch
Submits a batch job (Runs in the background)
Create a job file:
testjob.slurm
#!/bin/bash
#SBATCH --job-name=python-test
#SBATCH --output=test.log
#SBATCH -p general-cpu
#SBATCH -A compute2-account
python helloworld.pySubmit job using
sbatch
$ sbatch testjob.slurmCheck the output
$ cat test.log
Hello World!Storage allocations are mounted and you do not need to mount storage.
Storage allocations do need to be mounted if using a container. See the Using Containers section.
Using squeue
Displays running jobs
Basic options listed here
By user
-u, --user=user_name(s) comma separated list of users to viewBy partition/queue
-p, --partition=partition(s) comma separated list of partitions to view, default is all partitionsBy name
-n, --name=job_name(s) comma separated list of job names to viewBy account
-A, --account=account(s) comma separated list of accounts to view, default is all accountsBy default shows all jobs
$ squeue
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
103 general python-t elyn R 0:11 1 c2-node-001
104 general python-t elyn R 0:06 1 c2-node-001
105 general osu_bibw sleong R 0:06 2 c2-node-[001,079]
106 general python-t elyn R 0:03 1 c2-node-001Use the
-uoption to show only your jobs
$ squeue -u elyn
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
103 general python-t elyn R 1:42 1 c2-node-001
104 general python-t elyn R 1:37 1 c2-node-001
106 general python-t elyn R 1:34 1 c2-node-001Using scancel
Cancels or kills running jobs
Default use job ID to cancel a job
scancel job_id[_array_id]Use the
--meoption to kill all your jobs
scancel --meCan cancel via other options
Job name
-n, --name=job_name act only on jobs with this nameUser
-u, --user=user_name act only on jobs of this userPartion/Queue
-p, --partition=partition act only on jobs in this partition
Using sinfo
Provides partitions/queues available
$ sinfo
PARTITION AVAIL TIMELIMIT NODES STATE NODELIST
bigmem up infinite 82 idle c2-bigmem-[001-002],c2-node-[001-080]
general* up infinite 80 idle c2-node-[001-080]
general-short up 30:00 80 idle c2-node-[001-080]
gpu up infinite 2 down* c2-gpu-[010,012]
gpu up infinite 94 idle c2-gpu-[001-009,011,013-016],c2-node-[001-080]Using GPUs
GPUs are available in the general and general-short partitions
Use a script that uses pytorch to test for GPU usage called
test_gpu.py
#!/usr/bin/env python3
import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")Load
slurmandpy-torchmodule. Shown below is an example command.
ml load slurm py-torchRun a test using a python script
srun -A compute2-account -p general-interactive --gpus=1 python test_gpu.pyUsing Containers
A container (Docker image) can be used with the following option
--container-image=[USER@][REGISTRY#]IMAGE[:TAG]|PATH
[pyxis] the image to use for the container
filesystem. Can be either a docker image given as
an enroot URI, or a path to a squashfs file on the
remote host filesystem.${Home}directories are automatically mountedThe
${Home}directory is not default working directoryUse the full path name for files
$ srun -A compute2-account -p general-interactive --container-image='python:3.9.21-alpine' python /home/elyn/helloworld.py
pyxis: importing docker image: python:3.9.21-alpine
pyxis: imported docker image: python:3.9.21-alpine
Hello World!The working directory can be set with the following option
--container-workdir=PATH
[pyxis] working directory inside the containerExample setting the working directory
$ srun -A compute2-account -p general-interactive --container-image='python:3.9.21-alpine' --container-workdir='/home/elyn' python helloworld.py
pyxis: importing docker image: python:3.9.21-alpine
pyxis: imported docker image: python:3.9.21-alpine
Hello World!Storage allocations are not mounted by default
Storage allocations can be mounted with the following option.
--container-mounts=SRC:DST[:FLAGS][,SRC:DST...]
[pyxis] bind mount[s] inside the container. Mount
flags are separated with "+", e.g. "ro+rprivate"Example mounting a storage allocaiton
$ srun -A compute2-account -p general-interactive --container-image='python:3.9.21-alpine' --container-mounts='/storage2/fs1/elyn/Active:/storage2/fs1/elyn/Active' ls /storage2/fs1/elyn/Active
pyxis: importing docker image: python:3.9.21-alpine
pyxis: imported docker image: python:3.9.21-alpine
testfile.txtThese same options can be used with
sbatchsubmissions
Using Job Arrays
Job arrays can only be used with
sbatchjobsThey can be used with the following option
-a, --array=indexes job array index valuesAn example of using a job arrays
Job script
#!/bin/bash #SBATCH --job-name=array-test #SBATCH --array=1-5 #SBATCH --output=testarray.log #SBATCH -p general-cpu #SBATCH -A compute2-account /bin/bash -c "date;sleep 300;date"Job submission
$ sbatch testarrayjob.slurm Submitted batch job 243Looking at job arrays in the queue
$ sbatch testarrayjob.slurm Submitted batch job 243 [elyn@c2-login-001 ~]$ squeue -u elyn JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON) 243_2 general array-te elyn R 0:09 1 c2-node-001 243_3 general array-te elyn R 0:09 1 c2-node-001 243_4 general array-te elyn R 0:09 1 c2-node-001 243_5 general array-te elyn R 0:09 1 c2-node-001 243_1 general array-te elyn R 0:10 1 c2-node-001Job arrays can all be cancelled
Example for job array (job ID) 50 with 5 elements
scancel 50Or particular array elements can be cancelled
scancel 50_1 scancel 50_[1-3] scancel 50_1 50_4