Note: The video version of this segment of the tutorial is here. (The video was recorded under the previous scheduler, LSF; where it shows bsub/bjobs commands, use the Slurm commands shown on this page.)

Running from the scratch directory

You should keep source code and scripts in your home directory, but you should run your jobs from the scratch directory. Code and scripts need to be in permanent storage. Output files can be large, and they should be generated in the scratch directory. Since the scratch directory is purged every 30 days, you must move the outputs that you want to keep to a local computer or to another storage space. Step 7 covers the supplemental storage available to you.

Exercise 4.0: Always run from the scratch directory

The directory name, /share, is an alias to another directory, and it is not the same for every user. Usually you would not need to know this, but printing the working directory from /share may be confusing. Here is what the full path may look like for a typical user:

[unityID@login ~]$ cd /share/groupname/unityID/
[unityID@login unityID]$ pwd
/gpfs_common/share01/groupname/unityID

  • Copy the current working version of the guide directory from your home directory to your scratch space on /share.
    Hint: Add the argument '-r' (for recursive) to the usual copy command to copy an entire directory.
  • What is the full path of your scratch directory?

    Finding software specific information

    Click to open a new tab to the documentation on using R. Every application installed on Hazel has a page like this one; they are reached from the Software page. Note the sections on loading R (module load R) and on running R in batch mode.

    Using the software documentation

    Use the R documentation to complete the following exercises.

    Exercise 4.1: Submit an R job to a compute node using Slurm

    All of the required instructions are in the documentation on using R.
  • Navigate the file system so that the current working directory is the guide directory that you have just copied to /share.
  • Create a batch submission script submit.sh that runs the R script weather.R with a 10 minute time limit, then submit the job to Slurm with sbatch submit.sh.
  • Hint: In submit.sh, type the same text as contained in the example batch script in the R documentation. Change the wall clock time to #SBATCH --time=00:10:00, and change the name of the script from my_program.R to weather.R. Optional: Change the names of the output and error files, and add #SBATCH --job-name=weather to set the job name. This example is not memory intensive, so you may delete the #SBATCH --mem=20G line.
  • Type squeue -u $USER (or the shortcut sq). When you first submit a job to Slurm, the job will be listed with a PD (pending) state. Hit the [up] arrow key to repeat the command to check the state. After the job begins, the state will change to R (running), and the job disappears from the list when it finishes.
  • The R code should produce a PDF file called weather.pdf containing temperature. After the job is finished, display the contents of your directory to confirm that weather.pdf has been created, and type file weather.pdf to check that it is a valid PDF file. If not, check the Slurm output files out.<jobid> and err.<jobid> generated from your run. Modify and resubmit the job if necessary.
  • The error file err.<jobid> should be empty (no errors), and the output file out.<jobid> should contain the list of temperatures printed by the program. To see resource usage for a completed job (e.g., CPU time and memory), use sjs <jobid> — see job monitoring.
  • Copy the output file weather.pdf to your local computer.

    Exercise 4.2: Always save code and scripts to a directory with backups

  • Copy the submit.sh script to your guide directory on /home.
  • Keep the browser tab with the R instructions open for Step 5.

    Go to Step 5