How can I find out what hosts are available for given requirements (LongJobs, memory, staging)
- condor_status -compact -constraint "HasChtcStaging==true" -constraint 'DetectedMemory>500000' -constraint "CanRunLongJobs isnt Undefined"
It looks to me like most hosts at CHTC are setup to run LongJobs. The following shows a small list of about 20 hosts. Is the correct?
- condor_status -compact -constraint "CanRunLongJobs is Undefined"
How can I know if my job swapped?
- Actually, it looks like the e2000 nodes don't have swap so this may not be an issue.
Is there a ganglia server or some other monitor service at CHTC we can view?
Are there bugs in the condor.log output of a DAG node? For example, I have a condor.log file that clearly shows the job taking about three hours to run yet at the bottom lists user time of 13 hours and system time of 1 hour. https://open-confluence.nrao.edu/download/attachments/40541486/step07.py.condor.log?api=v2
Condor Annex processing in AWS. Is there support for spot market
What network mask should we use to do allow ssh from CHTC into NRAO? Is there it a class B or several class Cs?

Answered Questions:

JOB ID question from Daniel
- When I submit a job, I get a job ID back. My plan is to hold onto that job ID permanently for tracking. We have had issues in the past with Torque/Maui because the job IDs got recycled later and our internal bookkeeping got mixed up. So my questions are:
  - Are job IDs guaranteed to be unique in HTCondor?
  - How unique are they—are they _globally_ unique or just unique within a particular namespace (such as our cluster or the submit node)?
- A Job ID (ClusterID.ProcID)
- DNS name of the schedd and ctime of the job_queued.log file.
- It is unique to a schedd.
- We should talk with Daniel about this. They should craft their own ID. It could be seeded with a JobID but should not depend on just it.
UpgradingHTCondor without killing jobs?
- schedd can be upgraded and restarted without loosing state assuming the restart is less than the timeout.
- currently restarting execute services will kill jobs. CHTC is working on improving this.
- negotiator and collector can be restarted without killing jobs.
- CHTC works hard to ensure 8.8.x is compatible with 8.8.y or 8.9.x is compatible with 8.9.y.
Leaving data on execution host between jobs (data reuse)
- Todd is working on this now.
Ask about installation of CASA locally and ancillary data (cfcache)
- CHTC has a Ceph filesystem that is available to many of their execution hosts (notibly the larger ones)
- There is another software filesystem where CASA could live that is more used for admin usage but might be available to us.
- We could download the tarball each time over HTTP. CHTC uses a proxy server so it would often be cached.
Environment: Is there a way to have condor "login" when a job starts thus sourcing /etc/proflie and the user's rc files? Currently, not even $HOME is set.
- A good analogy is Torque does a su - _username_ while HTCondor just does a su _username_
- WORKAROUND: setting getenv = True which is like the -V option to qsub, may help. It doesn't source rc files but does inherit your current environment. This may be a problem if your current environment is not what you want on the cluster node. Perhaps the cluster node is a different OS or architecture.
- ANSWER: condor doesn't execute things with a shell. You could set your executable as /bin/bash and then have the arguments be the executable you used to have. I just changed our stuff to staticly set $HOME and I think that is good enough.

...

Space shortcuts

Page tree

Versions Compared

Old Version 111

New Version 112

Key

Answered Questions:

Space shortcuts

Page tree

Page History

Versions Compared

Old Version 111

New Version 112

Key

Answered Questions: