Friday, 6 July 2007
dteam Woes
In order to de-ATLAS my test jobs (especially for the RB tests) I joined the dteam. This was a disaster to start with as my ATLAS jobs failed at some sites due to VOMS/Gridmap file issues and bugs in edg-job-submit. These seem to more or less resolved now except that all the jobs I submit through the Glasgow RB never run. They get submitted to a CE but stay in the queue for ever. No-one seems to know why.
RB Tests
There is a new set of tests targeted at UK Resource Brokers. These send (non ATLAS) "Hello World" jobs to each UK RB every 10 minutes to execute on any UK CE. See
http://hepwww.ph.qmul.ac.uk/~lloyd/gridpp/rbtest.html.
http://hepwww.ph.qmul.ac.uk/~lloyd/gridpp/rbtest.html.
Tuesday, 19 June 2007
RB Logging Problems
Although I switched to lcgrb02 the logging was still going to lcgrb01. This is because it's hardwired into /opt/edg/etc/edg_wl_ui_cmd_var.conf and can't be overwritten by your own conf file (how dumb is that?). Commented out LoggingDestination and it should be OK now. Stephen B reminds me that it is written here:
http://www.gridpp.ac.uk/deployment/users/faq.html#chooserb
http://www.gridpp.ac.uk/deployment/users/faq.html#chooserb
More RB problems
lcgrb01 at RAL is now giving problems similar to that seen with lcgrb02 a couple of weeks ago. I've switched to using lcgrb02 only for the time being.
Monday, 18 June 2007
ATLAS Release 13 and replica problems
There have been a few problems in the last few days. Firstly some sites installed ATLAS release 13.0.10 which broke some of the tests. At the moment the tests are only running on version 12.0.6 even if 13.0.10 is installed. Secondly something has happened to the UI at QMUL in that I cannot find out about my file replicas any more. This prevented analysis jobs being submitted over the weekend. This is being worked around at the moment by using a fixed list and not trying to obtain replica information dynamically.
Tuesday, 22 May 2007
Several problems
On Saturday morning (19 May) the machine that submits the test jobs crashed. On Sunday the system was moved to a different machine but now ~50% jobs fail to complete. This appears to be correlated with them going through lcgrb02.gridpp.rl.ac.uk rather than lcgrb01.gridpp.rl.ac.uk and looks like an RB problem rather than a problem at my end. I have now switched to the IC RB to see if this improves things.
Friday, 11 May 2007
More hangups
System hung up again trying to retrieve the job output. I have changed the scrip so that this is now done in a separate thread which is killed after n minutes so as not to hang the whole script.
Subscribe to:
Posts (Atom)