elikvar
New Contributor III

Hi @Kaniz Fatma​ @Vidula Khanna​ @Suteja Kanuri​, I was not able to resolve the issue. I was monitoring it for a little while and it was behaving fine but it failed again today with the same issue. The only logs I have are from the Event log of the cluster shown below.

One thing I did notice is that the previous run which triggers 7 hours before had the same ADD_NODES_FAILED event but was able to eventually add the nodes and run the job. This makes me think their might be a race condition someone in the startup of the cluster? I'm not to sure how that sequence works. I also attached the first run event logs.