Anonymous
Not applicable

@Jay Yang​ :

Yes, you can add a sleep time between the original run and the retrial run to avoid the library installation error due to the driver being recently restarted or terminated. Here's an example of how you can modify your code to add a sleep time:

import time
 
for forecast_id in weekly_id_list:
  LOGGER.info(f"Get forecasts from {forecast_id}")
  # Get forecast_date of the forecast_id
  query1 = (f"SELECT F.forecast_date FROM cosmosCatalog.{cosmos_database_name}.Forecasts F WHERE F.id = '{forecast_id}';")
  
  try:
    forecast_date = spark.sql(query1).collect()[0][0]
  except:
    # If the driver fails, sleep for 5 seconds and try again
    time.sleep(5)
    forecast_date = spark.sql(query1).collect()[0][0]

In this example, the try block attempts to collect the forecast date as usual. If it encounters an error due to the driver being recently restarted or terminated, it will sleep for 5 seconds before attempting to collect the forecast date again. You can adjust the sleep time as needed.