Hi Sir,

Regarding your question "what you are using in your RF i.e. hyperparameters, depths etc.", I would like to share the following points about my training process:

  • Using 30,000 features for the TF-IDF matrix

  • Setting n_estimators = 500

Since I am working on a serverless environment in Databricks, I am not quite sure what you meant by "CPU based cluster with minor tweak using parallelism". Could you please provide more details on this so I can better understand your point?

Thank you for your support.

Best regards,