<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Tuning with Optuna and MlflowSparkStudy in Machine Learning</title>
    <link>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167886#M4699</link>
    <description>&lt;P&gt;&lt;SPAN&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/96460"&gt;@AdamIH123&lt;/a&gt; Did you try this?&lt;/SPAN&gt;&lt;/P&gt;&lt;DIV class=""&gt;&lt;BR /&gt;&lt;DIV&gt;from Synapse. ml.lightgbm import LightGBMClassifier&lt;/DIV&gt;&lt;BR /&gt;&lt;H1&gt;Convert a Spark Pandas DataFrame to a standard PySpark DataFrame.&lt;/H1&gt;&lt;DIV&gt;spark_df = spa&lt;SPAN class=""&gt;rk_pandas_df.to_spark()&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;H1&gt;&lt;SPAN class=""&gt;Train distributed LightGBM&lt;/SPAN&gt;&lt;/H1&gt;&lt;DIV&gt;&lt;SPAN class=""&gt;lgbm = LightGBMClassifier(learningRate=0.1, numIterations=100)&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN class=""&gt;model = lgbm.fit(spark_df)&lt;/SPAN&gt;&lt;/DIV&gt;&lt;/DIV&gt;</description>
    <pubDate>Tue, 08 Sep 2026 08:34:39 GMT</pubDate>
    <dc:creator>kunduruanil</dc:creator>
    <dc:date>2026-09-08T08:34:39Z</dc:date>
    <item>
      <title>Tuning with Optuna and MlflowSparkStudy</title>
      <link>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167830#M4698</link>
      <description>&lt;P&gt;I am following the guide for tuning a model with Optuna and MlflowSparkStudy. My compute is configured with autoscaling enabled, with 1–2 Spark workers, each with 8 cores and 32 GB of memory. I set n_jobs=2 and trials=100, in &lt;STRONG&gt;mlflow_study.optimize()&lt;/STRONG&gt;.&lt;/P&gt;&lt;P&gt;I have a few questions about how MlflowSparkStudy distributes the workload:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;How should I think about n_jobs in mlflow_study.optimize() vs. num_threads/n_jobs in LightGBM?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;For example, with 1–2 Spark workers and 8 cores per worker, would it make sense to set mlflow_study.optimize(n_jobs=2) and LightGBM n_jobs=7? My dataset is large, so ideally I would like to run only &lt;STRONG&gt;one model-training job per Spark worker&lt;/STRONG&gt; and use the remaining cores on that worker for the LightGBM training.&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;Does MlflowSparkStudy automatically trigger or make use of Spark autoscaling?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;If I configure my cluster with a minimum of 1 worker and a maximum of 2 workers, will MlflowSparkStudy cause Spark to scale up to 2 workers as needed when running multiple trials in parallel?&lt;/P&gt;&lt;/LI&gt;&lt;LI&gt;&lt;P&gt;&lt;STRONG&gt;What is the recommended way to make a large DataFrame available to each Spark worker?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Currently, I have a Pandas DataFrame that I pass to the objective function used by mlflow_study.optimize(). Should I broadcast the DataFrame, cache it in Spark, or use another approach to avoid repeatedly transferring the data to each worker for every trial?&lt;/P&gt;&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Any guidance on the recommended configuration or best practices would be greatly appreciated.&lt;/P&gt;&lt;P&gt;#mlflow #optuna #&lt;SPAN&gt;MlflowSparkStudy&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/databricks/machine-learning/automl-hyperparam-tuning/optuna" target="_self"&gt;https://learn.microsoft.com/en-us/azure/databricks/machine-learning/automl-hyperparam-tuning/optuna&lt;/A&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 07 Sep 2026 19:24:05 GMT</pubDate>
      <guid>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167830#M4698</guid>
      <dc:creator>AdamIH123</dc:creator>
      <dc:date>2026-09-07T19:24:05Z</dc:date>
    </item>
    <item>
      <title>Re: Tuning with Optuna and MlflowSparkStudy</title>
      <link>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167886#M4699</link>
      <description>&lt;P&gt;&lt;SPAN&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/96460"&gt;@AdamIH123&lt;/a&gt; Did you try this?&lt;/SPAN&gt;&lt;/P&gt;&lt;DIV class=""&gt;&lt;BR /&gt;&lt;DIV&gt;from Synapse. ml.lightgbm import LightGBMClassifier&lt;/DIV&gt;&lt;BR /&gt;&lt;H1&gt;Convert a Spark Pandas DataFrame to a standard PySpark DataFrame.&lt;/H1&gt;&lt;DIV&gt;spark_df = spa&lt;SPAN class=""&gt;rk_pandas_df.to_spark()&lt;/SPAN&gt;&lt;/DIV&gt;&lt;BR /&gt;&lt;H1&gt;&lt;SPAN class=""&gt;Train distributed LightGBM&lt;/SPAN&gt;&lt;/H1&gt;&lt;DIV&gt;&lt;SPAN class=""&gt;lgbm = LightGBMClassifier(learningRate=0.1, numIterations=100)&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV&gt;&lt;SPAN class=""&gt;model = lgbm.fit(spark_df)&lt;/SPAN&gt;&lt;/DIV&gt;&lt;/DIV&gt;</description>
      <pubDate>Tue, 08 Sep 2026 08:34:39 GMT</pubDate>
      <guid>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167886#M4699</guid>
      <dc:creator>kunduruanil</dc:creator>
      <dc:date>2026-09-08T08:34:39Z</dc:date>
    </item>
    <item>
      <title>Re: Tuning with Optuna and MlflowSparkStudy</title>
      <link>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167904#M4700</link>
      <description>&lt;P&gt;Great questions—especially the distinction between Optuna’s trial-level parallelism and LightGBM’s intra-trial threading. The interaction with Spark autoscaling and data locality is also something I’d love to see documented with a concrete example. Curious what configuration others have found most efficient for large datasets.&lt;/P&gt;</description>
      <pubDate>Tue, 08 Sep 2026 10:49:26 GMT</pubDate>
      <guid>https://community.databricks.com/t5/machine-learning/tuning-with-optuna-and-mlflowsparkstudy/m-p/167904#M4700</guid>
      <dc:creator>ThiamLee</dc:creator>
      <dc:date>2026-09-08T10:49:26Z</dc:date>
    </item>
  </channel>
</rss>

