Data Engineering

by KrishZ • Contributor

09-11-2022 7:49:10 AM

18091 Views
3 replies
3 kudos

[Pyspark.Pandas] PicklingError: Could not serialize object (this error is happening only for large datasets)

Context: I am using pyspark.pandas in a Databricks jupyter notebook and doing some text manipulation within the dataframe..pyspark.pandas is the Pandas API on Spark and can be used exactly the same as usual PandasError: PicklingError: Could not seria...

Data Engineering

18091 Views
3 replies
3 kudos

09-11-2022 7:49:10 AM

View Replies

Latest Reply

ryojikn
New Contributor III

01-14-2023 9:06:21 PM

3 kudos

@Krishna Zanwar , i'm receiving the same error.For me, the behavior is when trying to broadcast a random forest (sklearn 1.2.0) recently loaded from mlflow, and using Pandas UDF to predict a model.However, the same code works perfectly on Spark 2....

3 kudos

01-14-2023 9:06:21 PM

2 More Replies

by parthibsg • New Contributor II

08-29-2022 7:56:29 PM

1614 Views
1 replies
2 kudos

When to use Dataframes API over Spark SQL

Hello Experts,I am new to Databricks. Building data pipelines, I have both batch and streaming data.Should I use Dataframes API to read csv files then convert to parquet format then do the transformation? orwrite to table using CSV then use Spark SQL...

Data Engineering

1614 Views
1 replies
2 kudos

08-29-2022 7:56:29 PM

View Replies

Latest Reply

Debayan
Databricks Employee

08-30-2022 1:09:02 PM

2 kudos

Hi Rathinam, It would be better to understand the pipeline more in this situation. Writing to table using CSV and then using spark SQL will be faster in few cases than the other one.

2 kudos

08-30-2022 1:09:02 PM

Databricks Community

Forum Posts

[Pyspark.Pandas] PicklingError: Could not serialize object (this error is happening only for large datasets)

When to use Dataframes API over Spark SQL