Data Engineering

Forum Posts

Sorted by:

by Vindhya • New Contributor II

04-18-2023 3:41:51 PM

2979 Views
1 replies
0 kudos

Dataframes to Pandas conversion step is failing with exception ""java.lang.IndexOutOfBoundsException: index: 16384, length: 4 (expected: range(0, 16384))"

Dataframes to Pandas conversion step is failing with exception ""java.lang.IndexOutOfBoundsException: index: 16384, length: 4 (expected: range(0, 16384))", PFB screenshot for more details

Data Engineering

2979 Views
1 replies
0 kudos

04-18-2023 3:41:51 PM

View Replies

Latest Reply

Anonymous
Not applicable

04-23-2023 9:14:00 PM

0 kudos

Hi @Vindhya D Thank you for posting your question in our community! We are happy to assist you.To help us provide you with the most accurate information, could you please take a moment to review the responses and select the one that best answers you...

0 kudos

04-23-2023 9:14:00 PM

by User16776430979 • Databricks Employee

06-07-2021 9:51:14 AM

4852 Views
0 replies
0 kudos

How to optimize and convert a Spark DataFrame to Arrow?

Example use case: When connecting a sample Plotly Dash application to a large dataset, in order to test the performance, I need the file format to be in either hdf5 or arrow. According to this doc: Optimize conversion between PySpark and pandas DataF...

Data Engineering

4852 Views
0 replies
0 kudos

06-07-2021 9:51:14 AM

by User16776430979 • Databricks Employee

06-04-2021 3:57:27 PM

2003 Views
0 replies
0 kudos

How to optimize conversion between PySpark and Arrow?

Seems like you can convert between dataframes and Arrow objects by using Pandas as an intermediary, but there are some limitations (e.g. it collects all records in the DataFrame to the driver and should be done on a small subset of the data, you hit ...

Data Engineering

2003 Views
0 replies
0 kudos

06-04-2021 3:57:27 PM

Databricks Community

Dataframes to Pandas conversion step is failing with exception ""java.lang.IndexOutOfBoundsException: index: 16384, length: 4 (expected: range(0, 16384))"

How to optimize and convert a Spark DataFrame to Arrow?

How to optimize conversion between PySpark and Arrow?