EndreM
New Contributor III

After increasing the compute to one with 500 GB memory, the job was able to transfer ca 300 GB of data, but it produced a large amount of files, 26000. While the old table with partition and no liquid cluster had 4000 files with a total of 1.2 TB of data. Why does liquid cluster or unity catalog result in so many files being produced?

It looks like the job ran for more than 2 days, but there is no record of the job run in the logs, and subsequently running the job times out after 3 hours. The data is only 1.2 TB and even with 500 GB of memory it looks like it is not enough... How to resolve this issue? 

Its a lot less flexibility and tooling available in debugging any PySpark issues in databricks than debugging a Kafka pipeline. When using Kafka you have the option to do low level coding in Apache Kafka, while also getting a lot out of the box functionality from higher level library Apache Steams. Peeking under the hood of what is going on isnt possible in Databricks. Quite frustrating.