Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
09-01-2022 01:26 PM
As of my knowledge, there are not any options to optimize your code. https://github.com/databricks/spark-xml
It is the correct and the only way for reading XMLs, so on the databricks side, there is not much you can do except experiment with other cluster configurations.
Reading multiple small files is always slow. Therefore, it is common to know an issue called the "tiny files problem."
I don't know your architecture, but maybe when XMLs are saved, files can be appended to the previous one (or some trigger could merge them).
My blog: https://databrickster.medium.com/