- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
01-17-2023 12:28 AM
@Deepak U
It's hard to find an answer here without the proper monitoring of the entire process.
First of all, I would check how the workers behave during runtime (Ganglia - what's the overall usage of the workers, where's the bottleneck. It could be even more efficient to have a smaller number of more powerful nodes than multiple less powerful ones.
Second thing - networking. Maybe there's a bottleneck somewhere in there? Worth checking.
Third thing - Postgres instance - monitoring during the runtime.
Fourth thing - worth considering - exporting the data to Parquet files into storage, then using COPY FROM on Postgres to import the data. It's usually faster than using JDBC.