GroupBy in a multi node environment
Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
07-31-2021 02:48 PM
I have a group of rows with Information on a nested product calls.
example-Trxn1-product1-caller1-local1
Trxn1-Product1-local1-local2
Trxn1-Product1-local2-local3
here’s is a expected calls for a product
product1-caller1-local1
Product1-local1-local2
Product1-local2-local3
Product1-local3-local4
Product1-local4-local5
now I need to create one additional record to my dataframeTrxn1-product1-local3-local4. Picking only this coz, that’s the first missing link in the chain based on where it failed.
Like Trxn1, I have 10000s of them and I need to group these by transaction Id and apply the condition based on the product.
Trying if for loop is taking a
Labels:
- Labels:
-
Pyspark