Generate Group Id for similar deduplicate values of a dataframe column.
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
11-22-2022 10:37 PM
Inupt DataFrame
'''
KeyName KeyCompare Source
PapasMrtemis PapasMrtemis S1
PapasMrtemis Pappas, Mrtemis S1
Pappas, Mrtemis PapasMrtemis S2
Pappas, Mrtemis Pappas, Mrtemis S2
Micheal Micheal S1
RCore Core S1
RCore Core,R S2
'''
Names are coming from the different source after doing a union those applied fuzzy match on it. now irrespective of sources need a group Id for similar values.
I want to use pyspark.
Output should be like below.
'''
KeyName KeyCompare Source KeyId
PapasMrtemis PapasMrtemis S1 1
PapasMrtemis Pappas, Mrtemis S1 1
Pappas, Mrtemis PapasMrtemis S2 1
Pappas, Mrtemis Pappas, Mrtemis S2 1
Micheal Micheal S1 2
RCore Core S1 3
RCore Core,R S2 3
'''