Generate Group Id for similar deduplicate values of a dataframe column.

Adig
New Contributor III

Inupt DataFrame

'''

KeyName KeyCompare Source

PapasMrtemis PapasMrtemis S1

PapasMrtemis Pappas, Mrtemis S1

Pappas, Mrtemis PapasMrtemis S2

Pappas, Mrtemis Pappas, Mrtemis S2

Micheal Micheal S1

RCore Core S1

RCore Core,R S2

'''

Names are coming from the different source after doing a union those applied fuzzy match on it. now irrespective of sources need a group Id for similar values.

I want to use pyspark.

Output should be like below.

'''

KeyName KeyCompare Source KeyId

PapasMrtemis PapasMrtemis S1 1

PapasMrtemis Pappas, Mrtemis S1 1

Pappas, Mrtemis PapasMrtemis S2 1

Pappas, Mrtemis Pappas, Mrtemis S2 1

Micheal Micheal S1 2

RCore Core S1 3

RCore Core,R S2 3

'''