-werners-
Esteemed Contributor III

for pyspark you can use udf().

Here is an example on how to do this.