Re: How to remove more than 4 byte characters usin... - Databricks Community

assuming you are having a string type column in pyspark dataframe, one possible way could be

identify total number of characters for each value in column (say
identify no of bytes taken by each character (say b)
use substring() function to select first n characters where n = floor(4 / b)