-werners-
Esteemed Contributor III

Spark will run on the whole dataset in background and return 1000 rows of that. So it might be that, not necessarily the function itself.

You can test that by f.e. starting with a dataset of 1000 records and apply the function on that.