Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
02-15-2022 03:24 PM
I'm trying to read a file from a Google Cloud Storage bucket. The filename starts with a period, so Spark assumes the file is hidden and won't let me read it.
My code is similar to this:
from pyspark.sql import SparkSession
spark = SparkSession.builder.getOrCreate()
df = spark.read.format("text").load("gs://<bucket>/.myfile", wholetext=True)
df.show()The resulting DataFrame is empty (as in, it has no rows).
When I run this on my laptop, I get the following error message:
22/02/15 16:40:58 WARN DataSource: All paths were ignored:
gs://<bucket>/.myfileI've noticed that this applies to files starting with an underscore as well.
How can I get around this?
Labels:
- Labels:
-
Spark job