Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
11-15-2021 12:58 AM
I did with with boto3 instead with pyspark since its not a lot of files.
jsons_data = []
client = boto3.client('s3')
s3_resource = boto3.resource('s3')
bucket = s3_resource.Bucket(JARVIS_BUCKET)
for obj in bucket.objects.filter(Prefix=prefix):
file_name = obj.key
if re.search(ANCHOR_PATTERN, file_name):
json_obj = client.get_object(Bucket=JARVIS_BUCKET, Key=file_name)
body = json_obj['Body']
json_string = body.read().decode('utf-8')
jsons_data.append(json_normalize(json.loads(json_string,strict=False)))
df = pd.concat(jsons_data)