Brahmareddy
Esteemed Contributor II

Hi PeSe,

How are you doing today? As per my understanding, You're absolutely right to think through both options carefully. Option 1 runs into memory issues because it's trying to read the whole large file into memory at once, which doesn't work well for files over 100GB. Option 2 is technically correct, but it's painfully slow because the Databricks API only allows small chunk sizes (1MB), so uploading big files takes a lot of time. A better and much faster way would be to upload your large files to cloud storage first—like AWS S3, Azure Blob Storage, or GCP Cloud Storage—and then copy them into Databricks using dbutils.fs.cp or set up Auto Loader if you plan to do this regularly. Cloud platforms are designed for high-throughput data transfer, so this method will save you time and avoid memory or timeout issues. Let me know your cloud provider and I’d be happy to share exact steps.

Regards,

Brahma