Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
09-02-2025 01:55 PM
@TheOC @szymon_dybczak @BS_THE_ANALYST @Coffee77 the latest on this
Batch Ingestion (10 TB Synthetic Data)
- Data Generation
- Spark generates 100 billion synthetic rows (~10 TB).
- Columns: id, random_val, category, payload.
- Partitioned across 10,000 partitions for parallelism.
- Data Storage
- Uses Databricks Unity Catalog: Catalog = 10tb, Schema = bronze.
- Data is written as a managed Delta table: bronze.synthetic_10tb.
- This is batch ingestion — one-time write of massive dataset
Streaming Ingestion (~500 MB/day)
- Data Generation
- Spark structured streaming using rate source (~50 rows/sec → 500 MB/day).
- Columns: id, random_val, category, payload.
- Checkpointing
- Required for exactly-once guarantees.
- Stored on S3 bucket: s3://streamingdataproto735/checkpoints/synthetic_500mb_continuous.
- Data Storage
- Appends continuously to the same Bronze Delta table: 10tb.bronze.synthetic_10tb.
- Supports indefinite streaming ingestion with micro-batches (availableNow or timed trigger).
currently resolving error of AWS creds config . Request your thoughts in parallel ? WIll summarise all learnings and publish a knowledge article on collective learnings