ManojkMohan
Honored Contributor II

@TheOC @szymon_dybczak @BS_THE_ANALYST @Coffee77 the latest on this  

Batch Ingestion (10 TB Synthetic Data)

  1. Data Generation
    • Spark generates 100 billion synthetic rows (~10 TB).
    • Columns: id, random_val, category, payload.
    • Partitioned across 10,000 partitions for parallelism.
  2. Data Storage
    • Uses Databricks Unity Catalog: Catalog = 10tb, Schema = bronze.
    • Data is written as a managed Delta table: bronze.synthetic_10tb.
    • This is batch ingestion — one-time write of massive datasetManojkMohan_0-1756846467272.png

       

Streaming Ingestion (~500 MB/day)

  1. Data Generation
    • Spark structured streaming using rate source (~50 rows/sec → 500 MB/day).
    • Columns: id, random_val, category, payload.
  2. Checkpointing
    • Required for exactly-once guarantees.
    • Stored on S3 bucket: s3://streamingdataproto735/checkpoints/synthetic_500mb_continuous.
  3. Data Storage
    • Appends continuously to the same Bronze Delta table: 10tb.bronze.synthetic_10tb.
    • Supports indefinite streaming ingestion with micro-batches (availableNow or timed trigger). ManojkMohan_1-1756846485273.png

      currently resolving error of AWS creds config  . Request your thoughts in parallel ?  WIll summarise all learnings and publish a knowledge article on collective learnings