- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
09-09-2025 03:07 AM
We're iceberg's java lib to write managed iceberg tables in databricks. We actually can create these tables using databricks as iceberg REST catalog. But this only works when we provide a partitioning spec. This is then picked up as cluster_columns for databricks. Unfortunately the data files we put into the partition paths (e.g. 'tables/xxx-xxx/data/_kafka_date_day=2025-05-21/xxx.parquet') remain unmaintained.
databricks duplicates the data into it's own clustering scheme.
We were told that partitioning for managed iceberg tables is unsupported. But no one could tell us how we can create a table via iceberg REST catalog properly so that it can be filtered on `__kafka_date` with correct file pruning.
Could someone provide a sample CURL to databricks for table creation that achieves this?