<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>article From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1 in Technical Blog</title>
    <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/ba-p/90838</link>
    <description>&lt;P&gt;&lt;EM&gt;This is the first part of a two-part series blog on geospatial data processing on Databricks. The first part will cover ingesting and processing Overture Maps data on Databricks, while the second part will delve into a practical use case on dynamic segmentation.&lt;/EM&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Geospatial data is transforming how we understand and interact with our world, but processing this data efficiently at scale remains a significant challenge.&amp;nbsp;&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;A href="https://overturemaps.org/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Overture Maps&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; is a rich, open source geospatial dataset that promises to revolutionise mapping and location-based services. However, the sheer volume and complexity of this data can overwhelm traditional processing methods.&amp;nbsp; In this blog, we'll explore a practical approach to tackling this challenge with Databricks, focusing on filtering Overture map data for Victoria, Australia.&amp;nbsp; Along the way, we'll uncover techniques for parameterising notebooks, automating workflows, and benchmarking performance across different cluster sizes.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H1&gt;&lt;SPAN&gt;Overview of Overture Maps Data&lt;/SPAN&gt;&lt;/H1&gt;
&lt;P&gt;&lt;A href="https://overturemaps.org/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;The Overture Maps Foundation&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; has adopted the GeoParquet format for publishing its geospatial data. While it's possible to access specific data subsets using tools like &lt;/SPAN&gt;&lt;A href="https://docs.overturemaps.org/getting-data/sedona/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Apache Sedona&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; on Spark, in this blog we focus on downloading the entire dataset and applying our own filters and transformations. This approach offers several advantages:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Data Sovereignty and Control: By ingesting the entire dataset, you gain complete control over your data. This approach ensures that you're not dependent on external services or potential changes in data accessibility.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Performance Optimization: Having the data within your own environment allows for fine-tuned performance optimisations. You can structure the data in ways that best suit your specific use cases (e.g. liquid clustering), potentially leading to faster query times and more efficient processing.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Customisation Flexibility: With the full dataset at your disposal, you have the freedom to create custom data models, apply your own transformations, or combine Overture data with other proprietary datasets seamlessly.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;&amp;nbsp;&lt;/H1&gt;
&lt;H1&gt;&lt;SPAN&gt;Setting Up Databricks for Geospatial Data Ingestion&lt;/SPAN&gt;&lt;/H1&gt;
&lt;P&gt;&lt;SPAN&gt;Databricks is a data intelligence platform that provides robust analytics and machine learning for spatial and aspatial data. For geospatial data ingestion and processing, Databricks integrates seamlessly with various geospatial libraries and tools (e.g., Sedona, Databricks Mosaic, Geopandas, GDAL etc). &amp;nbsp; Recently Databricks has announced Spatial SQL, a native geospatial capability designed to enhance spatial data handling. As of this writing, Spatial SQL is in Private Preview for DBR 14.3+. Users interested in exploring this upcoming functionality should reach out to their Databricks account team to inquire about participating in the preview program.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H1&gt;&amp;nbsp;&lt;/H1&gt;
&lt;H1&gt;&lt;SPAN&gt;Downloading the data&lt;/SPAN&gt;&lt;/H1&gt;
&lt;P&gt;&lt;SPAN&gt;There are &lt;/SPAN&gt;&lt;A href="https://docs.overturemaps.org/getting-data/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;various ways&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; to download the Overture maps data.&amp;nbsp; We'll use azcopy to copy the data to a UC volume using a Databricks Notebook.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Firstly, install azcopy on the Databricks cluster if you haven’t done so:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;%sh
sudo bash -c 'cd /usr/local/bin; curl -L https://aka.ms/downloadazcopy-v10-linux | tar --strip-components=1 --exclude=*.txt -xzvf -; chmod +x azcopy'&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;This is my catalogue structure. I created a UC volume, danny_catalog.overture.raw, to store the geoparquet files. You can create yours with the UI or using SQL.&lt;/SPAN&gt;&lt;/P&gt;
&lt;DIV id="tinyMceEditordannywong_20" class="mceNonEditable lia-copypaste-placeholder"&gt;&amp;nbsp;&lt;/DIV&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_21-1726647973766.png" style="width: 278px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11308i7019860D6718087C/image-dimensions/278x270?v=v2" width="278" height="270" role="button" title="dannywong_21-1726647973766.png" alt="dannywong_21-1726647973766.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Create a folder in the raw UC volume, then copy the data using azcopy to that folder:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;%sh
mkdir -p /Volumes/danny_catalog/overture/raw/2024-08-20.0

azcopy copy "https://overturemapswestus2.dfs.core.windows.net/release/2024-08-20.0/" "/Volumes/danny_catalog/overture/raw/"  --recursive&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;Copying 431 GB of data (August 2024 release) from the Overture storage account to the UC volume took approximately 23 minutes.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_22-1726648091386.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11309i2D5CA675705104C4/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_22-1726648091386.png" alt="dannywong_22-1726648091386.png" /&gt;&lt;/span&gt;&lt;SPAN&gt;&lt;BR /&gt;However, it's important to note that transfer times can vary significantly based on several factors:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Geographic proximity: The time depends on the region of your Databricks workspace relative to the Overture data source.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Dataset size: Larger datasets will naturally take longer to transfer.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Network conditions: Transfer speeds can be affected by current network traffic and bandwidth.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Cluster resources: The specifications of your Databricks cluster may impact transfer speeds.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H1&gt;&lt;SPAN&gt;Create the filtering pipeline&lt;/SPAN&gt;&lt;/H1&gt;
&lt;P&gt;&lt;SPAN&gt;To automate the filtering process, we will leverage Databricks workflows and parameterised notebooks. This pipeline will run on a monthly schedule, ensuring that the filtered data remains up-to-date. The core of this pipeline involves filtering the Overture Maps data based on a multi-polygon boundary that closely approximates the Victoria region, with adjustments to include areas near but slightly outside the official boundary as those are also the areas of interest. This filtered data will then be persisted as a Delta table, making it readily available for downstream systems and processes to consume. The filtering logic utilises Spatial SQL's powerful ST_ functions to efficiently process the geospatial data. By automating this process, we ensure consistent and reliable data updates, which are crucial for our subsequent analyses and applications.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Environmental setup&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Let’s create the pipeline notebook.&amp;nbsp; Firstly we will create the notebook widgets:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;dbutils.widgets.text("catalog", "") # Catalog name
dbutils.widgets.text("schema", "") # Schema name
dbutils.widgets.text("table", "") # The filtered geo data by theme and type
dbutils.widgets.text("volume", "") # The volume that holds the raw files
dbutils.widgets.text("map_theme", "") # Overture map theme
dbutils.widgets.text("map_type", "") # Overture map type
dbutils.widgets.text("release", "") # Overture release 
dbutils.widgets.text("aus_polygon", "") # The geometry for the filter&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;You can provide default values to the notebook widgets such as:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;dbutils.widgets.text("catalog", "danny_catalog")&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;This is how it shows up on the notebook UI.&amp;nbsp; Populate the value that you desire for the widget and it will be referenced later:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_0-1726648458966.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11310i8E05A3F1DCC66CB1/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_0-1726648458966.png" alt="dannywong_0-1726648458966.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;My aus_polygon value is:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;MULTIPOLYGON(((146.224235224402 -35.4363358189844,148.225501405747 -35.647622056311,148.96662090591 -36.4718804131401,150.580844867015 -37.4059734691392,150.387767097845 -37.9947686179651,148.333245272028 -38.1583694792249,147.167982775491 -38.8351013798231,146.39794836319 -39.6562618160424,144.470264564547 -38.8532552527778,143.457314156555 -39.3875969932018,141.63332727031 -38.8880005741606,140.233033758295 -38.5024700559333,140.387650481347 -34.1564967083719,140.537933347279 -33.5832702977468,142.035376107123 -33.6588797210494,142.579114407923 -33.9833605220553,143.478747807493 -34.4440846825862,144.45098970567 -35.2900425539108,144.961566823313 -35.4071384438589,146.224235224402 -35.4363358189844)))&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;It is a multi-polygon that covers the State of Victoria in Australia with some buffers at the border. If you want to follow along, you can use this value or the geometry relevant to your use case.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Get the values from the notebook widget and store them as variables:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;mapTheme = getArgument("map_theme")
mapType = getArgument("map_type")
catalog = getArgument("catalog")
schema = getArgument("schema")
table = getArgument("table")
volume = getArgument("volume")
release = getArgument("release")
aus_polygon = getArgument("aus_polygon")&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;H2&gt;&lt;SPAN&gt;Processing&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Read the geoparquet file for the map theme and map type based on the input from the notebook widget.&amp;nbsp; Using the st_contains function from the Spatial SQL to filter data and only keeps those in the Victoria polygon.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="python"&gt;from pyspark.sql import functions as F

df = spark.read.parquet(f"/Volumes/{catalog}/{schema}/{volume}/{release}/theme={mapTheme}/type={mapType}/")

df_australia = df.filter(F.expr(f"st_contains(st_geomfromwkt('{aus_polygon}'), st_geomfromwkb(geometry))"))

df_australia.write.mode("overwrite").saveAsTable(f"{catalog}.{schema}.{table}")&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="dannywong_0-1726648709788.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11311i979B94594CEE7107/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_0-1726648709788.png" alt="dannywong_0-1726648709788.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Automate with Databricks workflow&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Click “Create Job” on the Workflow page.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_1-1726648742008.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11312i4EF11AB3F26EDF29/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_1-1726648742008.png" alt="dannywong_1-1726648742008.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Click “Edit parameters” under “Job parameters”.&amp;nbsp; Enter the value of the job level parameters.&amp;nbsp; These values will be passed to your notebook when your job runs.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_2-1726648776775.png" style="width: 444px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11313i82C46BA7738BD1BC/image-dimensions/444x356?v=v2" width="444" height="356" role="button" title="dannywong_2-1726648776775.png" alt="dannywong_2-1726648776775.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Add the task-level parameters.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;DIV id="tinyMceEditordannywong_16" class="mceNonEditable lia-copypaste-placeholder"&gt;&amp;nbsp;&lt;/DIV&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_3-1726648812004.png" style="width: 773px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11314iB382711602EEADBB/image-dimensions/773x540?v=v2" width="773" height="540" role="button" title="dannywong_3-1726648812004.png" alt="dannywong_3-1726648812004.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;This is the Overture map data structure, we are creating all the tasks based on this structure.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_4-1726648858345.png" style="width: 232px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11315i131ADF523D1B5652/image-dimensions/232x404?v=v2" width="232" height="404" role="button" title="dannywong_4-1726648858345.png" alt="dannywong_4-1726648858345.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Given the high degree of similarity among the tasks, switching to YAML code mode can streamline the task creation process. This allows you to copy, paste, and edit the tasks more efficiently, leveraging their similarities.&amp;nbsp; Alternatively you can use the &lt;/SPAN&gt;&lt;A href="https://learn.microsoft.com/en-us/azure/databricks/release-notes/product/2024/august#the-azure-databricks-jobs-for-each-task-is-ga" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;for each&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; task capability that was recently released.&lt;/SPAN&gt;&lt;/P&gt;
&lt;DIV id="tinyMceEditordannywong_17" class="mceNonEditable lia-copypaste-placeholder"&gt;&amp;nbsp;&lt;/DIV&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_5-1726648897819.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11316i2AEC985FD95596F0/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_5-1726648897819.png" alt="dannywong_5-1726648897819.png" /&gt;&lt;/span&gt;&lt;BR /&gt;&lt;SPAN&gt;Now, we have successfully created a workflow with 14 tasks that run in parallel on the same job clusters.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_6-1726648931896.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11317i7BE0A9E2EF33F354/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_6-1726648931896.png" alt="dannywong_6-1726648931896.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;TABLE style="width: 100%;"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;STRONG&gt;Cluster Size&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;STRONG&gt;Time&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P class="lia-align-center"&gt;&lt;STRONG&gt;Databricks Cost for this job&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;Single node&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;137 minutes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;$1.71&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;1 driver + 2 workers&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;69 minutes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;$2.59&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;1 driver + 4 workers&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;38 minutes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;$2.38&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;1 driver + 8 workers&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;21 minutes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;$2.36&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;1 driver + 16 workers&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;16 minutes&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD class="lia-align-center"&gt;
&lt;P&gt;&lt;SPAN&gt;$3.40&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;VM type for driver and workers: Standard_D4ds_v5&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;You can process the world’s geo data for less than the price of an ice cream scoop!&amp;nbsp; Databricks provides you with the flexibility to select the appropriate cluster size to align with the specific time constraints and budgetary considerations.&amp;nbsp; And you can now start working on your Overture data!&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Sample query to test it out:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%"&gt;&lt;LI-CODE lang="javascript"&gt;WITH H3_BUILDING AS (
SELECT
  id,
  level,
  height,
  names.primary AS primary_name,
  st_astext(ST_GeomFromWkb(geometry)) AS geometry, 
  inline(h3_tessellateaswkb(geometry, 10))
FROM danny_catalog.overture.overture_buildings_building
WHERE names.primary IS NOT NULL
)
SELECT * except(geometry) FROM H3_BUILDING&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-center" image-alt="dannywong_7-1726649044648.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/11318iFBE2333575EF52B9/image-size/large?v=v2&amp;amp;px=999" role="button" title="dannywong_7-1726649044648.png" alt="dannywong_7-1726649044648.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;It converts the geometry data of these buildings into a text format and applies a tessellation function to the geometry using the H3 geospatial indexing system. You can now do it with Spatial SQL on Databricks, without installing any additional libraries, with either a notebook or an SQL editor.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE style="border-style: hidden; width: 100%;" border="1" width="100%"&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD width="100%" height="355px"&gt;&lt;LI-CODE lang="javascript"&gt;CREATE TABLE danny_catalog.overture.overture_buildings_building_h3 AS (
  WITH H3_BUILDING AS (
    SELECT
      id,
      level,
      height,
      names.primary AS primary_name,
      st_astext(ST_GeomFromWkb(geometry)) AS geometry, 
      inline(h3_tessellateaswkb(geometry, 10))
    FROM danny_catalog.overture.overture_buildings_building
    WHERE names.primary IS NOT NULL
  )
  SELECT * EXCEPT(geometry) 
  FROM H3_BUILDING 
  CLUSTER BY (cellid)
)&lt;/LI-CODE&gt;&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;&lt;BR /&gt;This SQL code creates a new table that is clustered using Databricks’ &lt;/SPAN&gt;&lt;A href="https://docs.databricks.com/en/delta/clustering.html" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;liquid clustering&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; feature, specifically on these H3 cell IDs. By clustering data based on H3 indexes, the system ensures that spatially related information is stored together, greatly enhancing the speed and efficiency of spatial queries and analytics. This optimization is particularly valuable for big data applications, allowing for flexible, automatic, and efficient processing of massive geospatial datasets without the need for complex manual partitioning strategies.&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;H1&gt;&lt;SPAN&gt;Conclusion&lt;/SPAN&gt;&lt;/H1&gt;
&lt;P&gt;&lt;SPAN&gt;In this blog, we've navigated the exciting journey of transforming Overture Maps data using Databricks, showcasing how to efficiently process and refine geospatial data at scale. By leveraging Databricks' powerful capabilities, including parameterised notebooks and automated workflows, we've streamlined the process of filtering and analysing complex datasets.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Ready to elevate your geospatial analytics? Dive into Databricks' geospatial capabilities and experience the power of Spatial SQL.&amp;nbsp; Start your journey today and see how Databricks can uncover the stories behind your geospatial data!&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;In our &lt;A href="https://community.databricks.com/t5/technical-blog/dynamic-segmentation-in-geospatial-analytics-on-databricks-part/ba-p/91802" target="_self"&gt;next blog&lt;/A&gt;, we'll explore geospatial analytics in more depth, focusing on dynamic segmentation using Apache Sedona on Databricks. Don't miss the opportunity to see how these advanced techniques can further enhance your geospatial data analysis.&lt;/SPAN&gt;&lt;/P&gt;</description>
    <pubDate>Tue, 01 Oct 2024 15:49:40 GMT</pubDate>
    <dc:creator>dannywong</dc:creator>
    <dc:date>2024-10-01T15:49:40Z</dc:date>
    <item>
      <title>From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/ba-p/90838</link>
      <description>&lt;P&gt;&lt;SPAN&gt;This is the first part of a two-part series blog on geospatial data processing on Databricks. The first part will cover ingesting and processing Overture Maps data on Databricks, while the second part will delve into a practical use case on dynamic segmentation.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 01 Oct 2024 15:49:40 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/ba-p/90838</guid>
      <dc:creator>dannywong</dc:creator>
      <dc:date>2024-10-01T15:49:40Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/91488#M287</link>
      <description>&lt;P&gt;Nice article. Couple of questions: 1. Curious about the H3 near the end. Why did you choose H3 resolution 10 for the tessellation? That seems very coarse for buildings. You end up with maybe 20 buildings per H3 index. I guess it depends what other data you want to join to your buildings and what resolution that is. 2. Would it not make sense to keep the WKT representation of the buildings, so that higher resolution spatial joins can be done as needed - rather than totally abstract to the H3 cell id? Thanks.&lt;/P&gt;</description>
      <pubDate>Mon, 23 Sep 2024 20:28:58 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/91488#M287</guid>
      <dc:creator>GC-James</dc:creator>
      <dc:date>2024-09-23T20:28:58Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/91569#M288</link>
      <description>&lt;P&gt;Thank you for your questions!&amp;nbsp; You've raised some excellent points that deserve clarification:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;The choice of H3 resolution 10 in the example was primarily for demonstration purposes.&amp;nbsp; The resolution choice should be tailored to the specific use case. For broad area analysis like finding buildings within a suburb, a coarser resolution might suffice. For more precise operations, a finer resolution would be appropriate.&lt;/LI&gt;
&lt;LI&gt;H3 indexing provides an estimation that can be useful for quick spatial joins and proximity analysis. However, it's not intended to replace exact geometric operations.&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN&gt;A hybrid approach can be highly effective. Databricks supports combining H3 indexing with precise spatial functions like st_contains(). This allows for initial filtering using H3 followed by exact spatial operations where needed.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN&gt;Maintaining the original WKT/WKB representation alongside H3 indexes offers the best of both worlds. It allows for flexible querying strategies depending on the specific requirements of each analysis.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;Hope it clarifies!&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Tue, 24 Sep 2024 12:31:01 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/91569#M288</guid>
      <dc:creator>dannywong</dc:creator>
      <dc:date>2024-09-24T12:31:01Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/92299#M292</link>
      <description>&lt;P&gt;Nice one Danny&amp;nbsp;&lt;span class="lia-unicode-emoji" title=":smiling_face_with_sunglasses:"&gt;😎&lt;/span&gt;&lt;BR /&gt;Quick question: in the cost table, the transition from 8 workers to 16 workers is not linear. Is there a reason behind this?&lt;BR /&gt;&lt;BR /&gt;Will the same process work on serverless compute? The cost mentioned is probably for the DBUs only but there is still cost for cloud VMs. Hence it would be great to get an overall figure which could be simpler to calculate with serverless compute.&lt;/P&gt;</description>
      <pubDate>Mon, 30 Sep 2024 11:17:52 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/92299#M292</guid>
      <dc:creator>yousry-mohamed</dc:creator>
      <dc:date>2024-09-30T11:17:52Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/92302#M293</link>
      <description>&lt;P&gt;Thank you, Yousry. I apologize for the error in my previous calculations regarding single node, two worker nodes, and four worker nodes. I have made the necessary modifications, and the numbers should reflect a linear progression soon after the page is updated.&lt;/P&gt;
&lt;P&gt;&lt;BR /&gt;I intentionally excluded the cloud VM costs from the calculation due to their variability. These costs depend on several factors, including the cloud region of deployment, applicable discounts, and whether you're using a pay-as-you-go or commitment plan. Additionally, pricing differs across various cloud providers. To simplify the comparison, I focused solely on the Databricks costs.&lt;/P&gt;
&lt;P&gt;&lt;BR /&gt;For a rough estimate, you can assume that $1 in DBU (Databricks Units) will typically incur between $1 and $1.5 in underlying VM costs.&lt;/P&gt;</description>
      <pubDate>Mon, 30 Sep 2024 11:34:46 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/92302#M293</guid>
      <dc:creator>dannywong</dc:creator>
      <dc:date>2024-09-30T11:34:46Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93205#M304</link>
      <description>&lt;P&gt;Great Article. Nice to see more use-cases for geospatial on databricks and leveraging the new spatial sql preview.&amp;nbsp;&lt;/P&gt;&lt;P&gt;You mentioned running the pipeline on a monthly schedule to keep the filtered data up-to-date. Presumably running the azcopy component each time. Are you automating the overture release parameter, or just manually running the job and updating the parameter with the new release?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Wed, 09 Oct 2024 01:23:13 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93205#M304</guid>
      <dc:creator>mitchstares</dc:creator>
      <dc:date>2024-10-09T01:23:13Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93206#M305</link>
      <description>&lt;P&gt;It is manual currently, as the release dates are different every month:&lt;BR /&gt;&lt;A href="https://docs.overturemaps.org/release/latest/" target="_blank"&gt;https://docs.overturemaps.org/release/latest/&lt;/A&gt;&lt;BR /&gt;&lt;BR /&gt;We can potentially automate it by programmatically checking this page and getting the latest release version; if it is a new one, we can pass it as a parameter and trigger a run.&lt;/P&gt;
&lt;P&gt;Thanks!&lt;/P&gt;</description>
      <pubDate>Wed, 09 Oct 2024 01:41:11 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93206#M305</guid>
      <dc:creator>dannywong</dc:creator>
      <dc:date>2024-10-09T01:41:11Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93245#M307</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/60008"&gt;@mitchelh&lt;/a&gt;&amp;nbsp;-- I'm currently building an automated pipeline which I will run weekly to download Overture data. To know if there has been a new release there is this JSON file you can check:&amp;nbsp;&lt;A href="http://labs.overturemaps.org/data/releases.json" target="_blank"&gt;http://labs.overturemaps.org/data/releases.json&lt;/A&gt;&amp;nbsp;&lt;/P&gt;&lt;PRE&gt;{
    "latest": "2024-09-18.0",
    "releases": [
        "2024-09-18.0",
        "2024-08-20.0",
        "2024-07-22.0",
        "2024-06-13-beta.1",
        "2024-06-13-beta.0",
        "2024-05-16-beta.0",
        "2024-04-16-beta.0",
        "2024-03-12-alpha.0",
        "2024-02-15-alpha.0",
        "2024-01-17-alpha.0",
        "2023-12-14-alpha.0",
        "2023-11-14-alpha.0",
        "2023-10-19-alpha.0",
        "2023-07-26-alpha.0",
        "2023-04-02-alpha"
    ]
}&lt;/PRE&gt;&lt;P&gt;. I compare the 'latest' field with whatever version we have in our Lake, and then download the new one if they don't match.&lt;/P&gt;</description>
      <pubDate>Wed, 09 Oct 2024 10:08:52 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93245#M307</guid>
      <dc:creator>GC-James</dc:creator>
      <dc:date>2024-10-09T10:08:52Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93249#M308</link>
      <description>&lt;P&gt;That's awesome James!&lt;/P&gt;</description>
      <pubDate>Wed, 09 Oct 2024 10:56:17 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93249#M308</guid>
      <dc:creator>dannywong</dc:creator>
      <dc:date>2024-10-09T10:56:17Z</dc:date>
    </item>
    <item>
      <title>Re: From Raw to Refined: Processing Overture Maps Geospatial Data on Databricks - Part 1</title>
      <link>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93364#M309</link>
      <description>&lt;P&gt;Thanks&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/13171"&gt;@GC-James&lt;/a&gt;, that JSON is a great resource. I will definitely be building that into our pipeline!&lt;/P&gt;</description>
      <pubDate>Thu, 10 Oct 2024 02:25:44 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/from-raw-to-refined-processing-overture-maps-geospatial-data-on/bc-p/93364#M309</guid>
      <dc:creator>mitchstares</dc:creator>
      <dc:date>2024-10-10T02:25:44Z</dc:date>
    </item>
  </channel>
</rss>

