<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: How to bring the project id info into the databricks billing usage table on GCP in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170118#M56235</link>
    <description>&lt;P&gt;The reason you are not seeing the GCP Project ID directly in the system.billing.usage table is that Databricks system tables track DBU usage at the workspace, cluster, and workload level rather than underlying Google Cloud project infrastructure.&lt;/P&gt;&lt;P&gt;As such, Databricks workspaces in GCP are typically deployed either within a specific customer-managed VPC / GCP Project or across multiple projects and teams successfully implement showback/chargeback using a few proven strategies.&lt;/P&gt;&lt;P&gt;Solution 1: Workspace to Project Mapping Table (Easiest)&lt;/P&gt;&lt;P&gt;If your architecture utilizes specific Databricks Workspaces to specific GCP Projects (i.e., Project A hosts Workspace A, Project B hosts Workspace B), create a reference mapping table in Unity Catalog.&lt;/P&gt;&lt;P&gt;Create a Mapping Table and Join with System Billing Usage&lt;/P&gt;</description>
    <pubDate>Tue, 29 Sep 2026 07:21:07 GMT</pubDate>
    <dc:creator>Satyasai</dc:creator>
    <dc:date>2026-09-29T07:21:07Z</dc:date>
    <item>
      <title>How to bring the project id info into the databricks billing usage table on GCP</title>
      <link>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170114#M56233</link>
      <description>&lt;P&gt;Hi,&lt;BR /&gt;&lt;BR /&gt;We are using Databricks on GCP and are currently analyzing costs using the billing usage/system tables.&lt;/P&gt;&lt;P&gt;One challenge we are facing is that the billing usage data provides usage and cost information, but we are unable to identify the associated GCP Project ID for each usage record.&lt;/P&gt;&lt;P&gt;Our goal is to:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Map Databricks usage and costs to specific GCP projects.&lt;/LI&gt;&lt;LI&gt;Generate project-level chargeback/showback reporting.&lt;/LI&gt;&lt;LI&gt;Understand which GCP project is consuming the most Databricks resources.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Has anyone implemented a solution for this?&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;&lt;DIV&gt;&lt;P&gt;Any guidance, examples, or best practices would be greatly appreciated.&lt;/P&gt;&lt;/DIV&gt;</description>
      <pubDate>Tue, 29 Sep 2026 05:59:37 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170114#M56233</guid>
      <dc:creator>Danish11052000</dc:creator>
      <dc:date>2026-09-29T05:59:37Z</dc:date>
    </item>
    <item>
      <title>Re: How to bring the project id info into the databricks billing usage table on GCP</title>
      <link>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170118#M56235</link>
      <description>&lt;P&gt;The reason you are not seeing the GCP Project ID directly in the system.billing.usage table is that Databricks system tables track DBU usage at the workspace, cluster, and workload level rather than underlying Google Cloud project infrastructure.&lt;/P&gt;&lt;P&gt;As such, Databricks workspaces in GCP are typically deployed either within a specific customer-managed VPC / GCP Project or across multiple projects and teams successfully implement showback/chargeback using a few proven strategies.&lt;/P&gt;&lt;P&gt;Solution 1: Workspace to Project Mapping Table (Easiest)&lt;/P&gt;&lt;P&gt;If your architecture utilizes specific Databricks Workspaces to specific GCP Projects (i.e., Project A hosts Workspace A, Project B hosts Workspace B), create a reference mapping table in Unity Catalog.&lt;/P&gt;&lt;P&gt;Create a Mapping Table and Join with System Billing Usage&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 07:21:07 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170118#M56235</guid>
      <dc:creator>Satyasai</dc:creator>
      <dc:date>2026-09-29T07:21:07Z</dc:date>
    </item>
    <item>
      <title>Re: How to bring the project id info into the databricks billing usage table on GCP</title>
      <link>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170120#M56236</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/188867"&gt;@Danish11052000&lt;/a&gt;&amp;nbsp;System tables does not generally carry details such as GCP Project ID. You can create an &lt;STRONG&gt;static mapping table&lt;/STRONG&gt; (GCP Project id, other details, workspace id) in Unity Catalog and use the info with the system tables if your deployment is like every Databricks workspace maps to exactly one GCP project. You can load this mapping using &lt;STRONG&gt;Terraform deployment scripts&lt;/STRONG&gt;, since each workspace's resources land in a specific project. You can join the lookup table to system.billing.usage on workspace_id, filtering for cloud = 'GCP' and record_type = 'ORIGINAL'. You can also join system.billing.list_prices to convert the DBUs to dollars, giving the correct project level chargeback report to see which GCP project is consuming the most resources.&lt;/P&gt;&lt;DIV&gt;&lt;BR /&gt;&lt;DIV&gt;You can rely on sub-workspace attribution via &lt;STRONG&gt;custom tags&lt;/STRONG&gt; if you have multiple GCP projects sharing a single workspace. You can apply a custom tag like gcp_project_id = &amp;lt;project&amp;gt; to the clusters, SQL warehouses and pools so Databricks propagates it into the custom_tags column of system.billing.usage. To prevent untagged usage, enforce this through cluster policies so computes cannot even be created without it. For the &lt;STRONG&gt;serverless&lt;/STRONG&gt;, you can leverage &lt;STRONG&gt;serverless usage policies&lt;/STRONG&gt; to auto-tag the usage and it lands in the custom_tags column.&lt;/DIV&gt;&lt;BR /&gt;&lt;DIV&gt;You can extract custom_tags['gcp_project_id'] in the queries for the per-project cost breakdowns and that tags only apply to usage incurred after the tag is set - old records will remain untagged and hence get these policies in place quickly.&lt;/DIV&gt;&lt;/DIV&gt;</description>
      <pubDate>Tue, 29 Sep 2026 08:09:54 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170120#M56236</guid>
      <dc:creator>balajij8</dc:creator>
      <dc:date>2026-09-29T08:09:54Z</dc:date>
    </item>
    <item>
      <title>Re: How to bring the project id info into the databricks billing usage table on GCP</title>
      <link>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170122#M56237</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/188867"&gt;@Danish11052000&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;A feasible workaround is to use the Databricks Account Workspaces API to fetch the GCP project_id alongside workspace_id. The workspace API schema exposes it under cloud_resource_container.gcp.project_id. &lt;A href="https://docs.databricks.com/api/workspaces/v1/workspace" target="_self"&gt;Reference&lt;/A&gt;&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;url = f"https://accounts.gcp.databricks.com/api/2.0/accounts/{ACCOUNT_ID}/workspaces"
workspaces = requests.get(
    url,
    headers={"Authorization": f"Bearer {TOKEN}"}
).json()

rows = [
    (
        str(w["workspace_id"]),
        w.get("workspace_name"),
        w.get("cloud_resource_container", {})
         .get("gcp", {})
         .get("project_id")
    )
    for w in workspaces
]
df = spark.createDataFrame(
    rows,
    ["workspace_id", "workspace_name", "project_id"]
)&lt;/LI-CODE&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="data_pulse_0-1790669526481.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31573iF6A387901127E0AA/image-size/medium?v=v2&amp;amp;px=400" role="button" title="data_pulse_0-1790669526481.png" alt="data_pulse_0-1790669526481.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;Then persist in a mapping table like&amp;nbsp;workspace_project_map and enrich billing as:&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;SELECT
  u.*,
  m.project_id
FROM system.billing.usage u
LEFT JOIN finops.workspace_project_map m
  ON u.workspace_id = m.workspace_id&lt;/LI-CODE&gt;&lt;P&gt;At present,&lt;STRONG&gt; system.access.workspaces_latest&lt;/STRONG&gt; provides workspace metadata such as workspace_id, workspace_name, workspace_url, and status but doesn't expose the GCP &lt;STRONG&gt;project_id &lt;/STRONG&gt;as mentioned in &lt;A href="https://docs.databricks.com/gcp/en/admin/system-tables/workspaces#workspaces-table-schema" target="_self"&gt;docs&lt;/A&gt; too.&lt;/P&gt;&lt;P&gt;So the practical approach for now could be &lt;STRONG&gt;API → mapping table → billing&lt;/STRONG&gt; enrichment, while also watching for Databricks to add &lt;STRONG&gt;project_id&lt;/STRONG&gt; directly to a system table in the future.&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 08:16:42 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/how-to-bring-the-project-id-info-into-the-databricks-billing/m-p/170122#M56237</guid>
      <dc:creator>data_pulse</dc:creator>
      <dc:date>2026-09-29T08:16:42Z</dc:date>
    </item>
  </channel>
</rss>

