<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: HubSpot connector in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172311#M56625</link>
    <description>&lt;P&gt;Greetings &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/6280"&gt;@marcell_nagy&lt;/a&gt;, I did some digging and here is what I found.&lt;/P&gt;
&lt;P&gt;I think there are two separate issues here.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Missing columns (&lt;CODE&gt;portalSubscriptionStatus&lt;/CODE&gt;, &lt;CODE&gt;source&lt;/CODE&gt;, and custom fields)&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;In HubSpot's Email Events API, &lt;CODE&gt;portalSubscriptionStatus&lt;/CODE&gt; and &lt;CODE&gt;source&lt;/CODE&gt; aren't general event properties. They only exist on &lt;CODE&gt;STATUSCHANGE&lt;/CODE&gt; events. The managed connector flattens events into a fixed schema, and the docs say automated schema evolution for new and deleted columns isn't supported, so "Include future columns" won't pull in type-specific properties the connector doesn't already expose. The same applies to CRM custom properties: column selection only works on connector-exposed columns, and the pipeline definition has no way to request an arbitrary source property by name. So the straight answer is: not through the managed connector today.&lt;/P&gt;
&lt;P&gt;Two things to check before giving up on it, though. The connector has an incremental &lt;CODE&gt;email_subscription_change&lt;/CODE&gt; table that may already carry the subscription status data you're after. And the limitations page notes that nested or variable HubSpot fields are stored as strings, so look at whether your &lt;CODE&gt;email_events&lt;/CODE&gt; table has a string column holding the raw payload you could parse with &lt;CODE&gt;from_json&lt;/CODE&gt; downstream. I haven't seen a published column list for that table, so I'd want you to confirm what you actually got.&lt;/P&gt;
&lt;P&gt;On custom fields specifically: those live on CRM objects, and CRM Hub ingestion is in Beta behind the &lt;CODE&gt;hubspot_connector_crm_objects&lt;/CODE&gt; workspace preview. I can't confirm from the docs whether the Beta pulls custom properties. Your account team can get you into the preview and route a feature request for explicit column pass-through to the ingestion team. Your post describes the use case well.&lt;/P&gt;
&lt;P&gt;Workaround in the meantime: a small custom ingestion job that calls the Email Events API directly with &lt;CODE&gt;startTimestamp&lt;/CODE&gt; as the cursor (filter to &lt;CODE&gt;eventType=STATUSCHANGE&lt;/CODE&gt;), lands the raw JSON in a bronze Delta table, and flattens the properties downstream. Keep the HubSpot token in a secret scope or connection, not in the pipeline YAML.&lt;/P&gt;
&lt;OL start="2"&gt;
&lt;LI&gt;Is &lt;CODE&gt;email_events&lt;/CODE&gt; incremental?&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Per the reference page, yes. The first run is always a full historical load, but later runs should use a cursor. If you're seeing full pulls after that, two things will help pin it down. Query the pipeline event log and compare rows read against rows written for the &lt;CODE&gt;email_events&lt;/CODE&gt; flow across a few runs (&lt;CODE&gt;flow_progress&lt;/CODE&gt; events carry those metrics). That tells you whether it's source-side retrieval, destination writes, or change processing. And check whether the slowness is actually rate limiting: HubSpot caps API calls (the docs cite 100 requests per 10 seconds for free apps) and the connector retries silently, which can make a modest incremental pull look like a full one. If the event log shows full reads, open a support ticket with the pipeline ID, table, run timestamps, and row counts. That's a bug, not a feature gap.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;HubSpot connector overview and feature table: &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-overview" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;HubSpot connector reference (incremental vs batch tables): &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-reference" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-reference&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;HubSpot connector limitations: &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-limits" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-limits&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Troubleshoot the HubSpot connector (slow pipelines, rate limits): &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-troubleshoot" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-troubleshoot&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Ingest data from HubSpot (pipeline definition examples): &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-pipeline" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-pipeline&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;HubSpot Email Events API overview: &lt;A href="https://developers.hubspot.com/docs/api-reference/legacy/reporting/email-analytics/guide" target="_blank"&gt;https://developers.hubspot.com/docs/api-reference/legacy/reporting/email-analytics/guide&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Hope this helps.&lt;/P&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
    <pubDate>Thu, 08 Oct 2026 15:27:07 GMT</pubDate>
    <dc:creator>Louis_Frolio</dc:creator>
    <dc:date>2026-10-08T15:27:07Z</dc:date>
    <item>
      <title>HubSpot connector</title>
      <link>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172287#M56618</link>
      <description>&lt;P&gt;The &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ingestion/lakeflow-connect/hubspot-overview" target="_self"&gt;HubSpot Support Lakeflow connector&lt;/A&gt; is able to download basic datasets. However, I see no way to get non-basic columns:&lt;/P&gt;&lt;P&gt;1. The &lt;A href="https://developers.hubspot.com/docs/api-reference/legacy/reporting/email-analytics/guide#user-status-events" target="_self"&gt;suggested method&lt;/A&gt; to analyse email events is to use the&amp;nbsp;portalSubscriptionStatus and source properties. The Lakeflow connector does not get these properties, even if the "Include future columns" is checked.&lt;BR /&gt;2. Being able to download columns that are set up in HubSpot as Custom Field would be great.&lt;/P&gt;&lt;P&gt;Can we achieve this right now? If not yet, it would be useful to be able to explicity list column names in the pipeline YAML even if the connector is not aware of that, but could dynamically download additional columns.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;On the other hand, I think the email events dataset is not incremental, even if the &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ingestion/lakeflow-connect/hubspot-reference" target="_self"&gt;reference&lt;/A&gt; tells that. Seemingly the connector awlays downloads all records but only stores the difference, it makes the pipeline slower than reasonable.&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 12:43:26 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172287#M56618</guid>
      <dc:creator>marcell_nagy</dc:creator>
      <dc:date>2026-10-08T12:43:26Z</dc:date>
    </item>
    <item>
      <title>Re: HubSpot connector</title>
      <link>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172311#M56625</link>
      <description>&lt;P&gt;Greetings &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/6280"&gt;@marcell_nagy&lt;/a&gt;, I did some digging and here is what I found.&lt;/P&gt;
&lt;P&gt;I think there are two separate issues here.&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Missing columns (&lt;CODE&gt;portalSubscriptionStatus&lt;/CODE&gt;, &lt;CODE&gt;source&lt;/CODE&gt;, and custom fields)&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;In HubSpot's Email Events API, &lt;CODE&gt;portalSubscriptionStatus&lt;/CODE&gt; and &lt;CODE&gt;source&lt;/CODE&gt; aren't general event properties. They only exist on &lt;CODE&gt;STATUSCHANGE&lt;/CODE&gt; events. The managed connector flattens events into a fixed schema, and the docs say automated schema evolution for new and deleted columns isn't supported, so "Include future columns" won't pull in type-specific properties the connector doesn't already expose. The same applies to CRM custom properties: column selection only works on connector-exposed columns, and the pipeline definition has no way to request an arbitrary source property by name. So the straight answer is: not through the managed connector today.&lt;/P&gt;
&lt;P&gt;Two things to check before giving up on it, though. The connector has an incremental &lt;CODE&gt;email_subscription_change&lt;/CODE&gt; table that may already carry the subscription status data you're after. And the limitations page notes that nested or variable HubSpot fields are stored as strings, so look at whether your &lt;CODE&gt;email_events&lt;/CODE&gt; table has a string column holding the raw payload you could parse with &lt;CODE&gt;from_json&lt;/CODE&gt; downstream. I haven't seen a published column list for that table, so I'd want you to confirm what you actually got.&lt;/P&gt;
&lt;P&gt;On custom fields specifically: those live on CRM objects, and CRM Hub ingestion is in Beta behind the &lt;CODE&gt;hubspot_connector_crm_objects&lt;/CODE&gt; workspace preview. I can't confirm from the docs whether the Beta pulls custom properties. Your account team can get you into the preview and route a feature request for explicit column pass-through to the ingestion team. Your post describes the use case well.&lt;/P&gt;
&lt;P&gt;Workaround in the meantime: a small custom ingestion job that calls the Email Events API directly with &lt;CODE&gt;startTimestamp&lt;/CODE&gt; as the cursor (filter to &lt;CODE&gt;eventType=STATUSCHANGE&lt;/CODE&gt;), lands the raw JSON in a bronze Delta table, and flattens the properties downstream. Keep the HubSpot token in a secret scope or connection, not in the pipeline YAML.&lt;/P&gt;
&lt;OL start="2"&gt;
&lt;LI&gt;Is &lt;CODE&gt;email_events&lt;/CODE&gt; incremental?&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Per the reference page, yes. The first run is always a full historical load, but later runs should use a cursor. If you're seeing full pulls after that, two things will help pin it down. Query the pipeline event log and compare rows read against rows written for the &lt;CODE&gt;email_events&lt;/CODE&gt; flow across a few runs (&lt;CODE&gt;flow_progress&lt;/CODE&gt; events carry those metrics). That tells you whether it's source-side retrieval, destination writes, or change processing. And check whether the slowness is actually rate limiting: HubSpot caps API calls (the docs cite 100 requests per 10 seconds for free apps) and the connector retries silently, which can make a modest incremental pull look like a full one. If the event log shows full reads, open a support ticket with the pipeline ID, table, run timestamps, and row counts. That's a bug, not a feature gap.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;HubSpot connector overview and feature table: &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-overview" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;HubSpot connector reference (incremental vs batch tables): &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-reference" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-reference&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;HubSpot connector limitations: &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-limits" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-limits&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Troubleshoot the HubSpot connector (slow pipelines, rate limits): &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-troubleshoot" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-troubleshoot&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Ingest data from HubSpot (pipeline definition examples): &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-pipeline" target="_blank"&gt;https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/hubspot-pipeline&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;HubSpot Email Events API overview: &lt;A href="https://developers.hubspot.com/docs/api-reference/legacy/reporting/email-analytics/guide" target="_blank"&gt;https://developers.hubspot.com/docs/api-reference/legacy/reporting/email-analytics/guide&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Hope this helps.&lt;/P&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 15:27:07 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172311#M56625</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-08T15:27:07Z</dc:date>
    </item>
    <item>
      <title>Re: HubSpot connector</title>
      <link>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172444#M56641</link>
      <description>&lt;P&gt;Can you please explain these parts:&lt;/P&gt;&lt;P&gt;"&lt;SPAN&gt;Your account team can get you into the preview and route a feature request" - how-where should I contact them?&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;"If the event log shows full reads, open a support ticket with the pipeline ID, table, run timestamps, and row counts." - where to open it? We're in Azure, you mean Help+Support system inside Azure Portal?&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 09 Oct 2026 12:25:44 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172444#M56641</guid>
      <dc:creator>marcell_nagy</dc:creator>
      <dc:date>2026-10-09T12:25:44Z</dc:date>
    </item>
    <item>
      <title>Re: HubSpot connector</title>
      <link>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172464#M56643</link>
      <description>&lt;P&gt;Greetings &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/6280"&gt;@marcell_nagy&lt;/a&gt;, Fair questions, I should've been more specific.&lt;/P&gt;
&lt;P&gt;Account team. That's the Databricks people assigned to your company: account executive, solutions architect, or customer success contact. Separate from your HubSpot admin. If you don't know who they are, ask whoever at your company owns the Azure Databricks spend, or ask your Microsoft account rep to connect you. One thing to try first for the CRM preview: have a workspace admin check the Previews page under workspace settings. Some previews can be switched on there without involving anyone.&lt;/P&gt;
&lt;P&gt;Support ticket. Two routes, and which one applies depends on your support contract. If your company routes Databricks support through Microsoft (the common setup on Azure), use the Azure Portal: Help + support, Create a support request, service Azure Databricks. Microsoft escalates to Databricks engineering when it's a product issue. You'll need an Azure support plan that includes technical support. If your company has a direct Databricks support contract, an authorized contact can file through the Databricks Help Center instead, reachable from the Help menu in your workspace. If you're not sure which you have, your account contact can confirm.&lt;/P&gt;
&lt;P&gt;For the &lt;CODE&gt;email_events&lt;/CODE&gt; ticket, include: workspace URL and region, pipeline ID and destination table, the affected update IDs and timestamps, whether each run was an initial load, normal update, or full refresh, rows read versus rows written, and the relevant event log excerpt. Leave HubSpot tokens and other credentials out of the ticket and out of this thread.&lt;/P&gt;
&lt;P&gt;Feature request. Quickest route is in the workspace: click your username icon in the top bar, then Send feedback. Describe the use case (explicit column pass-through for the HubSpot connector) and tick the box allowing Databricks to contact you. That goes straight to the product team.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Azure support options: &lt;A href="https://azure.microsoft.com/support" target="_blank"&gt;https://azure.microsoft.com/support&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Azure support plans: &lt;A href="https://azure.microsoft.com/en-us/support/plans/" target="_blank"&gt;https://azure.microsoft.com/en-us/support/plans/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Submit product feedback from your workspace: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/resources/ideas" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/resources/ideas&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks Help Center (direct support contracts): &lt;A href="https://help.databricks.com" target="_blank"&gt;https://help.databricks.com&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks support process: &lt;A href="https://docs.databricks.com/aws/en/resources/support" target="_blank"&gt;https://docs.databricks.com/aws/en/resources/support&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Fri, 09 Oct 2026 13:47:28 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/hubspot-connector/m-p/172464#M56643</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-09T13:47:28Z</dc:date>
    </item>
  </channel>
</rss>

