Greetings @marcell_nagy, I did some digging and here is what I found.
I think there are two separate issues here.
- Missing columns (
portalSubscriptionStatus, source, and custom fields)
In HubSpot's Email Events API, portalSubscriptionStatus and source aren't general event properties. They only exist on STATUSCHANGE events. The managed connector flattens events into a fixed schema, and the docs say automated schema evolution for new and deleted columns isn't supported, so "Include future columns" won't pull in type-specific properties the connector doesn't already expose. The same applies to CRM custom properties: column selection only works on connector-exposed columns, and the pipeline definition has no way to request an arbitrary source property by name. So the straight answer is: not through the managed connector today.
Two things to check before giving up on it, though. The connector has an incremental email_subscription_change table that may already carry the subscription status data you're after. And the limitations page notes that nested or variable HubSpot fields are stored as strings, so look at whether your email_events table has a string column holding the raw payload you could parse with from_json downstream. I haven't seen a published column list for that table, so I'd want you to confirm what you actually got.
On custom fields specifically: those live on CRM objects, and CRM Hub ingestion is in Beta behind the hubspot_connector_crm_objects workspace preview. I can't confirm from the docs whether the Beta pulls custom properties. Your account team can get you into the preview and route a feature request for explicit column pass-through to the ingestion team. Your post describes the use case well.
Workaround in the meantime: a small custom ingestion job that calls the Email Events API directly with startTimestamp as the cursor (filter to eventType=STATUSCHANGE), lands the raw JSON in a bronze Delta table, and flattens the properties downstream. Keep the HubSpot token in a secret scope or connection, not in the pipeline YAML.
- Is
email_events incremental?
Per the reference page, yes. The first run is always a full historical load, but later runs should use a cursor. If you're seeing full pulls after that, two things will help pin it down. Query the pipeline event log and compare rows read against rows written for the email_events flow across a few runs (flow_progress events carry those metrics). That tells you whether it's source-side retrieval, destination writes, or change processing. And check whether the slowness is actually rate limiting: HubSpot caps API calls (the docs cite 100 requests per 10 seconds for free apps) and the connector retries silently, which can make a modest incremental pull look like a full one. If the event log shows full reads, open a support ticket with the pipeline ID, table, run timestamps, and row counts. That's a bug, not a feature gap.
References:
Hope this helps.
Regards, Louis.