3 weeks ago
Hello,
I'm trying to ingest data from my RDS instance and set up change data capture on Databricks. Everything on the Postgres side has been done so that replication is possible, but CDC is still greyed out. I have asked Claude, Gemini, and ChatGPT. I haven't got any concrete help. I read that it's an issue of Databricks needing to enable Lakeflow Connect for Postgres or something like.
P.S.: I might not have explained it well, but I really need help with it.
3 weeks ago
Hi Wola,
Don't worry about your Postgres side, that's most likely not the problem. The greyed-out option is coming from Databricks. The PostgreSQL CDC connector in Lakeflow Connect is still in Public Preview, and this one isn't self-service: the docs say "Contact your Databricks account team to request access." There is nothing to flip on the Previews page. Until the workspace is enrolled, the wizard only lets you pick Query-based capture, which is exactly what your screenshot shows.
https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/postgresql
So step one is a message to your account team asking to enable the PostgreSQL connector preview for your workspace.
While you wait, two things worth checking. The CDC pipeline needs serverless compute enabled in the workspace (the gateway runs on classic compute, the pipeline itself is serverless), and it only replicates from a primary instance, read replicas won't work.
https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/postgresql-pipeline
And the RDS checklist from the docs, just to compare with what you've done: rds.logical_replication = 1 in the parameter group, a replication user with the rds_replication role and the REPLICATION attribute, a publication, a replication slot created with pgoutput (the only plugin supported), and REPLICA IDENTITY set to DEFAULT or FULL on every table. The slot name and publication name are what the wizard asks for later, in the Database setup step.
https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/postgresql-source-setup
https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/postgresql-limits
If you can't wait, Query-based capture does work with Postgres today, no preview needed. Just know what you're getting: it doesn't read the WAL, it runs on a schedule and picks up changes through a cursor column (a timestamp or increasing id) per table. You won't see intermediate states of a row, and hard deletes are only tracked in Beta. Fine for a lot of use cases, but it isn't real CDC.
https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/query-based-overview
Good luck with it.
3 weeks ago
Thank you very much!
It's an individual account. So when I saw the message about reaching out to your Databricks account team so they can reach out to your Databricks rep, I was a bit confused, as I'm the account owner. I'm learning and didn't want to use a free account, so I set up a paid individual account.
Regarding the serverless compute for CDC, I read that serverless was enabled by default in most workspaces, but with the aid of Claude, I set a policy for the classic compute indicating the instance family and other things needed, if at all, classic compute was needed.
Once again, thank you very much.
I was already frustrated.
3 weeks ago
Quick question, being that my account is an individual account, how do I go about this "Contact your Databricks account team to request access"
Thursday
Greetings @Wola, I did some digging and here is what I found.
@ThomazNeto already pointed you at the two things that matter most: the PostgreSQL connector for Lakeflow Connect is in Public Preview and needs workspace enrollment, and logical replication only works against a primary instance. I'll build on that and answer your account team question directly.
First, the greyed-out CDC option. That's the preview gate, not your RDS configuration or your compute policy. The policy you built isn't wasted (the ingestion gateway runs on classic compute and will use it), but no change on your side flips that selector. Only enrollment does.
On "contact your account team" when you are the account: with a pay-as-you-go individual account you don't have a named rep, and the docs don't offer a self-service path for this preview. Three things to try, in order:
Once enrolled, the workspace side is mostly done already: Unity Catalog and serverless enabled (you're right that serverless is on by default in most workspaces, and the ingestion pipeline runs there while the gateway runs on classic). Confirm you also have CREATE CONNECTION on the metastore and USE CATALOG, USE SCHEMA, CREATE TABLE, and CREATE VOLUME on the target.
On the RDS side, you've likely covered most of this, but these are the ones that trip people up:
rds.logical_replication = 1 in the parameter group, reboot, and SHOW wal_level; returns logical.ALTER USER ... WITH REPLICATION, GRANT rds_replication TO your_user; (RDS-specific and easy to miss), plus CONNECT, USAGE, and SELECT on the tables.pgoutput plugin (the only one Databricks supports). Publication first, then slot, and create the slot while running as the replication user via SET ROLE.REPLICA IDENTITY FULL on tables without a primary key or with large TEXT/BYTEA columns; DEFAULT otherwise.max_slot_wal_keep_size set to a finite value so an idle slot can't bloat WAL.When the button lights up, the wizard is Data Ingestion > Add data > PostgreSQL, and it asks for the slot and publication names on the Database setup page.
If you need something while you wait, Lakeflow Connect's query-based connector can ingest from PostgreSQL with no WAL or CDC setup. It runs on a schedule using a cursor column, so it captures the latest state of changed rows rather than every intermediate change, and it isn't a substitute for log-based CDC. Hard-delete tracking is in Beta and needs API configuration.
Hang in there. You did the hard part on the Postgres side; the piece that's blocking you is the one piece you can't do yourself.
References:
Regards, Louis.