4 weeks ago
Hi all,
Running into what looks like a platform bug with Lakebase Postgres (Public Preview) and hoping a community engineer can help, since my Databricks support case (#01009170) was closed as out of entitlement (personal account, no support contract) before anyone could look at the actual issue.
Setup: project `ontobricks-demo`, branch `production` (default), endpoint `primary`. The branch auto-archived after about two weeks of inactivity ("Automatically archived ... due to inactivity" on the Branch overview page).
What I did: opened the branch, saw the "This branch is archived. Connecting to the branch will unarchive it." banner, and used the branch-level "Connect" button, which generated a working-looking connection string and showed the compute as "primary • Idle". I also enabled branch protection ("Protect").
What's still broken, over an hour later and after multiple retries/reloads:
- The Computes tab still shows `primary` as `SUSPENDED`.
- Monitoring → System operations shows "Timeline unarchive" as `OK`, but no "Start compute" operation ever follows it (compare to the branch's initial creation, where "Create timeline" was immediately followed by "Start compute").
- Every real Postgres connection — from my own application, and from the Databricks SQL Editor itself — is rejected with:
*** ERROR: The endpoint has been disabled. Enable it using the API and retry. ***
- In the SQL Editor, the compute selector shows "No computes" collapsed, and even though it lists `primary ● Idle` when expanded, selecting it doesn't let me actually run a query.
- A separate feature on the same project, Data API, independently returns "Temporarily Unavailable."
So it looks like the branch-level unarchive succeeded, but the compute endpoint itself never actually resumed, and something in that project may be in a stuck/inconsistent state (control plane says "Idle", data plane refuses connections).
Has anyone seen this, or know the right way to force a full compute resume after an archive → unarchive cycle? I'd rather not delete/recreate the `primary` compute blind, since I'm not sure whether that's safe here or would just reproduce the same stuck state.
Thanks!
4 weeks ago
Hi ,
Thank you for explaining the entire problem, the main key clue for me is “ERROR: The endpoint has been disabled. Enable it using the API and retry.”
This looks different from a normal Lakebase scale-to-zero suspension to me
A compute that has simply scaled to zero should be able to wake up from a new connection. However, if the endpoint is actually marked as disabled, connection attempts (including from the Lakebase UI / SQL Editor) will not wake it. It needs to be explicitly re-enabled through the Lakebase API.
For your project, I’d first check the endpoint state:
databricks postgres get-endpoint \
projects/ontobricks-demo/branches/production/endpoints/primary \
--output json | jq '.status.disabled'If that returns true, you can re-enable it with:
databricks postgres update-endpoint \
projects/ontobricks-demo/branches/production/endpoints/primary \
spec.disabled \
--json '{
"spec": {
"disabled": false
}
}'Then give the compute a moment to finish starting and retry the connection.
Once you’ve re-enabled it, give the compute a little time to come back up and then try connecting again.
I wouldn’t delete or recreate the primary compute at this stage. From what you’ve described, the branch itself seems to have been unarchived correctly, so I’d avoid making any destructive changes until we know what state the endpoint is actually in.
If status.disabled is already showing false, but you’re still getting the “endpoint has been disabled” error, then this does sound more like the endpoint is stuck in an inconsistent state. In that case, I’d try restarting the compute and check the System Operations timeline to see whether a proper start operation gets triggered.
The Data API being unavailable at the same time may also be related to the compute not being fully available, rather than a completely separate issue.
I also wouldn’t worry too much about branch protection here that shouldn’t be needed to bring the compute back online, so I’d treat that as separate from this problem.
Hope this helps and @jeremiasInetum If you can share what status.disabled returns after the unarchive, that should tell us pretty quickly which path we’re dealing with.
2 weeks ago
Hi Gokul,
Thank you — this worked perfectly. I ran the two CLI calls you described:
get-endpoint confirmed status.disabled was true, and after update-endpoint
with "disabled": false the compute came back as IDLE right away. Restarted
our app afterwards and everything is back — all domains and ontologies
intact, no data lost.
Really appreciate you spelling out that this needed the API specifically
(and not just a connection attempt or the "Connect" button in the UI) —
that distinction wasn't clear anywhere in the product itself, and it's
what had us stuck for two weeks. Also good call on not deleting/recreating
the compute blind; glad we didn't go down that path.
Thanks again!
3 weeks ago - last edited 3 weeks ago
I also faced the same error as @jeremiasInetum. So I applied the steps from @Gokul_Pillai1 and status.disabled became false. After that, I navigate to branches in Lakebase. In branch overview, click on Protect button under Disable archiving.
Click Set as Protected will unarchive branch. This resolve my issue though I'm unsure what is the root cause.
2 weeks ago
Hi Albert,
Thanks for confirming — good to know this wasn't a one-off on my end.
In my case the same CLI steps (disabled: true → false) were enough by
themselves and the compute came back right away, without needing to
touch "Protect" again at that point. That said, I can't fully rule out
that it mattered: I had already enabled branch protection days earlier,
during my first round of troubleshooting, so my case isn't a clean test
of whether "Protect" makes a difference on its own.
Might be worth someone from Databricks clarifying whether branch
protection state affects how re-enabling the endpoint behaves — could
save the next person a step either way.
Thanks for sharing your experience!