a week ago
Are there plans for a v.2.0 of Unity Catalog? I find the organization of tables in databricks to be very constrained and inflexible. Given that UC tables have their origins in data lakes, you would think they would have brought a lot more flexibile names into UC than what we are forced to use today. The rigid three-level names, in lowercase letters, are just terrible IMO. (Especially when data is arriving into UC from other technologies that have much more flexibile names than what we have in this UC.)
Try googling "how to name databricks catalogs and schemas", and you will find too MANY strategies that are being invented as a workaround, to overcome arbitrary limitations enforced in the databricks environment. They often involve concatenating unrelated concepts together with underscores (<environment>, <medallion>, <business_domain>, <business_unit>, <whatever>, etc).
The most challenging limitations are the restrictive three-level names, and being forced to use lowercase letters at the table level. These limitations are unnecessary. In a large UC environment, names should always be hierarchical, especially if the underlying cloud storage fully supports HNS.
Assuming there are certain legacy internals, within databricks, that cannot properly deal with hierarchical names, then why not just support aliases, defined during bootstrapping?
eg....
USING ALIAS mycatalog.presentation.sales = prod_catalog.GoldPresentation.NorthAmerica.Erp01.ExecutiveSummary.OutsideSales
SELECT * FROM mycatalog.presentation.sales
I think aliases are very important. It is silly for end end users and reporting analysts to constantly sprinkle terms like "production", or "gold" all over their queries and notebooks, when they never use anything different (and some probably don't even have access to anything different.)
Friday
Given the lack of flexibility in the naming of schema and tables, I decided to fall back on external volumes. NOTE: I will not use these external volumes for any of the structured tables in the gold/presentation layer (even though I really dislike the rigid names over there too). I will ONLY use external volumes for lower layers that have a smaller audience (silver/bronze/temp/etc).
The nice thing about external volumes is that I can still put any table format in there (delta/parquet), I'm not restricted to the three arbitrarly naming levels, and I'm no longer restricted to lowercase object names either. I suppose we'll lose some UC governance features and other things. But in any case, most of those can be better accommodated in the gold layer when needed. I think it is common for customers to do something slightly different when it comes to storing data in the lower medallion layers.
Below is what is possible after we have escaped the rigid managed table environment, and start hosting our bronze in Volumes. It unshackles us from the arbitrary naming strategies needed when creating UC managed tables.
%sql
-- Querying a Delta file from Volumes
SELECT * FROM delta.`/Volumes/prod_catalog/default/my_ext_volume/bronze/NorthAmerica/Erp01/ExecutiveSummary/OutsideSales`;
Friday
Hello,
I agree that the three-level naming structure can feel pretty restrictive, especially in large environments with existing naming conventions. The alias idea is interesting because it could give users cleaner, more meaningful names without requiring Databricks to completely change the underlying UC structure. Hopefully something like this gets considered in a future UC update.
Friday
Given the lack of flexibility in the naming of schema and tables, I decided to fall back on external volumes. NOTE: I will not use these external volumes for any of the structured tables in the gold/presentation layer (even though I really dislike the rigid names over there too). I will ONLY use external volumes for lower layers that have a smaller audience (silver/bronze/temp/etc).
The nice thing about external volumes is that I can still put any table format in there (delta/parquet), I'm not restricted to the three arbitrarly naming levels, and I'm no longer restricted to lowercase object names either. I suppose we'll lose some UC governance features and other things. But in any case, most of those can be better accommodated in the gold layer when needed. I think it is common for customers to do something slightly different when it comes to storing data in the lower medallion layers.
Below is what is possible after we have escaped the rigid managed table environment, and start hosting our bronze in Volumes. It unshackles us from the arbitrary naming strategies needed when creating UC managed tables.
%sql
-- Querying a Delta file from Volumes
SELECT * FROM delta.`/Volumes/prod_catalog/default/my_ext_volume/bronze/NorthAmerica/Erp01/ExecutiveSummary/OutsideSales`;
Saturday
Hi @DB1To3 ,
No official "Unity Catalog v2.0" has been announced. Databricks evolves UC through continuous feature releases rather than major version jumps, so there's no roadmap item labeled "v2.0" that I'm aware of (as of my knowledge cutoff).
On your specific complaints:
Three-level namespace (catalog.schema.table) โ this is a deliberate design choice, not a bug. It's fixed and unlikely to change; Databricks has stuck with this structure since UC's launch and it's baked into most tooling (SQL parsers, BI connectors, permissions model).
Lowercase-only naming โ also a known limitation people complain about. No public signal that this is changing.
Aliases for shorter/friendlier names โ this doesn't exist as a native "USING ALIAS" feature like you described. The closest workarounds today are:
Bottom line: Your frustration is a well-known and frequently-raised pain point in the Databricks community. There's no evidence of a fix on the near-term roadmap. Your suggestion (bootstrap-time aliases) is a reasonable idea worth submitting as formal product feedback โ this kind of structural request typically needs to go through Databricks' official feedback channels or your account team rather than get addressed via community forum visibility alone.
Monday - last edited Monday
There are cases where these issues have a more pronounced appearance of being buggy. Eg. when creating federation to an external database. If we must create links to SQL Server (with its database-scoped schemas and uppercase table names) then the federated database are almost unrecognizable. It is not a good experience, and was probably never good from its inception.
IMO, There are many examples of technologies that forget to implement the concept of namespaces on the first pass, and then circle back to introduce it later. I think everyone agrees that it is an unfriendly experience to concatenate unrelated concepts together with underscores. These naming workarounds lead to the ugliest names I have ever seen in any platform or technology. (especially for those of us that don't come from the postgres ecosystem. Databricks seems to have a special affinity to data originating there)
Even if we concede that the catalog and schema levels are a difficult place to introduce hierarchical names, I see no reason why that would apply to tables as well. It doesn't seem like there is anything preventing databricks from making improvements. We need a path from erp01_executivesales_outsidesales to Erp01.ExecutiveSummary.OutsideSales
Saturday
Fair points, especially for data coming from systems with richer naming. I can't speak to the roadmap, but here's what works for us today: