ThomazNeto
New Contributor II

Been digging into this since DAIS, so I'll share what I've found — with the big caveat up front that the Ontology is still in preview and the public docs are thin. Some of your questions genuinely don't have documented answers yet, and I think that's worth saying plainly rather than guessing.

On the "automated vs manual" framing: I'd argue it's a bit of a false dichotomy once you look at how it actually works. The graph is learned automatically (from tables, queries, dashboards, pipelines, connected apps), but it's anchored on manually curated artifacts. UC metric views, glossary terms, certified assets — those get treated as high-authority sources in the ranking. Databricks themselves still say a Genie Space is configured by a SQL-proficient analyst, and the manual surface area (measures with names, SQL calculations, synonyms) is not small. So it's less "the machine designs your ontology" and more "the machine fills in everything around the definitions you bothered to write down." The more semantics you model in UC, the more it has to learn from.

The drawbacks I'd actually worry about:

- OntoRank leans on usage signals (where a definition came from, how often it's used, how fresh it is). New org or new dataset = no usage history, so ranking falls back to creator authority and recency. Early results may be shakier than the demo suggests.
- Popular isn't the same as correct. The ranking can surface a widely-used definition that's wrong for the specific question, and whether a calculation is actually valid lives in transformation code — which a popularity graph doesn't read.
- The signals are public but the algorithm isn't. What happens when two high-authority sources genuinely disagree? Not published.
- The one that bothers me most from a governance angle: the ontology updates itself without surfacing what changed. Citations tell you which sources fed an answer, but there's no visible diff or approval flow for what it just "learned." It's a semantic layer that evolves on its own and only becomes visible one citation at a time.

Your practical questions:

Export/snapshot: I couldn't find any documented mechanism for this, and I'm not alone — others have raised exactly this (snapshot the state, prove who approved a change) and the honest answer right now is "not documented, ask your account team." I'd treat that as "not documented as of July 2026" rather than "doesn't exist," given how fast preview features move.

Inspecting sources: partially, at the answer level — you click the citation icons in a response to see which knowledge snippets fed it. What I could NOT find is ontology-level inspection: browsing the graph and seeing the provenance of a specific definition or relationship independent of a question. Maybe it exists behind the preview, but it's not in the docs I can see.

Preparation: it works with zero prep (that's the pitch), but quality clearly scales with what you've curated. Metric views are the big one — they're explicitly treated as a trusted, high-authority input. Glossary and Domains feed it too. There's also a simple lever people miss: workspace admins can drop a Markdown file with custom instructions (data conventions, preferred terminology) that applies to every chat — it's in Beta, capped at 20k characters. On the Knowledge Store specifically as an ontology input, I've only seen third-party writeups, no official doc describing its weight in the graph, so I won't claim details there.

Bottom line for me: the risk isn't really "the machine curated it wrong." It's that the curation is implicit and, today, not auditable. If you're in a regulated space, I'd push your account team hard on the export/lineage questions before trusting it with production workflows.

Thomaz A. Rossito Neto
Principal Data & AI — CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto