cancel
Showing results for 
Search instead for 
Did you mean: 
Knowledge Sharing Hub
Dive into a collaborative space where members like YOU can exchange knowledge, tips, and best practices. Join the conversation today and unlock a wealth of collective wisdom to enhance your experience and drive success.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

SumitSingh
by Contributor
  • 4645 Views
  • 7 replies
  • 13 kudos

From Associate to Professional: My Learning Plan to ace all Databricks Data Engineer Certifications

In today’s data-driven world, the role of a data engineer is critical in designing and maintaining the infrastructure that allows for the efficient collection, storage, and analysis of large volumes of data. Databricks certifications holds significan...

SumitSingh_0-1721402402230.png SumitSingh_1-1721402448677.png SumitSingh_2-1721402469214.png
  • 4645 Views
  • 7 replies
  • 13 kudos
Latest Reply
sandeepmankikar
Contributor
  • 13 kudos

As an additional tip for those working towards both the Associate and Professional certifications, I recommend avoiding a long gap between the two exams to maintain your momentum. If possible, try to schedule them back-to-back with just a few days in...

  • 13 kudos
6 More Replies
WarrenO
by New Contributor III
  • 3051 Views
  • 1 replies
  • 1 kudos

Resolved! Log Custom Transformer with Feature Engineering Client

Hi everyone,I'm building a Pyspark ML Pipeline where the first stage is to fill nulls with zero. I wrote a custom class to do this since I cannot find a Transformer that will do this imputation. I am able to log this pipeline using ML Flow log model ...

Knowledge Sharing Hub
Custom Transformer
feature engineering
ML FLow
pipeline
pyspark
  • 3051 Views
  • 1 replies
  • 1 kudos
Latest Reply
koji_kawamura
Databricks Employee
  • 1 kudos

Hi @WarrenO , thanks for sharing that with the detailed code! I was able to reproduce the error, specifically the following error: AttributeError: module '__main__' has no attribute 'CustomAdder'File <command-1315887242804075>, line 3935 evaluator = ...

  • 1 kudos
himoshi
by New Contributor II
  • 4771 Views
  • 3 replies
  • 0 kudos

Error code 403 - Invalid access to Org

I am trying to make a GET /api/2.1/jobs/list call in a Notebook to get a list of all jobs in my workspace but am unable to do so due to a 403 "Invalid access to Org" error message. I am using a new PAT and the endpoint is correct. I also have workspa...

  • 4771 Views
  • 3 replies
  • 0 kudos
Latest Reply
xmad2772
New Contributor II
  • 0 kudos

Hey did you make any progress on the error? I'm experiencing the same in my environment. Thanks! 

  • 0 kudos
2 More Replies
yadvendra_ksh
by New Contributor II
  • 339 Views
  • 0 replies
  • 0 kudos

The Hidden Security Risks in Stored Procedure Migrations—What Databricks Exposed

Your stored procedure migration to DB isn't just a 'copy-paste' job - it's a security nightmare waiting to happen.We discovered our 'trusted' stored procedures had hidden access patterns that nearly compromised our entire data governance model. Here'...

  • 339 Views
  • 0 replies
  • 0 kudos
Ajay-Pandey
by Esteemed Contributor III
  • 1375 Views
  • 1 replies
  • 1 kudos

📊 Simplifying CDC with Databricks Delta Live Tables & Snapshots 📊

In the world of data integration, synchronizing external relational databases (like Oracle, MySQL) with the Databricks platform can be complex, especially when Change Data Feed (CDF) streams aren’t available. Using snapshots is a powerful way to mana...

Pull-Based Snapshots.png
  • 1375 Views
  • 1 replies
  • 1 kudos
Latest Reply
BilalHaniff1
New Contributor II
  • 1 kudos

Hi AjayCan apply changes into snapshot handle re-processing of an older snapshot? UseCase:- Source has delivered data on day T, T1 and T2.  - Consumers realise there is an error on the day T data, and make a correction in the source. The source redel...

  • 1 kudos
ChsAIkrishna
by Contributor
  • 1082 Views
  • 1 replies
  • 4 kudos

Consideration Before Migrating Hive Tables to Unity Catalog

Databricks recommends four methods to migrate Hive tables to Unity Catalog, each with its pros and cons. The choice of method depends on specific requirements.SYNC: A SQL command that migrates schema or tables to Unity Catalog external tables. Howeve...

highresrollsafe piz.PNG
  • 1082 Views
  • 1 replies
  • 4 kudos
Latest Reply
Mantsama4
Valued Contributor
  • 4 kudos

This is a great solution! The post effectively outlines the methods for migrating Hive tables to Unity Catalog while emphasizing the importance of not just performing a simple migration but transforming the data architecture into something more robus...

  • 4 kudos
MichTalebzadeh
by Valued Contributor
  • 3645 Views
  • 3 replies
  • 3 kudos

Resolved! Feature Engineering for Data Engineers: Building Blocks for ML Success

For a  UK Government Agency, I made a Comprehensive presentation titled " Feature Engineering for Data Engineers: Building Blocks for ML Success".  I made an article of it in Linkedlin together with the relevant GitHub code. In summary the code delve...

Knowledge Sharing Hub
feature engineering
ML
python
  • 3645 Views
  • 3 replies
  • 3 kudos
Latest Reply
Mantsama4
Valued Contributor
  • 3 kudos

This is a fantastic post! The detailed explanation of feature engineering, from handling missing values to using Variational Autoencoders (VAEs) for synthetic data generation, provides invaluable insights for improving machine learning models. The ap...

  • 3 kudos
2 More Replies
Harun
by Honored Contributor
  • 8351 Views
  • 3 replies
  • 5 kudos

Comprehensive Guide to Databricks Optimization: Z-Order, Data Compaction, and Liquid Clustering

Optimizing data storage and access is crucial for enhancing the performance of data processing systems. In Databricks, several optimization techniques can significantly improve query performance and reduce costs: Z-Order Optimize, Optimize Compaction...

  • 8351 Views
  • 3 replies
  • 5 kudos
Latest Reply
Mantsama4
Valued Contributor
  • 5 kudos

I also have the same question!

  • 5 kudos
2 More Replies
Mantsama4
by Valued Contributor
  • 2649 Views
  • 0 replies
  • 0 kudos

How can Databricks AI/BI Genie, RAG, & LLMs seamlessly coexist with MS Copilot to drive innovation?

The future of enterprise productivity and analytics lies in the seamless integration of advanced tools like Databricks Genie AI/BI, RAG & LLMs and Microsoft Copilot. While each serves distinct purposes, their coexistence can unlock unparalleled value...

  • 2649 Views
  • 0 replies
  • 0 kudos
Mantsama4
by Valued Contributor
  • 827 Views
  • 0 replies
  • 1 kudos

How Databricks Empowers Scalable Data Products Through Medallion Mesh Architecture?

Unlock the Power of Your Data: Solving Fragmentation and Governance Challenges!In today’s fast-paced, data-driven enterprises, fragmented data and governance issues create roadblocks to decision-making and innovation. Traditional architectures strugg...

  • 827 Views
  • 0 replies
  • 1 kudos
Mantsama4
by Valued Contributor
  • 430 Views
  • 0 replies
  • 0 kudos

Rebuilding and Re-Platforming Your Databricks Lakehouse with Serverless Compute

Dear Databricks Community,In today’s fast-paced data landscape, managing infrastructure manually can slow down innovation, increase costs, and limit scalability. Databricks Serverless Compute solves these challenges by eliminating infrastructure over...

  • 430 Views
  • 0 replies
  • 0 kudos
nk25
by New Contributor II
  • 1357 Views
  • 3 replies
  • 0 kudos

Getting data from Databricks into Excel using Databricks Jobs API

If you have your data in Databricks, but like to analyse it in Excel, you can use Web API on Power Query. It allows you to not just query an existing table, but also trigger the execution of a PySpark notebook using Databricks Jobs API, and get the d...

  • 1357 Views
  • 3 replies
  • 0 kudos
Latest Reply
NandiniN
Databricks Employee
  • 0 kudos

Got it, yes you have specified the same in your message. Thanks for sharing.

  • 0 kudos
2 More Replies
Takuya-Omi
by Valued Contributor III
  • 1293 Views
  • 1 replies
  • 3 kudos

How to Grant Workspace Admin Permissions to an ID Using Parent Groups

Hello,There are several ways to grant Workspace Admin permissions in Databricks. While this may seem straightforward, I found it a bit confusing when I started using Databricks, so I’d like to share my experience. This guide is aimed at beginners.How...

TakuyaOmi_4-1735401774100.png TakuyaOmi_2-1735401216903.png TakuyaOmi_5-1735401862717.png TakuyaOmi_6-1735402496967.png
  • 1293 Views
  • 1 replies
  • 3 kudos
Latest Reply
Alberto_Umana
Databricks Employee
  • 3 kudos

Thanks for sharing this is great!

  • 3 kudos
jwu1
by Databricks Employee
  • 7555 Views
  • 3 replies
  • 1 kudos

Databricks Academy Labs Coupon Instructions

To see all Databricks training and enablement offerings, please visit our Learning Library and Certifications Catalog. To use your Databricks Academy Labs coupons, please -  Head to Databricks Academy Across the top navigation, select Subscriptions C...

  • 7555 Views
  • 3 replies
  • 1 kudos
Latest Reply
svfarande
New Contributor II
  • 1 kudos

Shekhar you will have to do below : Hi Shubham Thank you for reaching out to Databricks, and we're sorry to see the error message you're facing. Could you kindly raise a support ticket - https://help.databricks.com/s/contact-us?ReqType=training and g...

  • 1 kudos
2 More Replies

Join Us as a Local Community Builder!

Passionate about hosting events and connecting people? Help us grow a vibrant local community—sign up today to get started!

Sign Up Now