cancel
Showing results for 
Search instead for 
Did you mean: 
Machine Learning
Dive into the world of machine learning on the Databricks platform. Explore discussions on algorithms, model training, deployment, and more. Connect with ML enthusiasts and experts.
cancel
Showing results for 
Search instead for 
Did you mean: 

Forum Posts

ben_ai
by • New Contributor II
  • 295 Views
  • 2 replies
  • 2 kudos

Image Annotation

What is Image Annotation?What are the Steps of Image Annotation?What are the Different Techniques of Image Annotation?Types Used in Image AnnotationHow are Companies Handling Image Annotation?Features to Look for in Image Annotation Service Providers...

  • 295 Views
  • 2 replies
  • 2 kudos
Latest Reply
ThiamLee
Contributor
  • 2 kudos

Great overview of image annotation! The sections on different techniques, real-world use cases, and pricing factors are especially helpful for anyone getting started with AI/ML. Looking forward to exploring the future trends in image annotation! 

  • 2 kudos
1 More Replies
AdamIH123
by • New Contributor III
  • 380 Views
  • 2 replies
  • 1 kudos

Tuning with Optuna and MlflowSparkStudy

I am following the guide for tuning a model with Optuna and MlflowSparkStudy. My compute is configured with autoscaling enabled, with 1–2 Spark workers, each with 8 cores and 32 GB of memory. I set n_jobs=2 and trials=100, in mlflow_study.optimize()....

Machine Learning
mlflow
mlflow_study
mlflow_study.optimize
MlflowSparkStudy
optuna
  • 380 Views
  • 2 replies
  • 1 kudos
Latest Reply
ThiamLee
Contributor
  • 1 kudos

Great questions—especially the distinction between Optuna’s trial-level parallelism and LightGBM’s intra-trial threading. The interaction with Spark autoscaling and data locality is also something I’d love to see documented with a concrete example. C...

  • 1 kudos
1 More Replies
kartheek_rao
by • New Contributor III
  • 475 Views
  • 1 replies
  • 4 kudos

End-to-End Streaming NLP Pipeline with GDELT, Azure Data Factory, ADLS Gen2 and Databricks

Building an End-to-End Streaming NLP Pipeline with GDELT, Azure Data Factory, ADLS Gen2 and DatabricksI recently worked on an end-to-end streaming NLP project using GDELT news data, Azure Data Factory, ADLS Gen2 and Azure Databricks.The goal was not ...

  • 475 Views
  • 1 replies
  • 4 kudos
Latest Reply
kunduruanil
New Contributor III
  • 4 kudos

Great end-to-end project! To take it to the next level, consider exploring these native Databricks capabilities.Lakehouse Monitoring: Track data drift and automate model retraining when performance drops.Model Serving: Deploy your models behind serve...

  • 4 kudos
kartheek_rao
by • New Contributor III
  • 745 Views
  • 4 replies
  • 3 kudos

End-to-End Streaming NLP Pipeline with GDELT, Azure Data Factory, ADLS Gen2 and Databricks

I have been working on a project to understand Databricks end to end, rather than just loading some data and training a model.I picked GDELT news data and the use case is to identify supply chain disruption related news and eventually predict which e...

  • 745 Views
  • 4 replies
  • 3 kudos
Latest Reply
kunduruanil
New Contributor III
  • 3 kudos

@kartheek_rao, you are on the right track.Since you are doing clustering of new articles, it's unsupervised learning; you need to understand the feature engineering part more and the EDA part with MLFlow experiments. You have model monitoring as well...

  • 3 kudos
3 More Replies
ticusss
by • New Contributor II
  • 581 Views
  • 2 replies
  • 1 kudos

Model Serving An internal error occurred during feature store lookup all deploys failing

We can't deploy models to Model Serving. Failures started around 2026-08-16 and were intermittent at first — our current production config deployed cleanly on 08-20 — but since then every attempt fails at the feature store lookup setup step, with no ...

  • 581 Views
  • 2 replies
  • 1 kudos
Latest Reply
ticusss
New Contributor II
  • 1 kudos

Update — reproduced at minimum complexity, and the failure is unobservableWe stopped theorising and bisected with a deliberately trivial model: SimpleImputer + LogisticRegression, logged with fe.log_model, no custom code, on an online store created t...

  • 1 kudos
1 More Replies
MageshS
by • New Contributor II
  • 939 Views
  • 1 replies
  • 1 kudos

Resolved! Traffic split behavior when traffic_percentage values across served entities sum to more than 100%

Hi all,I'm configuring a Model Serving endpoint with two served entities and ran into some unexpected behavior while testing different traffic_config splits.When I set traffic_percentage to 100 for each of the two served entities (so the total sums t...

  • 939 Views
  • 1 replies
  • 1 kudos
Latest Reply
emma_s
Databricks Employee
  • 1 kudos

Hi,   Just been looking into this for you. All the docs I found suggested it would reject to I tried testing it myself. I tried setting traffic_config with various percentage combinations that don't sum to 100, on both Azure and AWS: Config (A / B) ...

  • 1 kudos
ASH1243434
by • New Contributor II
  • 1662 Views
  • 1 replies
  • 1 kudos

Resolved! Can Databricks Jobs Run on Kubernetes Clusters?

Context: We're exploring using Kubernetes (EKS) as our compute infrastructure instead of Databricks managed clusters. We want to understand if Databricks can orchestrate, deploy, and monitor jobs that run on a Kubernetes cluster.Questions:Is it possi...

  • 1662 Views
  • 1 replies
  • 1 kudos
Latest Reply
szymon_dybczak
Esteemed Contributor III
  • 1 kudos

Hi @ASH1243434 ,Unfortunately, the Databricks cannot natively route job execution into your EKS cluster. There is no "external compute" or "bring your own Kubernetes" option in Databricks Jobs configuration. If my answer was helpful, please consider ...

  • 1 kudos
tkfm_s
by • New Contributor II
  • 1117 Views
  • 2 replies
  • 0 kudos

Memory error in LightGBM training data processing

I am developing a LightGBM model on Databricks, and I am using the Native API because it offers the widest range of options and allows me to try various approaches.The training data is loaded from a table in the Catalog as a Spark DataFrame. However,...

  • 1117 Views
  • 2 replies
  • 0 kudos
Latest Reply
tkfm_s
New Contributor II
  • 0 kudos

Thank you.I will check the document.tkfm_s

  • 0 kudos
1 More Replies
KyraHinnegan
by • New Contributor II
  • 2909 Views
  • 2 replies
  • 1 kudos

Resolved! Which types of model serving endpoints have health metrics available?

I am retrieving a list of model serving endpoints for my workspace via this API: https://docs.databricks.com/api/workspace/servingendpoints/listAnd then going to retrieve health metrics for each one with: https://[DATABRICKS_HOST]/api/2.0/serving-end...

  • 2909 Views
  • 2 replies
  • 1 kudos
Latest Reply
johandoc
New Contributor II
  • 1 kudos

Your observation is correct—this behavior is expected.Endpoints with entity_type = FOUNDATION_MODEL_API do not expose health metrics via the /metrics endpoint, which is why you’re getting 404 responses. These endpoints are fully managed, multi-tenant...

  • 1 kudos
1 More Replies
thomasm
by • New Contributor III
  • 2551 Views
  • 7 replies
  • 3 kudos

MLFlow Detailed Trace view doesn't work in some workspaces

I've created a Databricks Model Serving Endpoint which serves an MLFlow Pyfunc model. The model uses langchain and I'm using mlflow.langchain.autolog().At my company we have some production(-like) workspaces where users cannot e.g. run Notebooks and ...

thomasm_1-1767785859607.png thomasm_0-1767785737567.png thomasm_2-1767785939124.png
  • 2551 Views
  • 7 replies
  • 3 kudos
Latest Reply
lkt1
New Contributor III
  • 3 kudos

Funnily enough, the problem also disappeard on my end this morning Previously, I saw a networking issue in my logs, but that also went away. Let's hope it stays that way! 

  • 3 kudos
6 More Replies
fede_bia
by • Databricks Partner
  • 2650 Views
  • 1 replies
  • 0 kudos

Databricks Model Serving Scaling Logic

Hi everyone,I’m seeking technical clarification on how Databricks Model Serving handles request queuing and autoscaling for CPU-intensive tasks. I am deploying a custom model for text and image extraction from PDFs (using Tesseract), and I’m struggli...

  • 2650 Views
  • 1 replies
  • 0 kudos
Latest Reply
AbhaySingh
Databricks Employee
  • 0 kudos

TLDR: Pre-provision min_provisioned_concurrency â‰¥ your peak parallel requests (in multiples of 4) with scale-to-zero disabled, and chunk large PDFs in your model code to bound per-request latency — reactive autoscaling can't help CPU-bound workloads ...

  • 0 kudos
Dali1
by • New Contributor III
  • 2702 Views
  • 4 replies
  • 1 kudos

Params with databricks Asset bundles

Hello,I am using Databricks Asset bundels to create jobs for machine learning pipelines.My problem is I am using SparkPython taks and defining params inside those. When the job is created it is created with some params. When I want to run the same jo...

  • 2702 Views
  • 4 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

Hi @Dali1, Great questions -- parameterizing ML pipelines in DABs is something a lot of people wrestle with, so let me break down the options. THE SHORT ANSWER No, you should not have to update the job definition every time you want different paramet...

  • 1 kudos
3 More Replies
fede_bia
by • Databricks Partner
  • 1904 Views
  • 1 replies
  • 1 kudos

Resolved! Model Serving Only Shows WARNING/ERROR Logs

Hi everyone,I’m deploying a custom model using mlflow.pyfunc.PythonModel in Databricks Model Serving. Inside my wrapper code, I configured logging as follows:logging.basicConfig( stream=sys.stdout, level=logging.INFO, format='%(asctime)s ...

  • 1904 Views
  • 1 replies
  • 1 kudos
Latest Reply
SteveOstrowski
Databricks Employee
  • 1 kudos

@fede_bia This is worth walking through carefully. this is a common source of confusion when deploying custom models on Databricks Model Serving. SHORT ANSWER The default root logging level for Model Serving endpoints is set to WARNING. That is why y...

  • 1 kudos
Dali1
by • New Contributor III
  • 1721 Views
  • 2 replies
  • 2 kudos

Resolved! Python environment DAB

Hello,I am building a pipeline using DAB.The first step of the dab is to deploy my library as a wheel.The pipeline is run on a shared databricks cluster.When I run the job I see that the job is not using exactly the requirements I specified but it us...

  • 1721 Views
  • 2 replies
  • 2 kudos
Latest Reply
stbjelcevic
Databricks Employee
  • 2 kudos

Hi @Dali1, +1 to @pradeep_singh, on shared clusters, tasks inherit cluster-installed libraries, so you won’t get a clean, versioned environment. Use a job cluster (new_cluster) or switch to serverless jobs with an environment per task for isolation. ...

  • 2 kudos
1 More Replies
Dali1
by • New Contributor III
  • 1266 Views
  • 1 replies
  • 0 kudos

Resolved! Install library in notebook

Hello ,I tried installing a custom library in my databricks notebook that is in a git folder of my worskpace.The installation looks successfulI saw the library in the list of libraries but when I want to import it I have : ModuleNotFoundError: No mod...

  • 1266 Views
  • 1 replies
  • 0 kudos
Latest Reply
Dali1
New Contributor III
  • 0 kudos

Just found the issue - The installation with editable mode doesnt work you have to install it as a library I don't know why 

  • 0 kudos
Labels