How to calculate correlation matrix (with all columns at once) in pyspark dataframe?

- - Certifications
- - Learning Paths

- - Community Discussions
- - Get Started Discussions
- - Summit 2024

- - Get Started Resources
- - Events
- - Product Platform Updates
- - Support FAQs
- - Technical Blog
- - Get Started Guides
- - Knowledge Sharing Hub
- - Announcements
- - DatabricksTV

- - Private Groups
- - Skills@Scale

- - Databricks Community Champions
- - Khoros Community Forums Support (Not for Databricks Product Questions)
- - Databricks Community Code of Conduct

Data Engineering

Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.

1 ACCEPTED SOLUTION

Accepted Solutions

got it use -

features = dataset.map(lambda row: row[0:])

from pyspark.mllib.stat import Statistics

corr_mat=Statistics.corr(features, method="pearson")

View solution in original post

1 REPLY 1

got it use -

features = dataset.map(lambda row: row[0:])

from pyspark.mllib.stat import Statistics

corr_mat=Statistics.corr(features, method="pearson")

never-displayed

You must be signed in to add attachments

never-displayed

Announcements

🔔 ALERT: Act Now to Protect Your Community Account; Secure Your Details Before It's Too Late!

Databricks Learning Festival (Virtual): 10 July - 24 July 2024

Data + AI Summit 2024: An Executive Summary for Data Leaders

Big Data Is Back and Is More Important Than AI

Announcing Mosaic AI Agent Framework and Agent Evaluation

Databricks Community

How to calculate correlation matrix (with all columns at once) in pyspark dataframe?

🔔 ALERT: Act Now to Protect Your Community Account; Secure Your Details Before It's Too Late!

Databricks Learning Festival (Virtual): 10 July - 24 July 2024

Data + AI Summit 2024: An Executive Summary for Data Leaders

Big Data Is Back and Is More Important Than AI

Announcing Mosaic AI Agent Framework and Agent Evaluation