Data Engineering

Forum Posts

Sorted by:

by User16826991422 • Contributor

12-02-2015 10:26:01 AM

11002 Views
12 replies
0 kudos

Resolved! How do I create a single CSV file from multiple partitions in Databricks / Spark?

Using sparkcsv to write data to dbfs, which I plan to move to my laptop via standard s3 copy commands. The default for spark csv is to write output into partitions. I can force it to a single partition, but would really like to know if there is a ge...

Data Engineering

11002 Views
12 replies
0 kudos

12-02-2015 10:26:01 AM

View Replies

Latest Reply

ChristianHomber
New Contributor II

01-21-2020 3:50:40 AM

0 kudos

Without access to bash it would be highly appreciated if an option within databricks (e.g. via dbfsutils) existed.

0 kudos

01-21-2020 3:50:40 AM

11 More Replies

by tunguyen90 • New Contributor

11-08-2019 7:41:42 AM

6681 Views
3 replies
1 kudos

How to change line separator for csv file exported from dataframe in databricks

Hello, Currently, I'm facing problem with line separator inside csv file, which is exported from data frame in Azure Databricks (version Spark 2.4.3) to Azure Blob storage. All those csv files contains LF as line-separator. I need to have CRLF (\r\n...

Data Engineering

6681 Views
3 replies
1 kudos

11-08-2019 7:41:42 AM

View Replies

Latest Reply

Nikhila
New Contributor II

12-24-2020 5:02:26 AM

1 kudos

Hi, Have you got the solution for above problem.Kindly let me know.

1 kudos

12-24-2020 5:02:26 AM

2 More Replies

by tismith1_558848 • New Contributor

05-23-2018 1:27:21 PM

6035 Views
2 replies
0 kudos

Resolved! Change size or aspect ratio of ggplot visualizations

I understand that plots in R notebooks are captured by a png graphics device. Is there a way to set the size or the aspect ratio of the canvas? I understand that I can resize the rendered .png by dragging the handle in the notebook, but that means I...

Data Engineering

6035 Views
2 replies
0 kudos

05-23-2018 1:27:21 PM

View Replies

Latest Reply

kassandra
New Contributor II

12-23-2020 1:49:15 PM

0 kudos

Hi @sdaza , the answer above didn't change the size somehow, or perhaps I was putting it in the wrong place? I entered it in a new cell before the %sql cell with the plot chart.

0 kudos

12-23-2020 1:49:15 PM

1 More Replies

by bhosskie • New Contributor

05-13-2016 1:33:41 PM

8136 Views
9 replies
0 kudos

How to merge two data frames column-wise in Apache Spark

I have the following two data frames which have just one column each and have exact same number of rows. How do I merge them so that I get a new data frame which has the two columns and all rows from both the data frames. For example, df1: +-----+...

Data Engineering

8136 Views
9 replies
0 kudos

05-13-2016 1:33:41 PM

View Replies

Latest Reply

AmolZinjade
New Contributor II

12-16-2020 9:36:04 AM

0 kudos

@bhosskie from pyspark.sql import SparkSession spark = SparkSession.builder.appName("Spark SQL basic example").enableHiveSupport().getOrCreate() sc = spark.sparkContext sqlDF1 = spark.sql("select count(*) as Total FROM user_summary") sqlDF2 = sp...

0 kudos

12-16-2020 9:36:04 AM

8 More Replies

by cfregly • Contributor

02-24-2015 3:51:39 PM

23701 Views
9 replies
0 kudos

Resolved! How do I avoid the "No space left on device" error where my disk is running out of space?

Data Engineering

23701 Views
9 replies
0 kudos

02-24-2015 3:51:39 PM

View Replies

Latest Reply

MichaelHuntsber
New Contributor II

07-17-2019 4:05:51 AM

0 kudos

I have 8 GB of internal memory, but several MB of them are free but I also have an additional memory with an 8 GB memory card. Anyway, there is no enough space and the memory card is completely empty.essay service

0 kudos

07-17-2019 4:05:51 AM

8 More Replies

by nmud19 • New Contributor II

09-08-2016 4:53:14 AM

56334 Views
8 replies
6 kudos

how to delete a folder in databricks mnt?

I have a folder at location dbfs:/mnt/temp I need to delete this folder. I tried using %fs rm mnt/temp & dbutils.fs.rm("mnt/temp") Could you please help me out with what I am doing wrong?

Data Engineering

56334 Views
8 replies
6 kudos

09-08-2016 4:53:14 AM

View Replies

Latest Reply

amitca71
Contributor II

11-30-2020 12:16:10 AM

6 kudos

use this (last raw should not be indented twice...): def delete_mounted_dir(dirname): files=dbutils.fs.ls(dirname) for f in files: if f.isDir(): delete_mounted_dir(f.path) dbutils.fs.rm(f.path, recurse=True)

6 kudos

11-30-2020 12:16:10 AM

7 More Replies

by vamsivarun007 • New Contributor II

03-13-2020 1:53:24 AM

25160 Views
2 replies
2 kudos

Driver is up but is not responsive, likely due to GC.

Hi all, "Driver is up but is not responsive, likely due to GC." This is the message in cluster event logs. Can anyone help me with this. What does GC means? Garbage collection? Can we control it externally?

Data Engineering

25160 Views
2 replies
2 kudos

03-13-2020 1:53:24 AM

View Replies

Latest Reply

Carlos_AlbertoG
New Contributor II

11-25-2020 11:57:30 PM

2 kudos

spark.catalog.clearCache() solve the problem for me

2 kudos

11-25-2020 11:57:30 PM

1 More Replies

by PraveenKumarB • New Contributor

04-24-2019 7:08:28 AM

5729 Views
5 replies
0 kudos

java.io.IOException: No FileSystem for scheme: null

Getting the error when try to load the uploaded file in py notebook.# File location and type file_location = "//FileStore/tables/data/d1.csv" file_type = "csv" # CSV options infer_schema = "true" first_row_is_header = "false" delimiter = ","# The app...

Data Engineering

5729 Views
5 replies
0 kudos

04-24-2019 7:08:28 AM

View Replies

Latest Reply

DivyanshuBhatia
New Contributor II

11-22-2020 6:29:46 AM

0 kudos

@naughtonelad if your issue is solved,please let me know as I am facing the same problem

0 kudos

11-22-2020 6:29:46 AM

4 More Replies

by nthomas • New Contributor

05-26-2016 11:27:38 AM

4972 Views
5 replies
0 kudos

Tips for properly using large broadcast variables?

I'm using a broadcast variable about 100 MB pickled in size, which I'm approximating with: >>> data = list(range(int(10*1e6))) >>> import cPickle as pickle >>> len(pickle.dumps(data)) 98888896Running on a cluster with 3 c3.2xlarge executors, ...

Data Engineering

4972 Views
5 replies
0 kudos

05-26-2016 11:27:38 AM

View Replies

Latest Reply

dragoncity
New Contributor II

11-06-2020 8:18:49 PM

0 kudos

The Facebook credit can be utilized by the gamers to purchase the pearls. The other route is to finished various sorts of Dragons in the Dragon Book. Dragon City Gems There are various kinds of Dragons, one is amazing, at that point you have the fund...

0 kudos

11-06-2020 8:18:49 PM

4 More Replies

by JulioManuelNava • New Contributor

11-02-2019 12:40:15 AM

5340 Views
2 replies
0 kudos

[pyspark] foreach + print produces no output

The following code produces no output. It seems as if the print(x) is not being executed for each "words" element: words = sc.parallelize ( ["scala", "java", "hadoop", "spark", "akka", "spark vs hadoop", "pyspark", "pysp...

Data Engineering

5340 Views
2 replies
0 kudos

11-02-2019 12:40:15 AM

View Replies

Latest Reply

john_nicholas
New Contributor II

11-03-2020 3:41:49 AM

0 kudos

Epson wf-3640 error code 0x97 is the common printer error code that may occur mostly in all printers but in order to resolve the error code, upon provides the best printer guide to all printer users.

0 kudos

11-03-2020 3:41:49 AM

1 More Replies

by dchokkadi1_5588 • New Contributor II

05-10-2016 3:36:19 PM

12129 Views
8 replies
0 kudos

Resolved! graceful dbutils mount/unmount

Is there a way to indicate to dbutils.fs.mount to not throw an error if the mount is already mounted? And viceversa, for unmount to not throw an error if it is already unmounted? I am trying to run my notebook as a job and it has a init section that...

Data Engineering

12129 Views
8 replies
0 kudos

05-10-2016 3:36:19 PM

View Replies

Latest Reply

Mariano_IrvinLo
New Contributor II

10-31-2020 7:59:00 PM

0 kudos

If you use scala to mount a gen 2 data lake you could try something like this /Gather relevant Keys/ var ServicePrincipalID = "" var ServicePrincipalKey = "" var DirectoryID = "" /Create configurations for our connection/ var configs = Map (...

0 kudos

10-31-2020 7:59:00 PM

7 More Replies

by Barb • New Contributor III

10-07-2019 9:05:30 AM

4204 Views
6 replies
0 kudos

SQL charindex function?

Hi all,I need to use the SQL charindex function, but I'm getting a databricks error that this doesn't exist. That can't be true, right? Thanks for any ideas about how to make this work!Barb

Data Engineering

4204 Views
6 replies
0 kudos

10-07-2019 9:05:30 AM

View Replies

Latest Reply

Traveller
New Contributor II

10-13-2020 10:14:32 PM

0 kudos

The best option I found to replace CHARINDEX was LOCATE, examples from the Spark documentation below > SELECT locate('bar', 'foobarbar', 5); 7 > SELECT POSITION('bar' IN 'foobarbar'); 4

0 kudos

10-13-2020 10:14:32 PM

5 More Replies

by SatheeshSathees • New Contributor

08-19-2020 11:31:33 AM

5411 Views
1 replies
0 kudos

how to dynamically explode array type column in pyspark or scala

HI, i have a parquet file with complex column types with nested structs and arrays. I am using the scrpit from below link to flatten my parquet file. https://docs.microsoft.com/en-us/azure/synapse-analytics/how-to-analyze-complex-schema I am able ...

Data Engineering

5411 Views
1 replies
0 kudos

08-19-2020 11:31:33 AM

View Replies

Latest Reply

shyam_9
Valued Contributor

09-18-2020 12:39:35 PM

0 kudos

Hello, Please check out the below docs and notebook which has similar examples, https://docs.microsoft.com/en-us/azure/synapse-analytics/how-to-analyze-complex-schemahttps://docs.microsoft.com/en-us/azure/databricks/_static/notebooks/transform-comple...

0 kudos

09-18-2020 12:39:35 PM

by zachary_jones • New Contributor

10-28-2019 7:49:59 AM

2729 Views
3 replies
0 kudos

Resolved! Python logging: 'Operation not supported' after upgrading to DBRT 6.1

My organization has an S3 bucket mounted to the databricks filesystem under /dbfs/mnt. When using Databricks runtime 5.5 and below, the following logging code works correctly:log_file = '/dbfs/mnt/path/to/my/bucket/test.log' logger = logging.getLogg...

Data Engineering

2729 Views
3 replies
0 kudos

10-28-2019 7:49:59 AM

View Replies

Latest Reply

lycenok
New Contributor II

09-09-2020 10:06:20 PM

0 kudos

Probably it's worth to try to rewrite the emit ... https://docs.python.org/3/library/logging.html#handlers This works for me: class OurFileHandler(logging.FileHandler): def emit(self, record): # copied from https://github.com/python/cpython/bl...

0 kudos

09-09-2020 10:06:20 PM

2 More Replies

by DimitrisMpizos • New Contributor

02-08-2016 7:45:52 AM

21204 Views
16 replies
0 kudos

Exporting data from databricks

I couldn't find in documentation a way to export an RDD as a text file to a local folder by using python. Is it possible?

Data Engineering

21204 Views
16 replies
0 kudos

02-08-2016 7:45:52 AM

View Replies

Latest Reply

Manu1
New Contributor II

03-25-2019 8:18:04 AM

0 kudos

To: Export a file to local desktop Workaround : Basically you have to do a "Create a table in notebook" with DBFS The steps are: Click on "Data" icon > Click "Add Data" button > Click "DBFS" button > Click "FileStore" folder icon in 1st pane "Sele...

0 kudos

03-25-2019 8:18:04 AM

15 More Replies

User

Count

1602

736

343

284

247

Databricks

Forum Posts

Resolved! How do I create a single CSV file from multiple partitions in Databricks / Spark?

How to change line separator for csv file exported from dataframe in databricks

Resolved! Change size or aspect ratio of ggplot visualizations

How to merge two data frames column-wise in Apache Spark

Resolved! How do I avoid the "No space left on device" error where my disk is running out of space?

how to delete a folder in databricks mnt?

Driver is up but is not responsive, likely due to GC.

java.io.IOException: No FileSystem for scheme: null

Tips for properly using large broadcast variables?

[pyspark] foreach + print produces no output

Resolved! graceful dbutils mount/unmount

SQL charindex function?

how to dynamically explode array type column in pyspark or scala

Resolved! Python logging: 'Operation not supported' after upgrading to DBRT 6.1

Exporting data from databricks

Best way to parse Google Analytics data in Databri...

DELTA_EXCEED_CHAR_VARCHAR_LIMIT

Not able to set run_as service_principal_name

Pyspark operations slowness in CLuster 14.3LTS as ...

[Databricks Assets Bundles] Workflow trigger on fi...