<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Best practices to use databrics to work with sql server data on premises in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/best-practices-to-use-databrics-to-work-with-sql-server-data-on/m-p/81744#M36402</link>
    <description>&lt;P&gt;I want to use Databricks to analyze data that is stored on an on-premises SQL Server. What is the best way to bring this data into Databricks?&lt;/P&gt;&lt;P&gt;I tried to configure Lakehouse Federation so I could query on-premises data directly from Databricks. However, due to the different networks, I cannot connect. I assume I’d have to configure an Azure Virtual Network or something similar.&lt;/P&gt;&lt;P&gt;The same issue arises when using the JDBC connector. I cannot simply connect and create a DataFrame from my data because it is on a completely different network.&lt;/P&gt;&lt;P&gt;Let's say I have a simple query that joins 10 tables and outputs just 200 rows. Do I need to bring all those 10 tables to Databricks? Or should I move them to Data Lake Storage first and then to Databricks?&lt;BR /&gt;&lt;BR /&gt;Thanks&lt;/P&gt;</description>
    <pubDate>Sat, 03 Aug 2024 15:13:06 GMT</pubDate>
    <dc:creator>serdia</dc:creator>
    <dc:date>2024-08-03T15:13:06Z</dc:date>
    <item>
      <title>Best practices to use databrics to work with sql server data on premises</title>
      <link>https://community.databricks.com/t5/data-engineering/best-practices-to-use-databrics-to-work-with-sql-server-data-on/m-p/81744#M36402</link>
      <description>&lt;P&gt;I want to use Databricks to analyze data that is stored on an on-premises SQL Server. What is the best way to bring this data into Databricks?&lt;/P&gt;&lt;P&gt;I tried to configure Lakehouse Federation so I could query on-premises data directly from Databricks. However, due to the different networks, I cannot connect. I assume I’d have to configure an Azure Virtual Network or something similar.&lt;/P&gt;&lt;P&gt;The same issue arises when using the JDBC connector. I cannot simply connect and create a DataFrame from my data because it is on a completely different network.&lt;/P&gt;&lt;P&gt;Let's say I have a simple query that joins 10 tables and outputs just 200 rows. Do I need to bring all those 10 tables to Databricks? Or should I move them to Data Lake Storage first and then to Databricks?&lt;BR /&gt;&lt;BR /&gt;Thanks&lt;/P&gt;</description>
      <pubDate>Sat, 03 Aug 2024 15:13:06 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/best-practices-to-use-databrics-to-work-with-sql-server-data-on/m-p/81744#M36402</guid>
      <dc:creator>serdia</dc:creator>
      <dc:date>2024-08-03T15:13:06Z</dc:date>
    </item>
  </channel>
</rss>

