Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-27-2024 12:34 PM
Seems like this thread has died, but for posterity, databricks provides the following code for installing poppler on a cluster. The code is sourced from the dbdemos accelerators, specifically the "LLM Chatbot With Retrieval Augmented Generation (RAG) and Llama 2 70B" (https://notebooks.databricks.com/demos/llm-rag-chatbot/index.html#) demo. In the 01-PDF-Advanced-Data-Preparation notebook there's code to remote execute the 00-init-advanced notebook and in that notebook, you'll find the code below.
#install poppler on the cluster (should be done by init scripts)
def install_ocr_on_nodes():
"""
install poppler on the cluster (should be done by init scripts)
"""
# from pyspark.sql import SparkSession
import subprocess
num_workers = max(1,int(spark.conf.get("spark.databricks.clusterUsageTags.clusterWorkers")))
command = "sudo rm -rf /var/cache/apt/archives/* /var/lib/apt/lists/* && sudo apt-get clean && sudo apt-get update && sudo apt-get install poppler-utils tesseract-ocr -y"
def run_subprocess(command):
try:
output = subprocess.check_output(command, stderr=subprocess.STDOUT, shell=True)
return output.decode()
except subprocess.CalledProcessError as e:
raise Exception("An error occurred installing OCR libs:"+ e.output.decode())
#install on the driver
run_subprocess(command)
def run_command(iterator):
for x in iterator:
yield run_subprocess(command)
# spark = SparkSession.builder.getOrCreate()
data = spark.sparkContext.parallelize(range(num_workers), num_workers)
# Use mapPartitions to run command in each partition (worker)
output = data.mapPartitions(run_command)
try:
output.collect();
print("OCR libraries installed")
except Exception as e:
print(f"Couldn't install on all node: {e}")
raise e