Kara Annanie and Stian Ulriksen is presenting at SQLSaturday in Denver on the topic of Machine Learning in Databricks. Come check it out!
On September 17th, I will be presenting an introduction to Azure Databricks for the Boulder SQL User Group at Datavail Corporation in Broomfield. Not sure what Databricks is, or if it is for you? Check out this short post about some of the key capabilities. Come check it out if you are in the area! https://www.meetup.com/BoulderSql/events/261149223/ Hope to see you there!
Pandas in Python is an awesome library to help you wrangle with your data, but it can only get you so far. When you start moving into the Big Data space, PySpark is much more effective in accomplishing what you want. This post aims at helping you migrate what you know about Pandas to PySpark. If you are new to Spark, checkout this post about Databricks, and go spin up a cluster to play around. Apache Spark and PySpark Before we get going, let’s take a step back and talk about Apache Spark. Spark is a fast and general engine for large-scale data processing. Spark uses distributed computing to accomplish higher speeds on large datasets. When you submit a request to Spark, the driver node distributes the workload to a number of worker nodes who processes parts of the request in parallel. Think of it as an improvement to original…
Databricks launched a new open source product at the Spark AI Summit 2019 called Delta Lake. Delta Lake touts that it brings ACID transactions to Apache Spark and big data workloads.
I recently had to connect my Azure Databricks instance to our Azure Data Lake Storage (Generation 1) and was running into some problems getting everything set up. I am sure I am not the only one out there having these problems so if you do as well, here is a little guide to get you connected.
Join Kara Annanie and I at the Denver SQL User Group meeting at Microsoft’s offices in Denver on Thursday. We will be presenting on Azure Databricks and how you can get started using it. https://www.meetup.com/Denver-SQL-Server-User-Group/events/jnhjcqyzgbxb/
In this blog post I will lay out five (5) reasons you should consider Databricks before starting your next data science project.