StreamSets and Databricks left port together today announcing a partnership to harness capabilities from StreamSets DataOps platform and Databricks’ Delta Lake, thus setting course to voyage into data lakes, accelerate analytics projects, and reduce cycle times. 

The partnership announcement comes less than a week after Databricks’s open-source software project Delta Lake raised the Linux Foundation’s open governance model flag, and StreamSets announced a beta release for its cloud data integration platform.

StreamSets’ platform aims to remove the guesswork of open source and hand-coding with a product tailored to address variables that come with ingesting and transforming modern data types.

Open-source tools are key to delivering data lake functionality. But to achieve the best return on investment for big data systems, companies need to marry commercial tools with open-source components, said Sean Anderson, director of product marketing at StreamSets. And that is why the new partnership makes sense, he explained.

"It is almost impossible to execute a data hub or lake without some components of the open-ecosystem involved," Anderson wrote in an email to SDxCentral. "We feel that with the combination of Apache Spark, the StreamSets Platform, and Delta Lake, we have designed an optimal Spark platform that meets the requirements of companies that practice DataOps and provide value for those that do not yet today."

Navigation Aid

StreamSets’ platform presides as an amicable co-captain in this endeavor alongside Databricks’ Delta Lake and Apache Spark. It enables users to develop data pipelines to deploy Apache Spark without the coding requirements associated with traditional dev/test/production systems. 

The integrated service aims to provide insight beyond error diagnoses so users can understand the health of their Apache Spark projects at a granular level when errors occur.

DataOps capabilities native to the StreamSets platform will bring enhanced visibility, control, error-handling, and troubleshooting to Databricks, Anderson claims. “As long as users understand the logical mechanisms of what they want to build, the StreamSets Platform can help execute that in a native Spark application.” 

Because companies are underinvesting in data applications due to a shortage of trained Apache Spark developers, stripped down coding requirements with the provided tools for operation in production can help companies “reduce the time that it takes to bring Customer 360, IoT, machine learning, data science, and streaming use cases to market,” he explained.