StreamSets announced a beta release for its new cloud offering, StreamSets Cloud. Its namesake cloud data integration platform aims to remove the guesswork of open source and hand-coding with a product tailored to address variables that come with ingesting and transforming modern data types.
The software-as-a-service (SaaS) offering is compatible with the likes of Snowflake, Azure Data Warehouse, Azure Data Lake Storage, and Amazon S3. It features built-in Kubernetes containerization and native cloud scaling, which eliminates the need for enterprises to manually construct data flow architectures capable of accommodating big data and streaming applications.
“We are squarely in the data ingest and data management space,” said Kirit Basu, VP of product at StreamSets, in an email to SDxCentral. “StreamSets was designed from the ground up to facilitate a DataOps world.”
Dual Planes In the CloudStreamSets Cloud offers data integration as a fully managed service through its data movement, ingestion, management, and protection platform that consists of two planes: the data plane and the control plane.
“By decoupling these two planes we can provide platform simplicity while fielding a wide variety of design patterns, including integrating data both on-premises and in the cloud,” Basu said.
The data plane is made up of pipelines from the open source StreamSets Data Collector tool, which streams and ingests real-time data movement, and StreamSets Data Transformer, which is a new UI tool to create native Apache Spark applications and pipelines for performing ETL, stream processing, and machine-learning operations.
The control plane hosts StreamSets' Control Hub, Dataflow Performance Manager, and Data Protector. Basu explained the control plane is a dashboard for DataOps while the data plane scales the executors seamlessly to handle the job.
StreamSets the Bar HighStreamSets has three large telco companies in its portfolio including RingCentral and Vodafone, along with mid- and top-tier oil and gas operations like Shell. It also counts customers in health care such as GlaxoSmithKline (GSK), Availity, Humana, and Voya.
Basu explained that GSK can now manage millions of data pipelines, “bridging operational silos such as genomics, clinical data and scientific research, and provide self-service data access and exploratory data science to over 8,700 scientists, analysts, and domain experts without the need for IT involvement,” Basu reported. “We have customers who are reducing sepsis rates, drastically reducing drug trial times, identifying cyberthreats for critical government agencies, and keeping consumer data safe as it transverses systems."
The San Francisco-based company was founded in 2014, and has raised $67.5M through its most recent Series C funding round.
Comments