Enterprises are collecting data in the order of petabytes, exabytes — even zettabytes.
But data is messy, often disparate and siloed. Many enterprises are hesitant to use it (and thus gain insights from it) in certain environments because it is highly proprietary; in regulated industries like telecommunications, much data can’t even be touched due to its highly sensitive nature.
For these reasons and others — including lack of available data that large scale needed for AI, biases in data or data drift — a growing number of enterprises are turning to synthetic data. This, as its name suggests, isn't real data, but closely resembles it.
“We have to make sure that customer data is completely kept private, that nothing is leaking out,” said Guenter Klas, senior manager for R&D, research clusters, AI and quantum at telecom giant Vodafone, which is beginning to leverage synthetic data.
Also, “we cannot wait for data to accumulate in one of our local markets and start projects,” he told SDxCentral. “We can fill the gap by generating artificial, synthetic data.”
Enhancing, protecting real-world dataSynthetic data reflects real-world data both mathematically and statistically. But rather than being collected from and measured in the real world, it is created by computer simulations, algorithms, simple rules, statistical modeling, simulation and other techniques based on small, anonymized real-world samples.
“While real data is almost always the best source of insights from data, real data is often expensive, imbalanced, unavailable or unusable due to privacy regulations,” Gartner VP analyst Alexander Linden said in a Q&A blog post. “Synthetic data can be an effective supplement or alternative to real data.”
Artificial data can help mitigate weaknesses in real data or can be used when no live data exists, when data is highly sensitive or otherwise biased, or can’t be used, shared or moved. But it doesn’t always have to be trained on real data, however: It can be generated just by looking at domain or institutional knowledge or traces of real data.
With the massive explosion in the use of data-hungry generative AI models and the necessity of privacy and security, enterprises across industry segments are recognizing the potential in synthetic data: Its global market was valued at just $168.9 million in 2021, but is projected to reach $3.5 billion by 2031, representing a CAGR of nearly 36%.
Gartner, for one, even estimates that by 2030, synthetic data will completely overshadow real data in AI models.
Some niche companies in the emerging space include Datagen, GenRocket, Mostly AI, Syntho, Gretel AI, Innodata and Tonic AI (which proudly proclaims itself “the fake data company,” typically a pejorative term).
Overcoming privacy hurdles with synthetic dataVodafone, as a multinational company operating in multiple different jurisdictions with varying rules and regulations, is naturally hampered in its use of data. Due largely to privacy concerns, access to data is often constrained and there are restrictions when it comes to it flowing across geographic borders.
“The hurdles are quite high in terms of bringing data to where it might be processed,” said Klas. “This is where we see synthetic data coming in.”
In this endeavor, Vodafone has teamed with London-based enterprise synthetic data startup Hazy. The company — which announced a $9 million series A seed funding round in March — has been working mainly with large organizations like Vodafone, Accenture, PwC, BMW Group and Wells Fargo because they have the biggest issues when it comes to data, explained Hazy CEO and cofounder Harry Keen.
These massive enterprises have “loads of sensitive data” and also “tons of data silos” fragmented across different geographies, he said.
His company’s tool takes structured datasets and uses machine learning (ML) to scan them to identify trends, patterns, correlations, differences and relations between columns.
“Wherever it lands, you can ask it to generate a realistic data point,” Keen said. “It will do that, and it will keep on doing that, until you say stop.”
The tool can generate much more data than would exist in a source data set, and do so in a safe environment that preserves data characteristics but doesn’t contain sensitive details. “No one wants to be leaking their data,” Keen pointed out.
Most holistic data analysis, accelerating machine learningVodafone, for its part, is looking to do more holistic data analysis — that is, looking at how different campaigns may work in one country compared to another and learning from those datasets, Keen explained.
The “grand plan,” he said, is to create synthetic data assets in each country and aggregate that in a central location to allow for more generalized, larger sets of analysis. For example, churn analysis, “why do customers leave?”
Other areas of interest are load prediction and fraud prediction, as well as detection and prediction of network outages.
One big use case for artificial data is ML: Speeding up internal development processes in creating and improving models and performing quick experiments.
Oftentimes there’s not enough access to data, and while they could work with open-source data, “it’s often not what we need, doesn't fit our circumstances,” said Klas. “We need to create synthetic data that is reflecting the reality in our networks.”
Artificial data helps to improve and accelerate access to data and get projects started faster, thus improving productivity and company agility, he said.
“Data is like the fuel for ML,” said Klas. “If you don't have data, you cannot do much with supervised learning.”
Fostering collaboration, bolstering automationVodafone’s massive ecosystem of suppliers for mobile networks are also innovating with ML — and if they want to train a new ML model, they need Vodafone data.
“It’s not easy for us to hand out Vodafone network data,” Klas said. “Instead we give them synthetic data. It can remove these hurdles.”
Software testing is another big use case. Vodafone is building more software in-house, Keen explained, and that needs to be tested. Artificial data can help determine when failures might occur, how loads on particular network software components are changing over time, how computing resources can be best allocated to software components and how energy consumption might be minimized.
Keen called testing the “bread and butter of every big business” that can take years, the biggest hindrance being access to representative production data.
Furthermore, synthetic data is important for network automation. “We want to automate as much as possible, do things in a predictive way,” said Klas.
Synthetic data considerations beyond telecomSynthetic data doesn’t just have use cases in telecom, of course. Keen pointed out that it is used by companies looking to fine-tune large language models (LLMs) without revealing company-specific data that is “super-super sensitive” to public models such as ChatGPT.
In banking, meanwhile, artificial data has been used as part of a sandbox system to help develop new technologies around fraud detection and money laundering. Meanwhile, BMW used synthetic data to make faster, more accurate decisions on how credit-worthy potential customers were, and Accenture built an application aimed at identifying vulnerable people based on their credit and debit transaction profiles so they could intervene early and prevent bad financial situations.
Similarly, the technology can be used to generate certain areas of a dataset to be more reflective of reality, Keen said. For instance, say a dataset only has 20% women; organizations can generate another 30% to better serve its user base.
Artificial data “increases the level of intensity of innovation that can happen in companies,” said Klas. “We can experiment, innovate real fast.”
Getting buy-in, determining enterprise maturityFrom a cultural standpoint, use of synthetic data can help put privacy officers at ease and dispel the perception that they inhibit innovation or are even enemies of data scientists.
“You can consider synthetic data as truly anonymous,” said Klas.
Still, because it changes the way that data moves around an organization, there must be buy-in from CISOs, CDOs, CEOs, security and legal teams and other C-suite and department leaders, Keen emphasized.
He suggested starting small and building proof points. To help support this, Hazy has created a synthetic data maturity model. Stages of maturity include exploration, evaluation, operationalization, scaling and embedding.
It’s also important to address the backlash that artificial data is “fake” or inaccurate.
“There are some myths that with synthetic, you’re going to lose some of the accuracy,” Keen acknowledged. “Synthetic data is never going to be as 100% accurate as real data.”
By making data private, there is some sacrifice in accuracy, he said, but there’s a lot of usefulness despite that slight dip.
Ultimately, synthetic data is having its moment, he said: Regulators are beginning to explore its possibilities, and as more enterprises embrace it, industry standards will emerge around data use and sharing.
“It’s an interesting time for synthetic data,” Keen said. “It's a complex product that hasn’t been very easy for enterprises to adopt. The next couple of years will be quite a pivot point.”
Comments