Large language models (LLMs) — the backbone of generative AI — are one of the most groundbreaking technologies to come along since the internet itself.

Yet because they are so complex, rapidly evolving and opaque — it can be difficult to know how they were produced or whether, when and how they were modified — they are ripe for security exploits.The recent PoisonGPT campaign demonstrates how easily a threat actor could implant bad information into an LLM supply chain without detection.

[ More AI coverage from SDxCentral ] 

Typical software security supply chains are complicated enough, and AI and LLM supply chains even more so. This calls for a new approach from developers, according to Chainguard, which today announced a new images bundle to help secure the AI software supply chain.

“As AI/ML workloads begin to move past chatbots and into more sensitive workloads, the security of the infrastructure they run on will begin to matter more and more,” Zack Newman, Chainguard principal research scientist, told SDxCentral.

Developers updating with abandon

In a typical software supply chain, software vulnerabilities are tracked in databases, often using common vulnerabilities and exposures (CVEs), Newman said. This model was designed for the “waterfall” software development era of linear sequential phases — typically infrequent releases from large software vendors and relatively small supply chains.

But in AI, there are no longer discrete releases each time there is a change, he said. Gigabytes or terabytes of training data are “entering the conversation,” and the expensive training step may use hundreds of thousands of dollars worth of computing resources.

“Software dependencies often need to be on the cutting edge for the best results, so developers update with abandon,” said Newman.

Therefore, identifying affected versions and tracking the provenance of AI systems as developers are continuously training models is a very complicated security domain from a software supply chain security perspective, he said.

Building a strong, secure base

Ultimately, it comes down to building on secure base images, which are the starting point for container-based development workflows. As their name suggests, they are layered over with the necessary libraries, binaries and configuration files necessary to run an application.

Clean base container images are “ground zero” for software supply chain security in modern software development, Newman said. As a developer, it is the discipline of establishing the provenance of artifacts you’re building with and knowing they are free from known vulnerabilities or unnecessary packages that could “accumulate vulnerabilities over time.”

Ultimately, base images have their own artifacts because supply chains are recursive.

“So, when you are using a base image, you are trusting not only the people you downloaded the software from, but also the security of the dependencies of the artifacts that software uses,” he said.

Using an out-of-date or vulnerability-laden base image means you’re starting out with “security debt  — “before you’ve even built your application or model, you’re at risk.”

Constantly updated base images key

More than 53% of data scientists plan to get large language applications into deployment “as soon as possible,” and they have great interest in building with popular AI languages including TensorFlow, PyTorch and Kubeflow, Newman pointed out.

But while great for data science use cases, these languages can pose challenges when deployed into “security critical” environments due to their large size and attack surface, Newman said. There are also package management issues.

He argued that developers creating LLMs and adopting new AI frameworks should be using base images that are always updated with the latest remediations to known vulnerabilities.

This can help protect them from a range of AI-specific vulnerabilities — copying vulnerable code from LLMs, deepfakes, backdoored training data or models, privacy issues where models spit out personal information from training data and “adversarial examples” that cause unexpected model behaviors.

AI Chainguard images

Chainguard aims to offer this capability with its new AI Chainguard Images, a collection of images for stages in the AI workload lifecycle — including development images, workflow management tools and vector databases for production storage.

Chainguard Images includes the following:

  • Python, Conda, OpenAI and Jupyter notebook images for developing models and using the OpenAI API
  • Kubeflow images for deploying production ML pipelines to Kubernetes-based platforms
  • Milvus and Weaviate vector database images for data storage

The images are hardened by default, a fraction of the size, and aim to meet the company’s standard zero-known CVE SLA through daily updates and patching.

The images have two advantages, said Newman: They’re “batteries-included” meaning ML developers can focus on the data and their analyses and models rather than operations.

Second, they’re secure by default, he said. Risks from dependencies go away so users can focus on the risks that they can control.

Still scratching the surface

Overall, developers are adopting open-source tooling like Sigstore and following best practices for locking down the provenance of software supply chains, Newman said.

There is a general industry-wide movement toward locking down backdoors to software and its dependencies, preventing typosquatting or malicious package inserts, build system compromises and compromises of the software distribution process itself.

So-called “MLBOMs” or “AIBOMs” (bills of materials for ML or AI) also allow “farm-to-table” tracing for the lifecycle of a particular model.

But even as developers are following these practices and organizations and government agencies alike are exploring and creating frameworks around risk, there is still much more work to be done.

“The industry has still barely scratched the surface to go beyond theoretical understandings of potential AI-specific security vulnerabilities,” he said.