Real-time fraud detection and prevention is just one of several capabilities promised by IBM’s latest mainframe processors. Telum, announced today at Hot Chips 2021, is the company’s first processor to feature integrated acceleration of artificial intelligence (AI) inference workloads.

By analyzing transactions in real time, IBM claims Telum can accelerate fraud detection and prevention, loan processing, clearing and settlement trades, anti-money laundering, and risk analysis workloads.

Fraud detection isn’t a new concept in the financial world. We’ve all received an automated call or text from our bank asking us to confirm whether a charge is legitimate at some point. But, the problem is those alerts usually come long after the transaction has taken place. They are too slow and they introduce privacy concerns, argues Christian Jacobi, distinguished engineer and chief architect for IBM Telum.

Telum features a discrete, on-die AI accelerator capable of six teraFLOPS of raw inferencing performance, which can be scaled up to a 32-chip mainframe capable of 200 teraFLOPS of performance, the company claims.

“The Telum processor's AI capabilities have been specifically designed for low-latency, real-time AI so that it can be embedded in transaction workloads,” Jacobi explained, adding that the technology enables AI models for credit card fraud to be applied before the transaction is completed.

Telum’s AI accelerator isn’t limited to fraud detection either, Jacobi notes. The accelerator can also be used to improve workload efficiency through “intelligent workload placement in the operating system, database query plan optimization, or anomaly detection for security.”

The AI accelerator at the heart of Telum is designed to run models developed using a variety of common AI frameworks through a combination of the Open Neural Network Exchange format (ONNX) and IBM’s own Deep Learning Compiler.

Enterprises “can use the tools that their data scientists are already familiar with… and then the trained models can be exported into the Open Neural Network Exchange format,” he said, adding that this is then fed into IBM’s Deep Learning Compiler, which, as its name suggests, recompiles it for direct execution on the AI accelerator.

To help enterprises get a headstart with these new capabilities, IBM also developed a series of proxy models with some of its largest customers. These pre-trained models can be used for a variety of real-world inferencing applications, including fraud detection, Jacobi said.

Telum Boosts IBM’s Mainframe Power

The AI acceleration capabilities at the heart of Telum only occupy a small part of the overall chip. Most of the die is taken up by eight processor cores with “deep super-scalar out-of-order instruction pipelines,” each of which is capable of running at clock speeds in excess of 5 GHz.

The chip is based on a 7-nanometer process node from Samsung Electronics and features an improved cache architecture, which provides up to 32 megabytes of level-2 (L2) cache per core for a total of 256 megabytes of cache per die.

These cache pools are interconnected by a ring structure, which allows for a virtual level-3 (L3) cache to be generated at lower latencies than would be possible using a physical cache, Jacobi explained. “We achieve a 256 megabyte distributed cache on the chip, with an average latency of only 12 nanoseconds. That is faster than the physical L3 that we had on Z15.”

The improvements to the cache architecture are completely transparent to the software or end user, he added. “From a software perspective and software performance perspective, it still feels like a traditional cache hierarchy, even though everything is built from the L2 caches.”

The company claims these improvements net Telum a 40% performance uplift over IBM's Z15 processor, which it replaces.

However, performance isn’t everything. Telum also features an improved trusted execution environment and support for memory encryption for confidential computing workloads. “The trusted execution environment enables clients to run containerized workloads in a way such that the hardware ensures that the system administrators and the hypervisor administrators cannot get to the data in those containers,” Jacobi explained.

Initial Telum Availability Slated for Early 2022

Telum will be available as a dual-die processor package in IBM’s Z-series and LinuxONE systems.

Up to eight Telum processor dies — four packages — can be installed per system drawer, with a four-drawer system featuring a total of 32 processors interconnected by a high-speed fabric.

IBM said the first Telum-based systems will begin shipping in early 2022.