Less than two months after IBM formed a high-performance computing (HPC) consortium to study COVID-19, the project has grown into a global effort connecting researchers with some of the most powerful supercomputers in the world.

It's the largest public/private computing partnership ever created, IBM claims. It counts 59 active projects, 40 members, 5 million CPU cores, 50,000 GPUs, and 483 petaFLOPs of cumulative performance.

The so-called COVID-19 HPC Consortium all started with a series of calls, said Dave Turek VP of technical computing at IBM Research, in an interview with SDxCentral. IBM's then CEO, Ginni Rometty, decided that it was imperative that the company do something to support COVID-19 research, Turek said.

"It just took maybe one phone call internally with Dario Gil, director of IBM research, to come to the conclusion that if we can only amass a huge amount of compute resource and make it available to the research community, that would be a really really good, positive step forward," Turek said.

So that's just what IBM set out to do. Within days of discussions with the U.S. Department of Energy, IBM had a website up and running, and by April 1, the first COVID-19 project had been assigned.

However, IBM couldn't do it on their own, and as Turek points out that was never the goal. "We knew that if everybody was willing, you can create a tremendous pool of computational resources to be made available for pursuit of the problem," he said.

The consortium counts some of the most powerful computing clusters in the world as members. One of the latest is the Swiss National Super Computing Center's Piz Daint, which is ranked No. 6 on the Top 500 supercomputing list.

Giving Researchers the Tools They Need

Computational research is nothing new to the scientific community, said Turek.

"A lot of modern science is done computationally, especially at the front end to try to ascertain useful pathways to pursue or to sort through alternatives or model phenomenon that might be too expensive or too dangerous to model in a laboratory," he explained.

Many labs are constrained by limited computation resources, and even if they have 14 promising leads they may only be able to pursue one, Turek said. This is the problem that the consortium set out to solve, by enabling researchers to leverage massive compute resources around the world to test all of their theories at once.

The consortium reports that researchers are studying how COVID-19 spreads between people indoors; investigating what traits make some people more susceptible; searching for new drug compounds with artificial intelligence (AI) or anti-viral compounds in native plants; and, as far fetched as it might sound, the use of atomic force fields to repel COVID-19.

"This computational approach is the best way to maximize the number of possible remedies to COVID-19," Turek said. "Therapies, the antibody tests, the vaccine, it dramatically increases the probability that things in those categories will actually be successful."

This democratization of computing resources also has another benefit, Turek said: No idea has to go untested.

"There is no geographic monopoly on genius. It could pop up anywhere," he said. "If that's the case, then to shortchange that genius by virtue of denying that person access to this kind of compute capability is a really pretty criminal kind of point of view."

Time is of the Essence, and Speed is Everything

However, to make this work the consortium had to move quickly. Turek explained that they couldn't have the kind of elaborate review programs that can take weeks or months to approve a project.

"Whatever we did, we had to move really really fast because people were dying," he said. "The last thing that we want to do is be bureaucratic and slow."

Instead, the idea was simple: Form a research review committee with people knowledgable in the field and give them three days to determine whether or not a proposal has merit. If a proposal passes muster, it moves onto another committee that pairs the project with the appropriate supercomputing resources. That determination, Turek said, is based on a variety of factors, including what resources are available, the resources required, or familiarity with particular architectures.

Once assigned to a compute cluster, a representative from that organization serves as a shepherd aiding the researchers to deploy their project.

"Their role is to help that project get up and running as quickly as possible, and if the principal investigator needs some help they'll find a way to provide him some help," Turek said. "It actually makes everybody's life a whole lot simpler. For example, as a researcher, yes, I do have to subscribe to certain security protocols and other policy issues with respect to the system. But beyond that, I'm just doing my research. I don't have to worry about anything else."

Only the Beginning

The next stage will be to see as many of these projects graduate from the computational phase to lab testing, Turek said.

"What you'll see forthcoming will be the transition of many of these projects from their computational state to their analog state," he said. "We're going to take these ideas into the lab now and flesh them out."

For now, the consortium's efforts are focused squarely on preventing more deaths on account of COVID-19, but the group's aspirations extend well beyond that. According to Turek, what IBM helped to start was a framework by which researchers, regardless of means, can get access to massive computing resources.

"IBM would prefer not to see this as a project, but as the beginning of the sustained effort, not only for COVID-19, but for future biological threats and other kinds of threats," Turek said.