OpenAI co-Founder and Chief Scientist Ilya Sutskever joined NVIDIA Founder and CEO Jensen Huang in a fireside chat at GTC to discuss the two-level training of the artificial intelligence (AI) startup’s generative pretrained transformer (GPT) model.

The first stage to train a large neural network is to make it accurately predict the next word in a series. “What we are doing is that we are learning a world model,” Sutskever said.

He explained that on the surface, it looks like the model is just learning statistical correlations in text to compress them, but it turns out that “what the neural network learns is some representation of the process that produced the text. This text is actually a projection of the world.”

“So what the neural network is learning is more and more aspects of the world of people, of the human conditions — their hopes and dreams and motivations, their interactions and the situations that we are in, and the neural network learn a compressed, abstract, usable representation of that,” Sutskever said. “This is what's being learned from accurately predicting the next word.”

"Furthermore, the more accurate you are at predicting the next word, the higher with fidelity, the more resolution you get in this process,” he added.

In this pretraining stage, what the company does not do is specify the desired behavior that it wishes its neural network to exhibit. Here comes the second stage, in which AI teachers communicate to the neural network about those behaviors and set up “guardrails” to make them more reliable and precise.

“We will follow certain guidance rules and not violate them that require additional training. This is where the fine-tuning and the reinforcement learning from human teachers and other forms of AI assistance, It is not just reinforcement learning from human teachers, [it’s] social reinforcement learning from human and AI collaboration,” Sutskever stated.

Those AI teachers give AI intended instructions and communication with the neural network to teach AI how to behave, including the bounding box, but they do not teach AI any new knowledge, he noted.

Sutskever emphasized the importance of the second stage. “The better we do the second stage, the more useful, the more reliable this neural network will be.”

Nvidia GPUs Powers OpenAI’s Language Model

In 2022, Sutskever demonstrated the AlexNet model with other AI pioneers to show the power of deep neural networks trained on massive datasets in an academic contest, Huang noted.

Some of Sutskever’s early work involved running their model on several Nvidia GeForce GTX 5080 graphics processing units (GPUs) in a University of Toronto lab.

“The ImageNet dataset and a convolutional neural network were a great fit for GPUs that made it unbelievably fast to train something unprecedented,” Sutskever said. “It was also very clear that the convolutional neural network is such a great fit for the GPU. So it should be possible to make it go unbelievably fast, and therefore train something which would be completely unprecedented in terms of its size.”

Nvidia touted now, Microsoft Azure cloud service uses tens of thousands of the latest NVIDIA A100 and H100 Tensor Core GPUs for training and inference on OpenAI's models like ChatGPT.

“In the 10 years we’ve known each other, the models you’ve trained [have grown by] about a million times,” Huang said. “No one in computer science would have believed the computation done in that time would be a million times larger.”

Looking ahead, Sutskever noted the current frontiers are centered around reliability.

It’s “around the system can be trusted, really get into a point where you can trust what it produces, really getting to a point where if it doesn't understand something that's for clarification, says that it doesn't know something, says that it needs more information. I think those are perhaps the biggest areas where improvement will lead to the biggest impact on the usefulness of real systems.,” he said.