Generative artificial intelligence (AI) — such as ChatGPT and Dalle-2 — is undoubtedly one of the most groundbreaking and discussed technologies in recent history. Its applications and related issues have gone mainstream.
But when looked at from a broader view, it is still a very young technology. And because it has been pushed out so rapidly, it has its limitations — notably when it comes to accuracy, scalability and its recollection capabilities.
All this has led to a growing call for what has been deemed “long-term memory” for AI applications.
As the name would suggest, this allows models to remember things for a long time — rather than just short intervals, as is the current standard.
“While AI models such as GPT from OpenAI are trained on billions of pieces of data, they don’t remember anything you show them or even anything they give back to you,” said Edo Liberty, founder and CEO of Pinecone. “AI models are stateless. They have no memory.”
And, clearly, “memory is useless if you can’t recollect anything,” said Liberty.
Limitations in generative AI context, scopeMany of today’s leading AI systems are recurrent neural networks (RNNs). This deep-learning method uses past information to improve performance on current and future inputs. One of the most applied of these is the Long short-term memory (LSTM) model.
These can process sequential information and retain context from a certain amount of previous inputs to process the next input, explained André Alcalde, co-founder and director of strategic development at CELUS.
“You can think about the usefulness of that when processing words in a sentence,” he said.
For instance, you may have words that are references from previous words, he said, so this model is able to keep the understanding for a limited amount of inputs.
But that’s the key word regarding LSTM models: They are limited.
Simply put, “neural networks are typically forgetful,” said Gartner analyst Erick Brethenoux.
They “remember” information from previous training sessions and take that into consideration for a short while, he said. When new training info comes in, the model will recognize what it has and hasn’t previously seen and use that to adapt and modify.
But if it sees something, then doesn't see it in the next session — or the session after that, or the session after that — the model will forget that information. In a sense, the model is considering: “How many sessions before [this current one] should I keep remembering?”
Alcalde also pointed out that LSTMs — and RNNs in general — may suffer from a couple of issues on practical applications, such as exponential (or vanishing) gradients during training, long training periods, and inefficient computational memory.
“To store a small amount of data in the neural network memory, you may need a large amount of parameters to be trained,” he explained.
Longer-term AI memory modelsLong-term memory models can address these issues by using either attention-augmented LSTMs or new neural network architecture such as transformers with attention mechanisms, said Alcalde. This allows models to better focus on important contextual information and treat it with higher importance
In turn, the long-term memory is then able to retain contextual information for longer sequences and do that in a more memory efficient way, thus saving computational resources, he said. This provides better performance for language models and on translation tasks, since they can better understand the context of the text being processed.
“Just like a human keeps long-term memories, [the model] can be able to capture and retain context for several months, or even years,” said Alcalde.
To advance such capabilities, for instance, one group of researchers recently introduced what they deemed “active long term memory networks.” They described this as “a model of sequential multitask deep learning that is able to maintain previously learned association between sensory input and behavioral output while acquiring new knowledge.”
They write that, “continual learning in artificial neural networks suffers from interference and forgetting when different tasks are learned sequentially.”
The power of vector databases to enhance AI memoryAnother tool being proposed to enhance AI memory is the vector database.
As Liberty of Pinecone explained, AI applications rely on models that understand inputs such as natural language or images. And just like the brain transmits information using chemical or electrical neural signals, AI models encode their understanding of whatever you show them in a numerical format called vector embeddings.
But, he said, traditional relational databases aren’t designed to store and search through vector embeddings; AI models require a specialized database — a vector database — that allows developers to search through “those memories” to find those most relevant, said Liberty.
A vector database stores large numbers of embeddings and associated metadata, such as labels or the original user input. It then uses algorithms to index those embeddings for fast retrieval. Then, when given a query (which is also in the form of an embedding) it quickly returns the most relevant results.
“The beauty of this is that the search is done by meaning, not exact match, and so you don’t need to have anything in the database that exactly matches the query to get useful answers,” said Liberty.
Applying meaning and contextThere are two basic issues that vector databases seek to solve, said Liberty. The first is lack of meaning in search-based applications. For decades, he pointed out, tools have relied on matching keywords from a query to keywords inside a database.
The second issue — which he described as “much more recent but even more painful,” as it prevents organizations from implementing advanced tools such as ChatGPT — is lack of context.
“A solution like ChatGPT can interpret and generate language, but the answers it generates are not always right,” he said. This is a phenomenon known as “hallucination.”
Vector databases can help solve lack of meaning in search-based applications by letting engineers build applications that search over embeddings rather than raw text. This semantic search can yield significantly better results, said Liberty.
“In a search application or recommendation system, that means more satisfied users,” he said. “In a security system, that means better detection of threats.”
Meanwhile, vector databases address the lack of context piece by letting engineers store and discover relevant context and feed it into the AI model along with the original input/question. Then the AI model will generate an answer to the question not only based on its understanding of language, but also based on “hyper-relevant” information.
“AI always gives an answer, but it’s not always right,” he said. “When you combine an AI model with a vector database for long-term memory, it gives the right answer.”
Human-like behavior — not human replacementUltimately, long-term memory can drive AI models towards a learning path that resembles much more human-like behavior, said Alcalde.
For instance, they can learn from their own interactions; remember and apply only important chunks of information on “consumed” books, news, audio and video; and have the ability to refer to this content.
Still, said Brethenoux, the path going forward should be a “hybrid learning capability” with humans in the loop.
Gartner proposes one advanced concept as “composite AI” or “hybrid AI.” This is “the combined application of different AI techniques to improve the efficiency of learning to broaden the level of knowledge representations and, ultimately, to solve a wider range of business problems in a more efficient manner.”
Naturally, humans have long-term memory, Brethenoux pointed out, and our brains use abstraction, perception, and context to make decisions. As he put it, machines should give us something we can interpret and act on in a symbiotic fashion.
“My hope is that we are learning not just to automate things, not just to have things presented to humans, but to find better ways to augment intelligence,” he said.
Comments