Artificial intelligence translates protein sequences into natural language descriptions

The BetaDescribe system, developed by researchers from the Technion and Tel Aviv University, takes a sequence of amino acids and produces a verbal description of the protein's function, chemical activity, and possible binding sites. The researchers hope the system will help choose which hypotheses to test in the lab and accelerate medical and biotechnological research.

A protein sequence may contain hundreds or even thousands of amino acids, but reading the sequence alone does not usually reveal what the protein does inside the cell. Researchers frompolytechnic AndTel Aviv University developed a system artificial intelligence which tries to bridge this gap: it accepts a protein sequence and produces a natural language description of its possible functions and characteristics.

The system, called BetaDescribe, presented in an article published in the journal Proceedings of the National Academy of Sciences - PNASIt was developed under the leadership of doctoral student Ido Dotan, under the joint supervision of Prof. Jonathan Blinkov from the Taub Faculty of Computer Science at the Technion and Prof. Tal Popko from the Faculty of Life Sciences at Tel Aviv University.

Prof. Eran Bachrach, Prof. Marcelo Erlich, and doctoral student Iris Lubman from the Faculty of Life Sciences at Tel Aviv University also participated in the study.

From a string of letters to a biological explanation

Proteins Are made up of chains of amino acids. The precise sequence of amino acids affects the three-dimensional structure of the protein, and the structure largely determines the actions it is able to perform. Proteins may, among other things, catalyze chemical reactions, transmit signals between cells, transport substances, recognize disease agents, or participate in tissue construction.

Modern sequencing technologies are discovering protein sequences much faster than the ability to test every protein in an experiment. Biological databases already contain billions of predicted or documented sequences, but only a relatively small fraction of proteins have been directly characterized in the laboratory.

Therefore, much of the information about new proteins is obtained by comparing them to known proteins. When a new sequence is very similar to a sequence with a known function, it can be assumed that their functions are also similar. The method is effective in many cases, but it becomes difficult when it comes to a protein that has no known relative in the databases.

BetaDescribe is designed to handle such cases as well. Rather than simply searching for a match to a known sequence, it uses a generative model and validation and evaluation mechanisms to formulate a detailed description of the protein.

According to the researchers, the description may include the protein's presumed biological function, the type of catalytic activity it performs, its involvement in metabolic pathways, and possible binding sites for other molecules.

Not just a short label

Traditional protein databases tend to attach labels, technical terms, or codes describing biological functions to sequences. In contrast, BetaDescribe is designed to produce more continuous and detailed text.

The difference is significant because a single protein may serve several functions, act only in certain tissues, or participate in a complex biological process. A verbal description may allow researchers to better understand why the system proposed a particular function and to design experiments that would test the hypothesis.

However, a description produced by an artificial intelligence system is no substitute for experimentation. Generative models may produce information that sounds convincing but is not true. Therefore, the researchers have integrated mechanisms into the system designed to test the descriptions, compare options, and reduce the risk of unfounded assertions.

The system does not prove the protein's function but suggests a testable hypothesis. Its potential importance lies in filtering through a vast number of sequences and directing researchers to the most promising proteins and experiments.

Test on uncharacterized proteins

The researchers tested the system on six previously uncharacterized proteins. According to the Technion announcement, the system was able to suggest functional descriptions for them that could be tested through further research.

Such examples are designed to test whether the model can go beyond memorization of information about known proteins and generate hypotheses about novel sequences. However, broader testing and experimental validation of the predictions will be needed to determine the accuracy and utility of the system across a range of protein families.

Future applications may include detection enzymes for industrial processes, searching for disease-related proteins, identifying drug target sites, and developing new proteins for biotechnology and agriculture.

The study also illustrates a broader shift in computational biology. AI systems are no longer used just to sort data or predict structures, but also to generate explanations and hypotheses in language that humans can read. The challenge will be to ensure that the current language does not hide uncertainty, and that any important predictions are ultimately tested in the lab.

The research was supported by the National Science Foundation.

The scientific article: BetaDescribe: Generating natural-language descriptions of protein sequences PNAS DOI: 10.1073 / pnas.2537345123

Questions and Answers

What did the researchers develop? An artificial intelligence system called BetaDescribe, which receives a protein sequence and produces a verbal description of its possible functions and characteristics.

Why is such a system needed? Databases include billions of protein sequences, but only a small fraction of them have been directly characterized in laboratory experiments.

How is it different from a regular search in protein databases? It does not rely solely on finding sequences similar to known proteins, but attempts to infer functions and present them as detailed text.

What information can she offer? Putative biological role, catalytic activity, involvement in metabolic pathways and possible binding sites.

Does the description prove what the protein does? No. This is a computational prediction intended to provide hypotheses for research. Full validation still requires biological experiments.

How was the system tested? The researchers demonstrated its action on six previously uncharacterized proteins.

What are the possible applications? Medical research, drug target discovery, enzyme detection, biotechnology and agriculture.

For the scientific article: Opening the scientific article

More on the subject on the science website

2 תגובות

Leave a Reply

Email will not be published. Required fields are marked *

This site uses Akismet to filter spam comments. More details about how the information from your response will be processed.