Back to Blog
Bioinformatics 5 min read

Beyond AlphaFold: The Expanding Role of AI in Biology

Shahroz Rahman

Shahroz Rahman

July 2, 2026

In 2020, DeepMind's AlphaFold solved a fifty-year grand challenge in biology: predicting a protein's three-dimensional structure from its amino acid sequence. This was a landmark achievement, providing static 3D structures for over 200 million proteins.

But the protein folding problem was just the beginning. Biology's most profound questions—how proteins move, bind, mutate, and interact within complex living systems—remain. Today, AI is expanding far beyond structure prediction, transforming how we discover drugs, design gene therapies, and interpret the vast, multi-layered "omics" data that define life itself. We are moving from a static picture to understanding the dynamic, functional language of biology.

From Static Structures to the Language of Life

AlphaFold gave us snapshots of proteins, but proteins are not static; they are dynamic machines. They change shape, bind to other molecules, and carry out the processes that sustain life. Diseases often arise not from a protein's static structure being wrong, but from its interactions or dynamics being perturbed.

This has led to a new wave of AI models. ESMFold, for instance, is not only faster than AlphaFold but approaches the problem differently by treating protein sequences as a language. This language-based approach is becoming a powerful paradigm across biology, extending well beyond proteins to the very code of life itself: DNA.

Reading DNA as a Language

The idea that DNA is a "language" is more than just a metaphor. It has grammatical rules, structural elements like "words" (genes), and hierarchical organization like "paragraphs" (gene clusters). If DNA is a language, the powerful Large Language Models (LLMs) that have revolutionized natural language processing can be adapted to decode and even write it.

One powerful example is GenomeOcean, a generative AI model trained not just on a few well-studied genomes but on a massive dataset of over 600 billion base pairs of DNA from diverse microbial communities in environments like oceans, soils, and the human gut. By learning the "grammar" of microbial genomes, GenomeOcean can do more than analyze existing sequences; it can generate novel DNA sequences, effectively "writing" new genes and pathways that don't exist in nature. This capability could have profound implications for synthetic biology, from designing new enzymes for breaking down plastic to creating microorganisms that can produce sustainable biofuels.

Other models like ATGC-Gen are specifically designed for controllable DNA generation, allowing researchers to design sequences with specific biological properties, such as binding to a particular protein. This is the next frontier in genomic engineering: moving from reading and editing nature's code to writing entirely new chapters.

AI in the Lab: CRISPR Gets a Copilot

The power to edit genes with technologies like CRISPR-Cas9 has revolutionized biology, but designing a successful and safe gene-editing experiment is complex and time-consuming. Enter CRISPR-GPT, an AI "copilot" developed by researchers at Stanford Medicine.

CRISPR-GPT is not a new biological tool but an AI agent powered by an LLM. It was trained on years of published research, online discussions, and experimental data, effectively giving it the collective knowledge of thousands of scientists. A researcher can describe their experimental goals to CRISPR-GPT in plain English, and the AI will suggest guide RNAs, predict potential off-target edits, draft protocols, and even troubleshoot designs.

This is a game-changer for accessibility and efficiency. In one case, a student with minimal experience used CRISPR-GPT to successfully edit genes in cancer cells on their first attempt, a feat that typically requires months of trial and error. By democratizing access to complex gene-editing techniques, tools like CRISPR-GPT could dramatically accelerate the development of new gene therapies, with the goal of developing new drugs in months instead of years.

A New Era for Drug Discovery

The traditional drug discovery pipeline is famously slow and expensive, often taking over a decade and billions of dollars to bring a single drug to market. AI is poised to upend this process from start to finish.

AI's role in drug discovery has expanded far beyond predicting a target protein's shape. It is now being used to analyze massive datasets to identify new disease targets. Once a target is found, generative AI models can design novel molecules specifically to bind to it. Models like DRUG-GAN, for instance, are being used to generate targeted, drug-like compound libraries, significantly outperforming traditional methods.

Companies like Insilico Medicine have already demonstrated this power in practice. By orchestrating AI-driven target discovery, generative chemistry, and automated validation, they have shrunk the timeline from project initiation to a viable preclinical drug candidate to 12-18 months, a process that traditionally takes 3-6 years.

This ambition is now expanding even further. The ultimate goal, outlined by researchers from Insilico and Eli Lilly, is an end-to-end "Prompt-to-Drug" system. In this future, a scientist could simply ask an AI to "design a drug for this disease," and a super-intelligent AI system would autonomously identify targets, design, synthesize, and validate a new drug, transforming pharmaceutical R&D.

Deciphering the Data Deluge with "Omics"

Perhaps the most transformative application of AI is in making sense of complex biological systems. Technologies like mass spectrometry can now measure thousands of molecules—proteins, metabolites, lipids—from a single sample. This "multi-omics" data provides a holistic view of a cell or organism, but it is incredibly complex and difficult to integrate.

AI and deep learning are stepping in to solve this "data deluge" problem. These models can automatically learn patterns from raw experimental data, enabling them to identify the millions of molecules that current bioinformatics tools fail to recognize. More importantly, they can integrate data from different "omics" layers—like genomics, proteomics, and metabolomics—to build a unified, systems-level picture of how a biological process works.

In the future, this AI-driven approach could democratize access to multi-omics, allowing researchers to interrogate complex datasets and extract meaningful biological insights without needing to be an expert in both biology and computational science. This ability to generate and test hypotheses on a "digital twin" of a biological system will lead to a massive acceleration in biodiscovery.

AI in biology has moved decisively beyond its first act. We are now entering an era where AI doesn't just model nature's static forms but participates in its dynamic conversation—reading, writing, and orchestrating the language of life to accelerate the discovery of new drugs, design sophisticated therapies, and ultimately, understand the complex systems that make life possible.

About the Author

Shahroz Rahman

Shahroz Rahman

Bioinformatics and Computational Biology Instructor

A bioinformatics enthusiast. Armed with a higher education degree in bioinformatics, I am passionate about decoding the secrets of life through computational biology. Join me on OmicSkills as I simplify the complexities of bioinformatics, guiding you through genomics, proteomics, and the exciting world where biology meets algorithms. Let's explore the wonders of this field together!

Share Post

Enjoyed this article?

Subscribe to our newsletter to get the latest bioinformatics tutorials and research insights delivered straight to your inbox.