Exploring Frontiers With Computational Power

Bridging the Gap in Protein Science: How RIS Empowers the Holehouse Lab

3D molecular structure model of a protein, specifically visualized using AlphaFold's confidence coloring scheme (pLDDT score)

Dr. Alex Holehouse’s lab is decoding one of biology’s most persistent blind spots: intrinsically disordered regions (IDRs) of proteins. Supported by the high-performance computing and storage resources of WashU IT’s Research Infrastructure Services (RIS), the lab recently published groundbreaking research in Nature. Their paper, “Accurate predictions of disordered protein ensembles with STARLING,” leverages a large protein-language model to achieve a new standard of accuracy. By removing the typical financial and administrative barriers to scaling such data-heavy research, RIS provided an essential foundation for this discovery.

The Challenge of Disordered Proteins 

About 70% of proteins inside a cell contain intrinsically disordered regions, segments that lack a fixed, stable structure and exist as dynamic, flexible ensembles. Proteins are thus not always precise, fixed machines, but can be more dynamic than traditionally thought. Dr. Holehouse explains, “Initially, people assumed they didn’t exist, and that all proteins would have to fold to work. However, in the last 10 to 15 years, we’ve really come to appreciate that disordered regions are ubiquitous and crucial for a wide variety of different ways in which cells do interesting things.”

This context underscores a central challenge: understanding how scientists, enabled by RIS, investigate and overcome the barriers posed by these essential but elusive protein regions.

WashU is an internationally known powerhouse in understanding IDRs. Driven by pioneering computational work from Rohit Pappu (WashU Biomedical Engineering) stretching across two decades, rules and principles that explain how disordered proteins behave have been slowly but surely elucidated, revealing a hidden “molecular grammar” underlying how these regions work (see Ruff et al. Cell 2026). Pappu’s work and the crucible of IDR researchers it has enabled has driven the recruitment of many outstanding scientists in this field, including Andrea Soranno, Meredith Jackrel, Alex Holehouse, Yifan Dai, and, most recently, Elizabeth Draganova and Anita Đonlić. 

While scientists are skilled at identifying IDRs through sequence information, determining their specific functions remains a significant hurdle. In particular, IDRs are notably prevalent in cellular systems responsible for information processing, systems that take in signals and make decisions. This connection highlights the complexity of linking sequence recognition to functional understanding. 

One of the most pressing challenges is the clinical blind spot created by these regions. While medicine is effective at identifying mutations in folded domains, researchers are often blind to mutations in disordered regions. While computational methods for predicting so-called “variant effects” are becoming highly accurate for folded domains, the same cannot be said for disordered regions. As Dr. Holehouse warns, “this will at some point likely become the limiting factor for personalized medicine – taking a person’s genome and inferring how somatic or germline mutations are causing disease.” 

A specific example of this risk is found in heart muscle proteins called troponins. Mutations in the IDRs of these proteins can lead to cardiomyopathies, a group of diverse, often genetic or acquired, diseases of the heart muscle that impair its ability to pump blood, frequently leading to heart failure, arrhythmias, and fluid retention. As Dr. Holehouse notes, “You might choose the wrong direction if you don’t understand what the mutation is doing.” Current research aims to predict whether these mutations make a protein more or less responsive, allowing for better therapeutic decision-making. In this context, and in close collaboration with Dr. Michael Greenberg and Dr. Andrea Soranno, Holehouse is working on ways to understand what drives troponin-related cardiomyopathies, funded by an R01 grant from the National Heart, Lung, and Blood Institute.  

Leveraging the Computational Power of Compute2 and Data Storage Infrastructure 

Deep Learning and Language Models: 

Overcoming these computational hurdles requires significant processing power. Leveraging the dedicated GPU and CPU resources of Compute2, the university’s high-performance computing cluster, the Holehouse lab successfully trained a large protein-language model designed to map evolutionary relationships within disordered protein regions. This level of access is a significant advantage; as Holehouse points out, the necessary GPU power is often inaccessible to researchers in many parts of the world. Reflecting on the impact at WashU, Dr. Holehouse adds: “The fact that a grad student can be wielding those kinds of resources at a home institution is really exceptional.” 

Workflow and Storage 

The transition from Compute1 to Compute2 offers great flexibility, allowing researchers to choose between containerized ecosystems like Docker and bare-metal analysis. RIS further sets itself apart with integrated storage that streamlines handling of the massive datasets needed for these models. Dr. Holehouse notes, “RIS Data Storage has been amazing. It’s almost too easy that when you get full, you can open a ticket and be like, ‘I’d like 5 more terabytes, please.” He adds, “Having one unified view of that storage, whether from our compute or local resources, is amazing.”

From Computation to Cure

By bridging massive computational power and advanced molecular biology, Dr. Holehouse’s lab, with vital support from the RIS infrastructure, is shedding light on intrinsically disordered regions. Their work with STARLING challenges traditional views of protein structure and paves the way for personalized medicine and targeted therapeutics for complex diseases like cardiomyopathies. As researchers continue decoding the “molecular grammar” of these flexible proteins, deep learning, robust storage, and high-performance computing will remain essential to turning these biological blind spots into medical breakthroughs. To explore their latest research, software tools, and ongoing projects, visit The Holehouse Lab.