Start Date
1-5-2026 12:00 PM
End Date
1-5-2026 1:00 PM
Description
Introduction & Problem
Intrinsically disordered proteins (IDPs) and their associated “fuzzy” complexes represent a major unresolved challenge in structural biology. Unlike traditional protein complexes that adopt stable, well-defined tertiary structures, fuzzy complexes retain conformational heterogeneity even in the bound state [1]. This dynamic behavior is functionally important in processes such as transcriptional regulation, signaling, and phase separation, but it complicates both experimental characterization and computational prediction. Recent advances in deep learning-based protein structure prediction, particularly AlphaFold2 and its successors, have demonstrated near-experimental accuracy for ordered proteins [2]. However, these models are fundamentally trained on static structural data, primarily derived from crystallography, which biases them toward predicting a single dominant conformation.
Standard model performance evaluation metrics such as DockQ or root-mean-square deviation assume a single “correct” structure, making them poorly suited for systems where disorder is biologically meaningful. As a result, current methods may appear accurate while systematically misrepresenting conformational ensembles [3]. The central problem addressed in this study is whether state-of-the-art structure prediction models can meaningfully capture or approximate the structural heterogeneity of fuzzy complexes, or whether they fundamentally fail under these conditions.
To address this, we perform a systematic benchmarking study of multiple next-generation protein structure prediction models on a curated dataset of experimentally characterized fuzzy complexes. This work aims to quantify model performance not only against static structural references, but also against experimental constraints that directly probe conformational variability, thereby providing a more biologically relevant evaluation framework.
Methods
We curated a dataset of 76 intrinsically disordered or fuzzy protein complexes from FuzDB, ensuring the availability of both high-resolution Protein Data Bank (PDB) structures and corresponding nuclear magnetic resonance (NMR) restraint data from the Biological Magnetic Resonance Data Bank (BMRB). This dataset represents a diverse set of binding modes and disorder-retaining interactions, enabling a broad assessment of model behavior across different classes of fuzzy complexes.
Four structure prediction models were evaluated: AlphaFold3, AlphaFold2-Multimer, Chai- 1, and Boltz-1. For each complex, predicted structures were generated and compared against experimental references using multiple complementary metrics. Structural agreement with PDB-derived complexes was quantified using DockQ scores and predicted interface TM (ipTM) values, capturing global and interface-level accuracy. However, to address the limitations of static benchmarking, we incorporated experimentally derived nuclear Overhauser effect (NOE) distance restraints as an orthogonal validation measure. NOE violation analysis was performed by comparing predicted interatomic distances against curated restraint datasets, providing a direct assessment of whether predicted conformations satisfy experimentally observed distance constraints.
This dual-metric framework allows for a more nuanced evaluation of model performance in the context of conformational disorder. Data processing, structure evaluation, and statistical analysis were conducted using Python-based pipelines, integrating structural bioinformatics tools and custom scripts for large-scale benchmarking.
Results & Findings
Across all evaluated models, predicted structures fall within the “Acceptable” DockQ range, with AlphaFold3 achieving the highest mean score (0.381), followed closely by AlphaFold2-Multimer (0.364), and more distantly by Boltz-1 (0.313) and Chai-1 (0.264). However, none of the models approach the “Medium quality” threshold (DockQ > 0.49), indicating that all predictors struggle to accurately model fuzzy or intrinsically disordered complexes despite apparent differences in ranking.
A second key finding is the lack of meaningful separation between models when evaluated against experimental NMR constraints. NOE violation distributions are nearly identical across all predictors, with median violation counts around 400 to 500 and long upper tails extending beyond 3,000 violations. This suggests that constraint violations are driven primarily by the inherent structural ambiguity of the target systems rather than differences in model architecture or predictive capability.
Additional analyses reveal several secondary patterns. Model performance exhibits high variance across targets, indicating that a subset of particularly difficult complexes consistently drives prediction failure across all methods. Chai-1 emerges as the weakest structural predictor, with the lowest mean DockQ and relatively low variability, suggesting consistently poor performance rather than occasional failure.
As a secondary outcome, this study produces a curated dataset of fuzzy protein complexes paired with NOE restraint data, which is released as a community resource to support future benchmarking efforts.
Beyond Folded Proteins: Assessing the Limits of AI Structure Prediction for Intrinsically Disordered Complexes
Introduction & Problem
Intrinsically disordered proteins (IDPs) and their associated “fuzzy” complexes represent a major unresolved challenge in structural biology. Unlike traditional protein complexes that adopt stable, well-defined tertiary structures, fuzzy complexes retain conformational heterogeneity even in the bound state [1]. This dynamic behavior is functionally important in processes such as transcriptional regulation, signaling, and phase separation, but it complicates both experimental characterization and computational prediction. Recent advances in deep learning-based protein structure prediction, particularly AlphaFold2 and its successors, have demonstrated near-experimental accuracy for ordered proteins [2]. However, these models are fundamentally trained on static structural data, primarily derived from crystallography, which biases them toward predicting a single dominant conformation.
Standard model performance evaluation metrics such as DockQ or root-mean-square deviation assume a single “correct” structure, making them poorly suited for systems where disorder is biologically meaningful. As a result, current methods may appear accurate while systematically misrepresenting conformational ensembles [3]. The central problem addressed in this study is whether state-of-the-art structure prediction models can meaningfully capture or approximate the structural heterogeneity of fuzzy complexes, or whether they fundamentally fail under these conditions.
To address this, we perform a systematic benchmarking study of multiple next-generation protein structure prediction models on a curated dataset of experimentally characterized fuzzy complexes. This work aims to quantify model performance not only against static structural references, but also against experimental constraints that directly probe conformational variability, thereby providing a more biologically relevant evaluation framework.
Methods
We curated a dataset of 76 intrinsically disordered or fuzzy protein complexes from FuzDB, ensuring the availability of both high-resolution Protein Data Bank (PDB) structures and corresponding nuclear magnetic resonance (NMR) restraint data from the Biological Magnetic Resonance Data Bank (BMRB). This dataset represents a diverse set of binding modes and disorder-retaining interactions, enabling a broad assessment of model behavior across different classes of fuzzy complexes.
Four structure prediction models were evaluated: AlphaFold3, AlphaFold2-Multimer, Chai- 1, and Boltz-1. For each complex, predicted structures were generated and compared against experimental references using multiple complementary metrics. Structural agreement with PDB-derived complexes was quantified using DockQ scores and predicted interface TM (ipTM) values, capturing global and interface-level accuracy. However, to address the limitations of static benchmarking, we incorporated experimentally derived nuclear Overhauser effect (NOE) distance restraints as an orthogonal validation measure. NOE violation analysis was performed by comparing predicted interatomic distances against curated restraint datasets, providing a direct assessment of whether predicted conformations satisfy experimentally observed distance constraints.
This dual-metric framework allows for a more nuanced evaluation of model performance in the context of conformational disorder. Data processing, structure evaluation, and statistical analysis were conducted using Python-based pipelines, integrating structural bioinformatics tools and custom scripts for large-scale benchmarking.
Results & Findings
Across all evaluated models, predicted structures fall within the “Acceptable” DockQ range, with AlphaFold3 achieving the highest mean score (0.381), followed closely by AlphaFold2-Multimer (0.364), and more distantly by Boltz-1 (0.313) and Chai-1 (0.264). However, none of the models approach the “Medium quality” threshold (DockQ > 0.49), indicating that all predictors struggle to accurately model fuzzy or intrinsically disordered complexes despite apparent differences in ranking.
A second key finding is the lack of meaningful separation between models when evaluated against experimental NMR constraints. NOE violation distributions are nearly identical across all predictors, with median violation counts around 400 to 500 and long upper tails extending beyond 3,000 violations. This suggests that constraint violations are driven primarily by the inherent structural ambiguity of the target systems rather than differences in model architecture or predictive capability.
Additional analyses reveal several secondary patterns. Model performance exhibits high variance across targets, indicating that a subset of particularly difficult complexes consistently drives prediction failure across all methods. Chai-1 emerges as the weakest structural predictor, with the lowest mean DockQ and relatively low variability, suggesting consistently poor performance rather than occasional failure.
As a secondary outcome, this study produces a curated dataset of fuzzy protein complexes paired with NOE restraint data, which is released as a community resource to support future benchmarking efforts.