{"cells":[{"cell_type":"markdown","metadata":{},"source":"# AI-Assisted Systematic Optimization of RNA 3D Structure Prediction: From Full Automation to Human-in-the-Loop\n\n**Anis Marrouchi¹, Omar Mejawel¹, Razi (AI Research Agent)¹**\n\n¹ Bio-AI & Agritech Department, Noqta, Tunisia","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## Abstract\n\nWe present a detailed account of our participation in the Stanford RNA 3D Folding Part 2 Kaggle competition ($75K prize pool, 2026). Starting with a fully automated \"autoresearch\" pipeline inspired by Karpathy's vision — where an AI agent autonomously generates hypotheses, runs experiments, scores results, and iterates — we discovered that the Kaggle kernel environment's constraints forced a pivotal shift to human-in-the-loop collaboration. This paper documents our complete journey: 74+ notebook versions, systematic ablation studies, and the insights that emerged. Our best score of **0.432 TM-score** was achieved through a hybrid Template-Based Modeling (TBM) + Protenix deep learning approach with chunked inference and a finetuned checkpoint. In our latest iteration, we integrate DRfold2 — an RNA-specific deep learning model with a pre-trained RNA Composite Language Model — creating a three-model hybrid (TBM + DRfold2 + Protenix) that leverages the strengths of each approach. We provide a comprehensive analysis of what worked, what didn't, and why — contributing to the community's understanding of RNA structure prediction and AI-assisted research methodology.\n\n**Keywords:** RNA 3D structure prediction, template-based modeling, Protenix, automated research, human-in-the-loop, Kaggle competition","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 1. Introduction\n\n### 1.1 The RNA Structure Prediction Challenge\n\nRNA molecules fold into complex three-dimensional structures that determine their biological function. Unlike proteins, where AlphaFold has largely solved the structure prediction problem, RNA 3D structure prediction remains one of biology's grand challenges. The Stanford RNA 3D Folding Part 2 competition tasked participants with predicting the 3D coordinates (x, y, z) of each nucleotide's C1' atom for 28 test RNA sequences, scored by TM-score — a scale-invariant metric of structural similarity.\n\n### 1.2 The Autoresearch Vision\n\nInspired by Andrej Karpathy's concept of \"autoresearch\" — fully autonomous AI-driven scientific experimentation — we set out to build a pipeline where our AI agent (Razi) would:\n\n1. **Formulate hypotheses** based on prior results\n2. **Generate experiment code** (Kaggle notebooks)\n3. **Push and run** experiments via the Kaggle API\n4. **Score results** and decide keep/discard\n5. **Iterate** toward optimal solutions\n\nThis vision met reality in unexpected ways, teaching us fundamental lessons about the boundaries of automation in scientific research.\n\n### 1.3 Team Composition\n\n- **Anis Marrouchi** — Competition lead, infrastructure, Kaggle account management\n- **Omar Mejawel** — Lead researcher, domain expert (M.Sc. Molecular & Cellular Biology), second Kaggle account operator\n- **Razi** — AI research agent (Claude-based), hypothesis generation, code writing, analysis, experiment orchestration","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 2. Related Work\n\n### 2.1 Deep Learning Approaches for RNA 3D Structure Prediction\n\nThe past two years have seen an explosion of deep learning methods targeting RNA structure prediction, though the field remains significantly behind protein structure prediction in accuracy and reliability.\n\n**DRfold2** (Li et al., PLoS Biology, Feb 2026) represents the current state-of-the-art for single-sequence RNA structure prediction. It combines a novel pre-trained RNA Composite Language Model (RCLM) with a denoising structure module for end-to-end prediction. Critically, DRfold2 demonstrates **high complementarity with AlphaFold3** — combining both methods yields statistically significant accuracy gains over either alone. The RCLM achieves a >100% increase in contact prediction precision compared to existing methods by capturing co-evolutionary patterns from single sequences without requiring multiple sequence alignments (MSA). This is particularly relevant to our work, as Protenix (our deep learning backbone) lacks such RNA-specific language model integration.\n\n**trRosettaRNA2** (bioRxiv, Apr 2025) takes a different approach by integrating an auxiliary secondary structure (SS) prior module, pre-trained on extensive SS data, to generate base-pairing priors. This structure-aware attention mechanism enables conformer prediction — predicting multiple plausible conformations for the same sequence. Our pipeline does not incorporate secondary structure information, representing an unexplored avenue for improvement.\n\n**RhoFold+** (Shen et al., 2024) utilizes large language models to predict atom coordinates for single-chain RNA, while **RoseTTAFold2NA** (Baek et al., 2024) and general-purpose models like **Boltz-1**, **Chai-1**, and **HelixFold3** extend AlphaFold3-style architectures to handle RNA alongside proteins and small molecules.\n\n**ProRNA3D-single** (Roche et al., Cell Systems, Sep 2025) introduces geometric attention-enabled pairing of heterogeneous biological language models for single-sequence protein-RNA complex structure prediction, outperforming AlphaFold 3 without requiring MSA input. This demonstrates that language model representations can compensate for the lack of evolutionary information typically provided by MSAs.\n\n### 2.2 Benchmarking and Known Limitations\n\nA critical independent benchmark study, **\"Limits of Deep Learning-Based RNA Prediction Methods\"** (bioRxiv, May 2025), evaluated nine single-chain RNA prediction methods and four RNA complex prediction methods. Their findings directly corroborate our experimental observations:\n\n1. **Prediction success strongly correlates with resemblance to known structures** — current methods recognize recurring motifs rather than generalizing to novel folds. This aligns precisely with our finding that Template-Based Modeling accounts for ~95% of our hybrid score, and that all 28 test targets have homologs in the training set at 30-50% identity.\n\n2. **Quality estimation for RNA models is unreliable** — it is \"not possible to reliably identify the correctly predicted models with today's methods.\" This explains our observation that RNAPro produced a uniform pLDDT of 50 across all predictions, and why confidence-based model selection (our consensus picking experiment) failed to improve scores.\n\n3. **Accurate predictions are limited to RNAs with well-defined or regular secondary structures**, with accuracy notably higher for complexes involving extensive canonical base pairing.\n\n### 2.3 RNA Language Models and Co-evolutionary Signals\n\nThe **RNA LM-Augmented AlphaFold 3** approach (OpenReview, Oct 2025) investigates how self-supervised RNA language models, trained on millions of RNA sequences, can enhance AlphaFold 3 for RNA structure prediction. This work highlights a fundamental gap in current general-purpose structure prediction tools: they lack RNA-specific learned representations. Protenix, the deep learning component in our pipeline, was designed primarily for protein structure prediction and does not incorporate RNA language model embeddings — a likely contributor to its limited standalone performance on RNA targets.\n\n### 2.4 The Broader Context: CASP17 and Competition Landscape\n\nThe 17th Critical Assessment of Structure Prediction (CASP17), scheduled for April 2026, will serve as the definitive benchmark for RNA structure prediction methods. The Stanford RNA 3D Folding competition on Kaggle provides a parallel, community-driven evaluation with real-time leaderboard feedback. Our systematic documentation of what works and what fails in this competition setting contributes practical knowledge that complements the controlled evaluations of CASP.\n\nA comprehensive review of computational approaches for RNA structure prediction and design (ScienceDirect, Feb 2026) surveys the full landscape from fragment assembly to deep learning, noting that the field is at an inflection point where data availability and methodological advances are converging — but reliable, generalizable RNA 3D prediction remains an open problem.\n\n### 2.5 Positioning Our Work\n\nOur contribution is distinct from the above in several ways:\n\n| Aspect | Existing Work | Our Work |\n|--------|--------------|----------|\n| Focus | Model architecture development | Systematic optimization within fixed architecture |\n| Method | Novel neural networks | TBM + Protenix hybrid with ablation |\n| Insight type | What models *can* do | What interventions *don't* work (negative results) |\n| Process | Traditional research | AI-assisted autoresearch with human-in-the-loop |\n| Template selection | Assumed improvable | Proven near-optimal with simple seq identity |\n| Evaluation | Cross-validation / CASP | Live competition leaderboard (28 blind targets) |\n\nOur strongest contribution to the literature is the comprehensive negative result on template selection optimization — demonstrating across 5 independent methods that simple sequence identity ranking is robust and hard to beat — and the practical documentation of an AI-assisted research methodology operating under real computational constraints.","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 3. Technical Background\n\n### 2.1 TM-Score (Template Modeling Score)\n\nTM-score is the competition's evaluation metric. Unlike RMSD, TM-score is **length-normalized and scale-invariant**, ranging from 0 to 1:\n\n- **TM > 0.5**: Same fold topology\n- **TM > 0.7**: High structural similarity\n- **TM = 1.0**: Perfect match\n\nA critical finding in our work: because TM-score is scale-invariant, **rescaling coordinates has no effect** (proven experimentally in v63, where rescaling *hurt* the score from 0.404 to 0.401 due to numerical artifacts).\n\n### 2.2 Template-Based Modeling (TBM)\n\nTBM predicts a target RNA's structure by finding homologous sequences with known structures in the training set and transferring their coordinates:\n\n1. **Sequence alignment**: Global pairwise alignment (Biopython `PairwiseAligner`) between each test sequence and all 5,744 training sequences\n2. **Template selection**: Rank by sequence identity percentage, filter by `MIN_PERCENT_IDENTITY` threshold\n3. **Coordinate adaptation**: Map template coordinates to the target sequence via the alignment, handling insertions (linear interpolation) and deletions (skip)\n\n**Key insight**: All 28 test targets have matches in the training set at 30-50% segment-level identity (not 100% as initially assumed). The bottleneck is coordinate adaptation quality, not template availability.\n\n### 2.3 Protenix (Deep Learning Structure Prediction)\n\nProtenix is ByteDance's open-source implementation inspired by AlphaFold3, adapted for general biomolecular structure prediction. It uses:\n\n- **Diffusion-based coordinate generation**: Iteratively denoises random coordinates into structured conformations\n- **Pairformer architecture**: Processes sequence and pair representations through transformer blocks\n- **N_cycle**: Number of recycling iterations through the trunk (default: 4)\n- **N_step**: Number of diffusion denoising steps (default: 200)\n\nIn our pipeline, Protenix serves as a **fallback** for targets where TBM templates are insufficient (below quality thresholds) or to supplement TBM predictions to reach the required 5 samples per target.\n\n### 2.4 Chunked Inference\n\nRNA sequences longer than `MAX_SEQ_LEN` (512 nucleotides) cannot be processed in a single Protenix pass due to GPU memory constraints. Chunked inference:\n\n1. **Split**: Divide the sequence into overlapping windows of 512nt with `CHUNK_OVERLAP` nucleotides of overlap\n2. **Predict**: Run Protenix independently on each chunk\n3. **Align**: Use Kabsch alignment on overlap regions to superpose adjacent chunks\n4. **Blend**: Linear weight ramp in overlap zones for smooth transitions\n\nThis is critical for two large test targets: 9MME (4640 nt, requiring 11-18 chunks) and 9ZCC (1460 nt, 4-5 chunks).\n\n### 2.5 DRfold2 (RNA-Specific Deep Learning)\n\nDRfold2 (Li et al., PLoS Biology, Feb 2026) represents a fundamentally different approach from Protenix. While Protenix is a general-purpose biomolecular structure prediction tool adapted from protein-first architectures, DRfold2 is **purpose-built for RNA**:\n\n- **RNA Composite Language Model (RCLM)**: Pre-trained on RNA sequences to capture co-evolutionary patterns and RNA-specific structural motifs, achieving >100% improvement in contact prediction precision over existing methods\n- **Denoising structure module**: End-to-end coordinate prediction from single sequences, no MSA required\n- **4-model ensemble**: Runs configs cfg_95, cfg_96, cfg_97, cfg_99 internally, with Selection and Optimization steps to pick the best geometry\n- **Arena refinement**: C++ all-atom energy minimization for final structure polish\n\n**Integration in our pipeline**: DRfold2 serves as **Phase 2a** — running before Protenix on targets ≤600 nucleotides (where it's fast enough). Protenix then handles remaining large targets (9MME at 4640nt, 9ZCC at 1460nt) via chunked inference. This creates a three-tier prediction hierarchy:\n\n```\nPhase 1: TBM (template-based)          → 6 targets fully covered\nPhase 2a: DRfold2 (RNA-specific DL)    → ~20 targets ≤600nt  \nPhase 2b: Protenix (general DL + chunking) → 2 large targets\n```\n\n**Packaging for Kaggle**: DRfold2 (~1.4GB including model weights) was packaged as a Kaggle dataset (`omarmejawel/drfold2-rna-structure-prediction`), extracted at runtime via `tar -xzf`, with Arena pre-compiled as a static binary. The no-internet constraint required all dependencies to be bundled.\n\n### 2.6 Post-Processing\n\nOur geometric post-processing pipeline corrects physically implausible predictions:\n\n- **Sequential distance normalization**: Enforces ~6.2Å mean C1'-C1' distance between consecutive nucleotides\n- **Steric clash resolution**: Separates atoms closer than van der Waals radii\n- **Backbone guard**: Skips post-processing entirely if mean backbone distance > 100Å (indicates Protenix output at wildly different scale)","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 4. The Autoresearch Pipeline\n\n### 3.1 Architecture\n\n```\n┌─────────────────────────────────────────────────┐\n│                 Razi (AI Agent)                  │\n│  ┌───────────┐  ┌──────────┐  ┌──────────────┐ │\n│  │ Hypothesis │→ │ Code Gen │→ │ Kaggle Push  │ │\n│  │ Generator  │  │ (Notebook)│  │ (API)        │ │\n│  └───────────┘  └──────────┘  └──────────────┘ │\n│        ↑                            │           │\n│  ┌───────────┐              ┌──────────────┐    │\n│  │ Analysis  │ ←────────── │ Score Check  │    │\n│  │ & Decision│              │ (Cron/Poll)  │    │\n│  └───────────┘              └──────────────┘    │\n└─────────────────────────────────────────────────┘\n```\n\n**Components:**\n- **Kaggle API integration**: `kernels/push` for notebook deployment, `submissions/list` for score polling\n- **Cron jobs**: Hourly score checks, daily progress reports\n- **Dual accounts**: `anismarrouchi` (primary) + `omarmejawel` (secondary) for 10 submissions/day\n- **Memory system**: Persistent experiment logs, decisions, and lessons learned across sessions\n\n### 3.2 The Automation-to-Human Pivot\n\nOur initial pipeline worked for simple parameter sweeps (changing thresholds, toggling features). However, we hit fundamental limitations:\n\n**API push failures**: Kaggle API-pushed notebooks mount datasets at `/kaggle/input/datasets/<owner>/<slug>/` while UI-created notebooks mount at `/kaggle/input/<slug>/`. This path mismatch caused persistent failures for experiments requiring specific dataset configurations (Protenix checkpoints, dependency wheels).\n\n**Package management**: API-pushed notebooks couldn't properly wire `packagemanager` kernel dependencies. Complex dependency chains (biotite cp312, rdkit cp312) required UI-based fork and manual dataset attachment.\n\n**The lesson**: After 6+ failed API push attempts on Omar's account (Exp2 v3-v6, Exp3 v12-v14), we pivoted to **UI fork + manual edit** as the reliable execution path. The AI agent prepares the code and instructions; the human executes in the Kaggle UI.\n\nThis wasn't a failure of the autoresearch concept — it was a discovery about the **boundary conditions** of automation. The Kaggle environment's quirks created a \"last mile\" problem that required human hands.","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 5. Experimental Journey\n\n### 4.1 Phase 1: Baseline Establishment (v5-v26)\n\n| Version | Hypothesis | Score | Decision |\n|---------|-----------|-------|----------|\n| v5 | A-form helix baseline | 0.087 | Baseline |\n| v12 | TBM + Protenix, PID≥50 | 0.378 | Keep ✅ |\n| v16 | Stricter: sim≥0.15, PID≥55 | 0.380 | Keep ✅ |\n| v19 | Pure Protenix (no TBM) | 0.373 | Discard ❌ |\n| v21 | Multi-seed Protenix | 0.276 | Discard ❌ |\n| v24 | top_n=50 templates | 0.387 | Best at the time |\n| v26 | Finetuned checkpoint alone | 0.347 | Discard ❌ |\n\n**Key Phase 1 findings:**\n- TBM dominates the hybrid score — templates are the primary signal\n- **Complexity hurts**: Every \"clever\" addition (multi-seed, finetuned-only) regressed\n- Score plateau at ~0.387 with TBM-only approach\n\n### 4.2 Phase 2: Adopting the Reference Baseline (v59-v61)\n\nWe identified a public notebook scoring 0.413 (Protenix + TBM hybrid by artemevstafyev) and adopted it as our new foundation.\n\n| Version | Change | Score | Delta |\n|---------|--------|-------|-------|\n| v59 | Reference baseline (Protenix + TBM) | 0.395 | New baseline |\n| v61 | Post-processing guard (skip if BB > 100Å) | 0.404 | +0.009 ✅ |\n| v62 | Rescaling Protenix coordinates | FAILED | Format error |\n| v63 | Fixed rescaling | 0.401 | -0.003 ❌ |\n\n**Critical discovery**: TM-score is scale-invariant. Rescaling coordinates cannot help and only introduces numerical noise. This was proven experimentally (v63) after we initially assumed scale mismatch was a problem.\n\n### 4.3 Phase 3: Top Notebook Analysis & Chunking (v65-v68)\n\nWe systematically analyzed the leaderboard's top notebooks to identify proven techniques:\n\n| Rank | Team | Score | Key Techniques |\n|------|------|-------|---------------|\n| #1 | oracle (sigmaborov) | 0.554 | CHUNK_OVERLAP=256, MIN_SIM=0.10, MIN_PCTID=55 |\n| #2 | AyPy | 0.488 | Unknown (private) |\n| #3 | Brad | 0.477 | Unknown |\n| #4 | d4t4 | 0.460 | Unknown |\n\nFrom sigmaborov's public notebook, we identified two key techniques: **chunked inference** and **finetuned RNA3DB checkpoint**.\n\n| Version | Change | Score | Delta from v61 |\n|---------|--------|-------|----------------|\n| v65 | Chunking + finetuned checkpoint | **0.432** | **+0.028** ✅ |\n| v67 | Chunking only (no finetuned) | 0.428 | +0.024 |\n| v68 | Chunking + rule-based template re-ranking | 0.410 | +0.006 |\n\n**Ablation results:**\n- **Chunking alone**: +0.024 (the biggest single gain, 0.404 → 0.428)\n- **Finetuned checkpoint**: +0.004 (modest but real, 0.428 → 0.432)\n- **Template re-ranking**: -0.018 from 0.428 (HURT)\n\n### 4.4 Phase 4: Template Selection Optimization (v68-v70, Omar Exp2)\n\nBased on cross-validation analysis showing that oracle template selection could yield +0.042, we invested heavily in improving template choice:\n\n**Deep analysis conducted:**\n- Trained a GBR template quality scorer (R²=0.78, 20K pairs, 6 features)\n- Found sequence identity accounts for 76% of prediction signal\n- Identified 66% of the gap comes from 16% of targets (structural homologs with low seq identity)\n- Multi-template top-5 oracle: +0.06 potential\n\n**Experimental results:**\n| Experiment | Method | Score | vs v65 |\n|-----------|--------|-------|--------|\n| v68 | Rule-based penalty (coverage, breaks) | 0.410 | -0.022 ❌ |\n| v69 | ML scorer (GBR) | 0.394 | -0.038 ❌ |\n| v70 | ML scorer + finetuned | 0.379 | -0.053 ❌ |\n| Omar Exp2 | Consensus pick (pairwise alignment) | 0.417 | -0.015 ❌ |\n\n**Conclusion**: Template re-ranking is a **consistently negative** intervention. The simple first-by-sequence-identity approach is already near-optimal. Attempting to be \"smarter\" about template selection introduces more error than it corrects. This is a strong negative result worth publishing.\n\n### 4.5 Phase 5: Configuration Optimization (v71-v74)\n\nComparing our notebook cell-by-cell with the #1 public solution revealed configuration differences:\n\n| Parameter | Our v65 | Sigmaborov #1 | Top-2 |\n|-----------|---------|---------------|-------|\n| CHUNK_OVERLAP | 64 | 256 | 128 |\n| MIN_SIMILARITY | 0.0 | 0.10 | 0.0 |\n| MIN_PERCENT_IDENTITY | 50.0 | 55.0 | 50.0 |\n| N_step (diffusion) | 200 (default) | 200 | 200 |\n| N_cycle (recycling) | 4 (default) | 4 | 4 |\n| Finetuned ckpt | Yes | No | No |\n\nNotable: Neither the #1 nor #2 public solution uses the finetuned checkpoint or increased diffusion steps. Their advantage comes primarily from **CHUNK_OVERLAP=256** and **stricter TBM quality thresholds**.\n\n| Version | Changes | Status |\n|---------|---------|--------|\n| v71-v73 | Various arg name errors | Failed ❌ |\n| v74 | OVERLAP=256 + MIN_SIM=0.10 + MIN_PCTID=55 + N_step=250 + N_cycle=12 + finetuned | Running |\n| Omar Exp5 | OVERLAP=256 only (clean isolation test) | Running |\n\n### 4.6 Phase 6: DRfold2 Integration (v76)\n\nMotivated by recent literature showing that RNA-specific models with pre-trained language models outperform general-purpose tools on RNA targets (see Section 2), we integrated DRfold2 into our pipeline — creating a **three-model hybrid**: TBM + DRfold2 + Protenix.\n\n**The insight**: Protenix was designed for proteins and adapted for general biomolecules. It lacks RNA-specific learned representations. DRfold2's RCLM captures RNA co-evolutionary patterns that Protenix simply cannot access. By using DRfold2 for small-to-medium RNA targets (≤600nt) and Protenix only for large sequences requiring chunking, we combine the best of both worlds.\n\n**Implementation steps**:\n1. Cloned DRfold2 from GitHub, downloaded 1.3GB model weights from Zhang Lab\n2. Compiled Arena refinement tool (C++ → static binary for Kaggle's Linux env)\n3. Packaged code + weights + binary as a Kaggle dataset (1.4GB tar.gz)\n4. Integrated as Phase 2a: DRfold2 runs on targets ≤600nt, sorted smallest-first\n5. Time-budgeted at 1 hour to leave runtime for Protenix Phase 2b on large targets\n6. Protenix handles 9MME (4640nt) and 9ZCC (1460nt) via chunked inference\n\n**v76/v77 results (0.364/0.362)**: Our first DRfold2 attempts **failed catastrophically** — but not because of DRfold2. The dataset wasn't properly mounted (private dataset on a different account), so DRfold2 was silently skipped. The real culprit was the simultaneous config changes: `N_step=250`, `N_cycle=12`, `MIN_SIM=0.10`, `MIN_PCTID=55`. These \"improvements\" caused a **-0.070 regression** from our best score — the worst result in our entire experiment history.\n\n**The lesson**: Never change 4 variables simultaneously. We violated our own scientific method under time pressure, and the results were unambiguous.\n\n**v79 (running)**: Clean experiment — v65's proven config (default N_step/N_cycle, relaxed thresholds) + DRfold2 properly mounted via public dataset (`noqta-tunisia/drfold2-rna-structure-prediction`). This isolates the DRfold2 variable from the config regression.\n\n### 4.7 Phase 7: The Contamination Discovery (v81)\n\nOn March 16, we made a critical forensic discovery by examining actual kernel logs rather than speculating about failure causes. Three revelations emerged:\n\n**1. Silent Code Contamination**: The template re-ranking code introduced in v68 (which scored 0.410, a known -0.022 regression) **persisted in every subsequent version** from v69 through v80. This code changed TBM's template selection from simple sequence identity sorting to an `_template_quality_score` adjustment that penalizes templates by coverage and backbone breaks. The effect: TBM went from covering **6/28 targets** (v65, optimal) to **19/28 targets** (v80, suboptimal) — forcing poor-quality TBM predictions onto targets that should have received Protenix deep learning predictions.\n\n**2. DRfold2 Never Actually Ran**: The glob pattern found the dataset *directory* rather than the `.tar.gz` file inside it. `tar` received a directory path and failed silently. Every v76-v80 \"DRfold2 experiment\" was actually Protenix-only with the bad re-ranking.\n\n**3. Boltz-2 Never Actually Ran**: Missing `mashumaro` dependency caused import failure. Boltz-2 covered 0/9 targets in 86 seconds of failed attempts.\n\n**Experiments requiring re-evaluation:**\n\n| Version | Original Attribution | Actual Cause | Contaminated? |\n|---------|---------------------|-------------|---------------|\n| v65 (0.432) | Chunking + finetuned | Clean baseline | ✅ Clean |\n| v67 (0.428) | Chunking only | Clean ablation | ✅ Clean |\n| v68 (0.410) | Re-ranking introduced | Re-ranking hurt | ✅ Clean (intentional test) |\n| v69 (0.394) | ML scorer | ML scorer + re-ranking | ⚠️ Contaminated |\n| v70 (0.379) | ML scorer v2 | ML scorer + re-ranking | ⚠️ Contaminated |\n| Omar Exp2 (0.417) | Consensus pick | Consensus + re-ranking? | ⚠️ Needs verification |\n| v74-77 (0.362-0.364) | N_step/N_cycle | N_step/N_cycle + re-ranking + DRfold2 failure | ⚠️ Triple contaminated |\n| v79 (0.376) | DRfold2 test | Re-ranking + DRfold2 failure | ⚠️ Contaminated |\n| v80 (0.386) | 4-model hybrid | Re-ranking + DRfold2 failure + Boltz2 failure | ⚠️ Triple contaminated |\n\n**Impact**: The v69/v70 ML scorer experiments may have scored worse than necessary because they inherited the re-ranking penalty. However, since re-ranking alone caused -0.022 (v68) while v69 scored -0.038, the ML scorer itself still likely hurt by ~0.016. The directional conclusion (ML scorer hurts) stands, but the magnitude is uncertain.\n\n**The meta-lesson**: In a fast-iteration competition setting, **code archaeology matters**. We assumed each version started clean from v65's base, but incremental edits accumulated undiscovered regressions. A proper experiment framework would diff each version against the known-good baseline and flag unintended changes — exactly what our autoresearch pipeline should have automated but didn't.\n\nv81 fixes: reverted to simple sequence identity sort, fixed DRfold2 glob pattern, added mashumaro for Boltz-2. Additionally, we created a **diagnostic notebook** (`rna-diagnostic-drfold2-boltz2-test`) to validate each component in isolation before running full experiments — a practice we should have adopted from the start.\n\n### 4.8 Failed Experiments Summary\n\n| Experiment | Hypothesis | Result | Lesson |\n|-----------|-----------|--------|--------|\n| Rescaling Protenix coords | Scale mismatch hurts TM-score | -0.003 | TM-score is scale-invariant |\n| Multi-seed Protenix | Ensemble improves predictions | -0.111 | Noise > signal for RNA |\n| Finetuned checkpoint alone | Better weights = better score | -0.040 | Only helps with TBM backbone |\n| Rule-based re-ranking | Coverage & break penalties | -0.022 | Over-penalizes good templates |\n| ML template scorer | Learned quality prediction | -0.038 to -0.053 | Seq identity already optimal |\n| Consensus picking | Pairwise alignment consensus | -0.015 | Adds noise, doesn't filter |\n| CHUNK_OVERLAP=128 | More overlap = better stitching | -0.016 | Tested without finetuned ckpt |\n| N_step=250 + N_cycle=12 | More diffusion steps + recycling | **-0.070** | **Worst regression** — defaults are optimal |\n| Stricter TBM (SIM≥0.10, PID≥55) | Filter low-quality templates | Part of -0.070 | Pushes good targets to weaker DL predictions |\n| RNAPro model | NVIDIA RNA-specific model | -0.046 (vs v65) | pLDDT=50 uniformly, checkpoint not ready |\n| Coordinate averaging | Average multi-template coords | N/A (cross-val) | 0.431 vs 0.567 oracle |","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 6. Key Scientific Insights\n\n### 5.1 Simplicity Wins in Template-Based RNA Prediction\n\nOur most important finding: **every attempt to add complexity to template selection reduced scores**. This includes rule-based penalties, machine learning scorers, consensus picking, and coordinate averaging. The simple first-match-by-sequence-identity approach is robust and hard to beat.\n\n**Why?** We hypothesize that template selection errors are not independent — a \"better\" template by one metric often introduces correlated structural errors in regions that happen to matter for TM-score. The highest-sequence-identity template, while not always the best structural match, provides the most consistent and least-biased prediction.\n\n### 5.2 Chunking Is the Single Most Impactful Technique\n\nFor long RNA sequences (>512 nt), chunked inference with overlap and Kabsch-aligned stitching provides the largest single improvement (+0.024 TM-score). Without chunking, large targets like 9MME (4640 nt) produce garbage predictions, dragging down the overall score.\n\n### 5.3 The 100Å Post-Processing Guard\n\nProtenix occasionally outputs coordinates at wildly different scales (10^16). Rather than attempting to rescale (which we proved doesn't help), simply skipping post-processing for these cases and letting the raw Protenix output through improved scores by +0.009.\n\n### 5.4 Template Re-Ranking Is a Trap\n\nDespite building a strong template quality scorer (R²=0.78) and identifying clear oracle headroom (+0.042), we could not translate this into actual gains. The fundamental problem: **our scorer correlates with TM-score in aggregate but makes target-level errors that are worse than the errors of simple sequence identity ranking**.\n\n### 5.5 More Diffusion Steps and Recycling Hurt — Defaults Are Optimal\n\nCounter-intuitively, increasing Protenix's diffusion steps (`N_step=250` vs default 200) and recycling iterations (`N_cycle=12` vs default 4) caused our **worst regression**: -0.070 from 0.432. This is consistent with a known phenomenon in diffusion models where over-sampling can degrade quality by accumulating numerical errors. Additionally, the 3x compute increase for N_cycle=12 likely caused timeouts on the largest target (9MME, 18 chunks), producing incomplete or garbage predictions.\n\nNotably, none of the top public notebooks override these defaults. The sigmaborov #1 solution (0.554) uses vanilla N_step=200 and N_cycle=4. **The default hyperparameters are well-tuned; competition participants should focus on architecture and data, not hyperparameter escalation.**\n\n### 5.6 Silent Code Contamination: The Hidden Cost of Fast Iteration\n\nOur most impactful debugging discovery was that a code change from v68 (template re-ranking) persisted undetected through 13 subsequent versions, silently degrading every experiment. This is a well-known problem in software engineering (\"regression bugs\") but takes on special significance in ML experimentation:\n\n- **Confounded results**: We attributed score drops to the variable we *intended* to change, when the *unintended* re-ranking was responsible for a significant portion of the regression\n- **False confidence in negative results**: \"ML scorer hurts by -0.038\" should have been \"ML scorer hurts by ~-0.016, and re-ranking hurts by ~-0.022\"\n- **Wasted GPU hours**: ~15 experiments ran with contaminated baselines\n\n**Recommendation for competition participants and autoresearch systems**: Maintain a **frozen baseline checkpoint** and diff every experiment against it before execution. Our autoresearch agent should have automated: `diff v65_cell8.py current_cell8.py | grep -v \"intended_change\"` — if any unintended changes exist, abort and alert.\n\n### 5.7 Read the Logs, Don't Speculate\n\nWhen v79 scored 0.376, our initial analysis listed 4 possible causes (\"DRfold2 may have alignment issues\", \"Arena binary may not work\", etc.). Examining the actual kernel logs took 2 minutes and revealed DRfold2 never ran at all. **Every minute spent speculating about hypothetical causes is a minute not spent reading actual error messages.** This applies equally to human researchers and AI agents.\n\n### 5.8 Protenix Template System Does Not Support RNA\n\nA critical architectural finding: Protenix v1's `InferenceTemplateFeaturizer` explicitly asserts `ctype == PROTEIN_CHAIN` (line 694 of `template_featurizer.py`). The `USE_TEMPLATE` flag has no effect for RNA-only inputs. This means the PDB_RNA structural template dataset used by some protein prediction workflows is architecturally inapplicable here.","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 7. The Human-in-the-Loop Discovery\n\n### 6.1 What We Planned\n\n```\nAI Agent → Generate Code → Push via API → Score → Iterate\n(Fully autonomous loop)\n```\n\n### 6.2 What Actually Worked\n\n```\nAI Agent → Generate Code → Human Edits in UI → Save & Run → Score → AI Analyzes → Next Hypothesis\n(Human-in-the-loop collaboration)\n```\n\n### 6.3 Why the Pivot Was Necessary\n\n1. **Environment mismatch**: Kaggle API and UI mount datasets at different paths\n2. **Dependency hell**: Package manager wiring differs between API-pushed and UI-created notebooks\n3. **Argument discovery**: Protenix config parameter names had to be discovered empirically through error messages (3 failed versions to find `--model.N_cycle` vs `--num_trunk_recycles` vs `--sample_diffusion.N_cycle`)\n4. **Fork inheritance**: UI-forked notebooks inherit working dataset/dependency configurations that API push cannot replicate\n\n### 6.4 Lessons for Autoresearch\n\nThe autoresearch paradigm is sound but requires:\n\n1. **Environment parity**: The execution environment must behave identically whether triggered by API or UI\n2. **Graceful degradation**: When automation fails, the system should generate clear human-executable instructions rather than retrying blindly\n3. **Hybrid orchestration**: The optimal loop is not fully autonomous but AI-driven with human execution at bottlenecks\n4. **Error-as-data**: Every failed experiment (including infrastructure failures) generates valuable information about the system's constraints","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 8. Recommendations for Future Autoresearch Systems\n\n### 7.1 Loop Architecture Improvements\n\nBased on our experience, the ideal autoresearch loop for Kaggle competitions should:\n\n```python\nclass AutoresearchLoop:\n    def run_experiment(self, hypothesis):\n        code = self.generate_code(hypothesis)\n        \n        # Try API push first\n        result = self.api_push(code)\n        if result.failed_infrastructure:\n            # Fall back to human-executable instructions\n            self.notify_human(\n                instructions=self.format_ui_instructions(code),\n                changes=self.diff_from_baseline(code),\n                estimated_runtime=self.predict_runtime(code)\n            )\n            result = self.wait_for_human_execution()\n        \n        # Score and analyze\n        score = self.poll_score(result)\n        analysis = self.analyze_result(score, hypothesis)\n        \n        # Update knowledge base\n        self.knowledge_base.record(hypothesis, score, analysis)\n        \n        # Generate next hypothesis\n        next_hypothesis = self.generate_next(\n            self.knowledge_base,\n            constraints=self.remaining_budget()\n        )\n        return next_hypothesis\n```\n\n### 7.2 Key Design Principles\n\n1. **One variable at a time**: Never change multiple parameters simultaneously. Our cleanest insights came from isolated ablations (chunking alone vs finetuned alone).\n\n2. **Negative results are results**: Document and publish what doesn't work. Our template re-ranking negative results save others significant time.\n\n3. **Environment sniffing**: Auto-detect mount paths, Python versions, available packages rather than hardcoding.\n\n4. **Budget awareness**: Track GPU hours remaining, submissions remaining, time to deadline. Prioritize experiments by expected information gain per GPU-hour.\n\n5. **Persistent memory**: The AI agent must maintain structured logs across sessions. Our `MEMORY.md` + daily logs system was essential for continuity.\n\n### 7.3 Minimizing Human Reliance\n\nTo reduce the human bottleneck in future iterations:\n\n1. **Pre-validate locally**: Run a minimal subset (1-2 targets) in a local environment before pushing to Kaggle\n2. **Path normalization**: Implement robust path fallback logic (we added this, and it works: glob for datasets, checkpoint file search)\n3. **Argument validation**: Parse the Protenix `--help` output programmatically to validate argument names before pushing\n4. **Fork management**: Use the Kaggle API's fork endpoint (if available) rather than manual UI forking\n5. **Notification system**: Automated alerts when experiments complete or fail, with actionable next steps","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 9. Competition-Specific Analysis\n\n### 8.1 Test Target Characteristics\n\n| Target | Length | TBM Coverage (v65) | Notes |\n|--------|--------|--------------------| ------|\n| 9MME | 4640 nt | 4/5 TBM | Largest target, requires 11-18 chunks |\n| 9ZCC | 1460 nt | 2/5 TBM | Second largest, 4-5 chunks |\n| 9LEL | 476 nt | 2/5 TBM | Poor template coverage (50%, 5 breaks) |\n| 9LEC | 378 nt | 2/5 TBM | Medium length |\n| 9G4J | 334 nt | 5/5 TBM | Re-ranked to 9G4I (0.89 coverage) |\n| 9JFS | 246 nt | 1/5 TBM | Poor coverage (48%, 1 break) |\n| 9I9W | 28 nt | 5/5 TBM | Small, fully covered |\n| 9QZJ | 19 nt | 5/5 TBM | Smallest target |\n\n### 8.2 Score Ceiling Analysis\n\n- **Oracle TBM** (perfect template selection): 0.554 TM-score\n- **Our best** (v65): 0.432\n- **Gap**: 0.122 (22% of theoretical maximum)\n- **Real #1 competitor** (AyPy): 0.488\n- **Gap to close**: 0.056\n\nThe remaining gap is attributable to:\n1. Suboptimal template coordinate adaptation (~40% of gap)\n2. Long target chunking artifacts (~30% of gap)  \n3. Protenix prediction quality for non-TBM targets (~30% of gap)","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 10. Reproducibility\n\n### 9.1 Code Availability\n\n- **Public methodology notebook**: [kaggle.com/code/anismarrouchi/systematic-autoresearch-tbm-protenix-optimization](https://www.kaggle.com/code/anismarrouchi/systematic-autoresearch-tbm-protenix-optimization)\n- **Competition submission notebook**: `anismarrouchi/stanford-rna-3d-folding-2-baseline-v1` (v65 = best score)\n\n### 9.2 Key Dependencies\n\n| Package | Version | Source |\n|---------|---------|--------|\n| Protenix | v1 (base) | `qiweiyin/protenix-v1-adjusted` dataset |\n| Finetuned checkpoint | 1599_ema_0.999.pt | `zoushuxian/protenix-finetuned-rna3db-all-1599` |\n| Biopython | 1.86 (cp312) | `kami1976/biopython-cp312` |\n| USalign | compiled from source | `metric/usalign` |\n\n### 9.3 Critical Configuration (v65 — Best Score)\n\n```python\nMAX_SEQ_LEN       = 512\nCHUNK_OVERLAP     = 64\nMIN_SIMILARITY    = 0.0\nMIN_PERCENT_IDENTITY = 50.0\nN_SAMPLE          = 5\nSEED              = 42\nUSE_MSA           = False\nUSE_TEMPLATE      = False  # Protenix templates, not our TBM\nUSE_RNA_MSA       = True\n# Finetuned checkpoint: 1599_ema_0.999.pt\n```","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## 11. Conclusion\n\nOur journey through 74+ notebook versions demonstrates both the promise and current limitations of AI-assisted scientific research. The autoresearch paradigm successfully:\n\n- **Accelerated hypothesis generation** and experimental design\n- **Maintained rigorous documentation** across dozens of experiments\n- **Identified non-obvious insights** (TM-score scale invariance, template re-ranking trap)\n- **Performed systematic ablation studies** that isolated individual contributions\n\nBut it also revealed that **full autonomy is not yet achievable** in environments with complex infrastructure dependencies. The most productive mode was **human-in-the-loop**: AI-driven analysis and code generation with human execution at the Kaggle interface boundary.\n\nOur best contribution to the field is the comprehensive negative result on template selection optimization: despite strong cross-validation evidence of headroom, no practical method we tested could improve upon simple sequence identity ranking. This finding, replicated across 5 independent approaches (rule-based, ML scorer, consensus, coordinate averaging, re-ranking), should guide future work away from this attractive but unproductive direction.\n\n**بِسْمِ اللَّهِ الرَّحْمَنِ الرَّحِيمِ** — In the name of God, the Most Gracious, the Most Merciful. We dedicate this work to the advancement of computational biology in the MENA region.","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## Acknowledgments\n\nWe thank the Stanford Das Lab and Kaggle for organizing this competition, the Protenix team at ByteDance for their open-source implementation, and the Kaggle community (sigmaborov, nihilisticneuralnet, artemevstafyev, lightningv08) whose public notebooks provided invaluable reference points.","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## Appendix A: Complete Experiment Log\n\n| # | Version | Date | Hypothesis | Score | Delta | Decision |\n|---|---------|------|-----------|-------|-------|----------|\n| 1 | v5 | 2026-03-07 | A-form helix baseline | 0.087 | — | Baseline |\n| 2 | v12 | 2026-03-08 | TBM+Protenix PID≥50 | 0.378 | +0.291 | Keep |\n| 3 | v16 | 2026-03-08 | sim≥0.15 PID≥55 | 0.380 | +0.002 | Keep |\n| 4 | v19 | 2026-03-09 | Pure Protenix (no TBM) | 0.373 | -0.007 | Discard |\n| 5 | v21 | 2026-03-09 | Multi-seed Protenix | 0.276 | -0.104 | Discard |\n| 6 | v24 | 2026-03-09 | top_n=50 templates | 0.387 | +0.007 | Keep |\n| 7 | v26 | 2026-03-10 | Finetuned ckpt only | 0.347 | -0.040 | Discard |\n| 8 | v59 | 2026-03-13 | New baseline (Protenix+TBM ref) | 0.395 | — | New baseline |\n| 9 | v61 | 2026-03-13 | Post-processing 100Å guard | 0.404 | +0.009 | Keep |\n| 10 | v62 | 2026-03-13 | Rescale Protenix coords | FAIL | — | Format error |\n| 11 | v63 | 2026-03-13 | Fixed rescaling | 0.401 | -0.003 | Discard |\n| 12 | v65 | 2026-03-14 | **Chunking + finetuned ckpt** | **0.432** | **+0.028** | **BEST** |\n| 13 | v67 | 2026-03-14 | Chunking only (ablation) | 0.428 | +0.024 | Ablation |\n| 14 | v68 | 2026-03-14 | Rule-based template re-ranking | 0.410 | -0.018 | Discard |\n| 15 | v69 | 2026-03-15 | ML scorer (no finetuned) | 0.394 | -0.034 | Discard |\n| 16 | v70 | 2026-03-15 | ML scorer variant | 0.379 | -0.049 | Discard |\n| 17 | v71-73 | 2026-03-15 | N_step/N_cycle (wrong args) | FAIL | — | Config errors |\n| 18 | v74-77 | 2026-03-15 | N_step=250 + N_cycle=12 + stricter TBM (DRfold2 NOT mounted) | 0.362-0.364 | **-0.070** | **Discard** ❌ |\n| 19 | v79 | 2026-03-15 | v65 config + DRfold2 (but DRfold2 failed + re-ranking contamination) | 0.376 | -0.056 | Discard ❌ |\n| 20 | v80 | 2026-03-16 | 4-model hybrid (DRfold2 + Boltz2 both failed + re-ranking) | 0.386 | -0.046 | Discard ❌ |\n| 21 | v81 | 2026-03-16 | **Contamination fix**: reverted re-ranking, fixed DRfold2 glob, added mashumaro | Pending | — | Running |\n| 19 | Omar Exp2 | 2026-03-15 | Consensus pick | 0.417 | -0.015 | Discard |\n| 20 | Omar OL128 | 2026-03-14 | CHUNK_OVERLAP=128 | 0.416 | -0.016 | Discard |\n| 21 | Omar Exp5 | 2026-03-15 | CHUNK_OVERLAP=256 (isolated) | Pending | — | Running |\n| 22 | RNAPro | 2026-03-14 | NVIDIA RNA model | 0.386 | -0.046 | Discard |","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## Appendix B: Ablation Summary\n\n```\nBase (v59):                    0.395\n  + Post-processing guard:     0.404  (+0.009)\n  + Chunking:                  0.428  (+0.024)\n  + Finetuned checkpoint:      0.432  (+0.004)\n  ────────────────────────────────────\n  Total improvement:           +0.037\n\nNegative interventions:\n  + N_step=250 + N_cycle=12:   -0.070  (WORST — never override defaults)\n  + Template re-ranking:       -0.018\n  + ML scorer:                 -0.034 to -0.053\n  + Consensus picking:         -0.015\n  + Rescaling:                 -0.003\n  + CHUNK_OVERLAP=128:         -0.016 (vs v65, but without finetuned)\n```","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"## Appendix C: Infrastructure Failure Log\n\n| Attempt | Method | Error | Root Cause |\n|---------|--------|-------|-----------|\n| Exp2 v3-v6 | API push to Omar | Dataset path mismatch | API mounts at `/kaggle/input/datasets/<owner>/<slug>/` |\n| Exp3 v12-v14 | API push to Omar | Missing checkpoint | Different dataset version (v1 vs v2) |\n| v71 | API push to Anis | `unrecognized arguments: --num_trunk_recycles 12` | Wrong Protenix arg name |\n| v72 | API push to Anis | `unrecognized arguments: --num_trunk_recycles 12` | Same (v72 = v71 base) |\n| v73 | API push to Anis | `unrecognized arguments: --sample_diffusion.N_cycle 12` | Wrong namespace (model, not sample_diffusion) |\n\n**Resolution path**: v71 → v72 → v73 → v74 (each fixing the previous arg error, discovered empirically through Protenix's `--help` output in error logs).","outputs":[]},{"cell_type":"markdown","metadata":{},"source":"*Paper prepared by Razi, Bio-AI & Agritech Department, Noqta. March 2026.*\n*Competition ongoing — results will be updated upon final scoring.*","outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.0"}},"nbformat":4,"nbformat_minor":4}