{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":118765,"databundleVersionId":15231210}],"dockerImageVersionId":31286,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Stanford RNA 3D Folding Part2 Japanese Tutorial (日本語チュートリアル)\nこのノートブックでは、スタンフォード RNA 3D フォールディング競合データを調査し、RNA 3D 構造を予測するためのベースライン モデルを実装します。","metadata":{}},{"cell_type":"markdown","source":"## コンペティション概要\n**Stanford RNA 3D Folding Part 2** は、Stanford University をはじめとする研究機関が主催する Kaggle 上の機械学習コンペティションです。本コンペは RNA 分子の **一次配列（文字列情報）からその 3 次元構造（3D 立体座標）を予測すること** を競います。第一弾のチャレンジでは、自動モデルが人間の専門家レベルに到達するなど重要な進展を示しており、その続編としてさらに難易度の高い課題が出題されています。\n\nこのタスクは、RNA の構造が生物学的機能に深く関連しているという生体分子科学の基本原理にもとづき、**構造予測における計算的・機械学習的手法の性能向上** を目的としています。\n\n---","metadata":{}},{"cell_type":"markdown","source":"### 課題設定・目的\n1. **RNA 配列から 3D 構造を予測するモデル構築**  \n入力は RNA の塩基配列（A, C, G, U）。出力はその配列に対応する 3 次元構造（各原子または基準原子の座標）です。\n\n2. **未知構造の RNA 分子にも対応可能な汎化性能の獲得**  \n第一弾では類似構造が既知のデータが比較的容易な課題となりましたが、第二弾では テンプレート構造の存在しない RNA や新規構造など、より一般化が求められる設定になっています。\n\n3. **評価指標に基づくモデル評価**  \n予測結果は一般に立体構造の全体的な一致度や局所誤差に対して頑健な指標（例：TM-score など）により評価され、モデルの構造予測精度を測定します。\n\n---","metadata":{}},{"cell_type":"markdown","source":"### モチベーションとチャレンジ点\n- **生物学・医療へのインパクト**  \nRNA の立体構造はその機能に直結するため、精度の高い予測モデルは新薬設計・ワクチン開発・分子機構の理解に貢献します。実験的手法では時間・コストが高くつく 3D 構造決定を、計算モデルで補完できる可能性があります。\n\n- **機械学習と構造生物学の融合**  \nRNA 構造予測は、深層学習やテンプレートベース手法など最先端の機械学習技術の応用領域となっており、ここでの成功は他の複雑構造予測問題（例：タンパク質折り畳み）の進展にもつながることが期待されています。\n\n- **高次元予測問題**  \n一次配列から直接 3D 座標を推定することは、非線形かつ高次元な関数近似問題です。RNA 分子は部分的な二次構造（塩基対）や長距離相互作用など複雑な物理的制約を持つため、正確なモデリングが難しいという特性があります。\n\n- **一般化と未知構造への対応**  \n第一弾以上に、未知の構造やテンプレートが存在しないケースへの対応が求められるため、モデルの汎化性能 と ロバストな設計 が重要な課題となっています。\n\n- **データや評価の複雑性**  \nRNA の構造は単一の正解が存在しない場合もありうるため、TM-score などの構造類似性評価指標に基づく精度評価を適切に扱うことが求められます。\n---","metadata":{}},{"cell_type":"markdown","source":"## 準備","metadata":{}},{"cell_type":"code","source":"# ライブラリーインポート\n#　基本的ライブラリー\nimport numpy as np\nimport pandas as pd\nimport polars as pl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom IPython.display import display\nimport warnings\nwarnings.filterwarnings(\"ignore\", category=RuntimeWarning)\n\n# ディスプレイオプション\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', 100)\n\n# Set professional plotting style\nplt.style.use('ggplot')\nplt.rcParams['figure.facecolor'] = 'white'\nplt.rcParams['font.family'] = 'Arial'\ncustom_palette = [\"#3498db\", \"#e74c3c\", \"#2ecc71\", \"#f39c12\", \"#9b59b6\"]\nsns.set_palette(custom_palette)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:32:47.407287Z","iopub.execute_input":"2026-02-26T08:32:47.407751Z","iopub.status.idle":"2026-02-26T08:32:47.417354Z","shell.execute_reply.started":"2026-02-26T08:32:47.407681Z","shell.execute_reply":"2026-02-26T08:32:47.415806Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# データ読み込み\n# BASE_DIR = \"./\"    ##自分の環境のルートディレクトリを指定する\nBASE_DIR = \"/kaggle/input/competitions/stanford-rna-3d-folding-2\"\nOUTPUT_DIR = \"/kaggle/working\"\n# OUTPUT_DIR = BASE_DIR + \"/output\"\n\nDATA_DIR = BASE_DIR\n# DATA_DIR = \"/input/stanford-rna-3d-folding-2\"\ntrain_seq_df = pd.read_csv(DATA_DIR + \"/train_sequences.csv\")\ntrain_lbl_df = pd.read_csv(DATA_DIR + \"/train_labels.csv\")\nval_seq_df = pd.read_csv(DATA_DIR + \"/validation_sequences.csv\")\nval_lbl_df = pd.read_csv(DATA_DIR + \"/validation_labels.csv\")\ntest_seq_df = pd.read_csv(DATA_DIR + \"/test_sequences.csv\")\n\ndatasets_seq = {\n    \"train seq data\" : train_seq_df,\n    \"val seq data\" : val_seq_df,\n    \"test seq data\"  : test_seq_df,\n}\n\ndatasets_lbl = {\n    \"train label data\" : train_lbl_df,\n    \"val label data\" : val_lbl_df,\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:32:47.422026Z","iopub.execute_input":"2026-02-26T08:32:47.422949Z","iopub.status.idle":"2026-02-26T08:32:58.718973Z","shell.execute_reply.started":"2026-02-26T08:32:47.422878Z","shell.execute_reply":"2026-02-26T08:32:58.717771Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 初期データ観察","metadata":{}},{"cell_type":"code","source":"print(\"<shape>\")\nfor name, df in datasets_seq.items():\n    print(f\"{name}: {df.shape}\")\n\nfor name, df in datasets_lbl.items():\n    print(f\"{name}: {df.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:32:58.720533Z","iopub.execute_input":"2026-02-26T08:32:58.720944Z","iopub.status.idle":"2026-02-26T08:32:58.728753Z","shell.execute_reply.started":"2026-02-26T08:32:58.720884Z","shell.execute_reply":"2026-02-26T08:32:58.727328Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"<null count>\")\nfor name, df in datasets_seq.items():\n    print(f\"{name}: \")\n    display(df.isnull().sum())\nfor name, df in datasets_lbl.items():\n    print(f\"{name}: \")\n    display(df.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:32:58.73177Z","iopub.execute_input":"2026-02-26T08:32:58.732176Z","iopub.status.idle":"2026-02-26T08:33:00.111558Z","shell.execute_reply.started":"2026-02-26T08:32:58.732147Z","shell.execute_reply":"2026-02-26T08:33:00.110523Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 提供データの構成\n- **[train/validation/test]_sequences.csv** : RNA分子のターゲット配列\n- **[train/validation]_labels.csv** : 実験構造ラベルデータ\n- **sample_submission.csv** : 提出用のcsvフォーマット\n- **MSA/** : \n- **PDB_RNA/** : \n- **extra/** : \n","metadata":{}},{"cell_type":"code","source":"# 生データ観察\nprint(\"<raw data (sequence)>\")\nfor name, df in datasets_seq.items():\n    print(f\"{name}: \")\n    display(df.sample(5))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.112783Z","iopub.execute_input":"2026-02-26T08:33:00.113057Z","iopub.status.idle":"2026-02-26T08:33:00.153949Z","shell.execute_reply.started":"2026-02-26T08:33:00.113033Z","shell.execute_reply":"2026-02-26T08:33:00.149966Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### RNA分子のターゲット配列データ([train/validation/test]_sequences.csv)\n- *target_id* : 任意の識別子\n- *sequence* : ターゲット内の全てのRNA鎖配列\n- *temmporal_cutoff* : 鎖配列が公開された、または公開される予定の日付(yyyy-mm-dd)\n- *description* : 鎖配列の起源に関する詳細。PDBエントリーの場合はエントリータイトル\n- *stoichiometry* : 科学量論ターゲットに使用されるチェーン。\"chain : number\"作成者定義チェーンとall_sequencesで対応している\n- *all_sequences* : 実験的に解読された構造に含まれるすべてのFASTA形式分子鎖配列\n- *ligand_ids* : 実験構造で解読された任意の小分子リガンドのPDB科学成分辞書に登録された3文字の名前\n- *ligand_SMILES* : 実験構造で解読された任意の小分子リガンドの科学構造を示すSMILES文字列","metadata":{}},{"cell_type":"code","source":"# 生データ観察\nprint(\"<raw data label>\")\nfor name, df in datasets_lbl.items():\n    print(f\"{name}: \")\n    display(df.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.156991Z","iopub.execute_input":"2026-02-26T08:33:00.157993Z","iopub.status.idle":"2026-02-26T08:33:00.346946Z","shell.execute_reply.started":"2026-02-26T08:33:00.157952Z","shell.execute_reply":"2026-02-26T08:33:00.345882Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### RNA分子のターゲット配列データ([train/validation/test]_labels.csv)\n- *ID* : target_idで区切られた残基番号\n- *resname* : 残基のRNAヌクレオチド(A, C, G, U)\n- *resid* : 残基番号\n- *x_1,,,,* : 角実験RNA構造のC1'原子の座標(オングストローム単位)\n- *chain* : 残基の鎖ID\n- *copy* : 残基がどの鎖コピーに含まれているか","metadata":{}},{"cell_type":"markdown","source":"### RNA構造の理解","metadata":{}},{"cell_type":"markdown","source":"**RNA構造階層:**\n1. **一次構造:** ヌクレオチド配列(A, C, G, U)\n2. **二次構造**: 塩基対号パターン(ステム、ループ、バルジ)\n3. **三次構造**: 空間における3次元配置 ←今回の予測\n\n**C1'原子:**\n- 各ヌクレオチドのC1'原子の位置を予測する\n- C1'はリボース糖を塩基に結びつける炭素原子である\n- これは3D空間におけるヌクレオチドの位置をよく表している","metadata":{}},{"cell_type":"markdown","source":"**TM-score Metric:**\n$$TM\\text{-}score = \\max\\left[\\frac{1}{L_{ref}} \\sum_{i=1}^{L_{align}} \\frac{1}{1 + (d_i/d_0)^2}\\right]$$\n\n- $L_{ref}$: 参照構造中の残基数\n- $d_i$: 整列した残基ペア間の距離\n- $d_0$: 配列の長さに基づく正規化係数","metadata":{}},{"cell_type":"markdown","source":"- TM-score > 0.5: 位相的に相似\n- TM-score > 0.17: ランダムな構造類似はない<br>\n<br>\n**目標はTM-scoreを1.0に最大化すること**","metadata":{}},{"cell_type":"markdown","source":"## データ考察と戦略立案","metadata":{}},{"cell_type":"markdown","source":"**テンプレートベースのアプローチを試してみる**\nこのコンペティションのパート1ではテンプレートベースの手法がdenovo構造予測よりも優れた結果を示しました。とりあえずで試すアプローチとしては最適です。\n\n**戦略の詳細:**\n1. 既知の3D構造を持つPDB内の相同配列を見つける\n2. ターゲット配列をテンプレート配列に合わせる\n3. アライメントに基づいてテンプレートからターゲットに座標を転送する\n4. 構造を改良して衝突を取り除き、ジオメトリを最適化する\n\n**実装手順**\n1. MSAファイルを解析して関連シーケンスを検索する\n2. 関連する配列の構造をPDBで検索\n3. 配列構造アライメントを実行する\n4. 座標の抽出と変換","metadata":{}},{"cell_type":"markdown","source":"## モデル構築と予測","metadata":{}},{"cell_type":"code","source":"# ENVIRONMENT SETUP - MEMORY OPTIMIZED FOR KAGGLE\nimport os\nimport sys\nimport warnings\nimport gc\nfrom pathlib import Path\nfrom typing import Dict, List, Tuple, Optional\nfrom collections import defaultdict\nfrom tqdm.auto import tqdm\nimport time\nimport plotly.graph_objects as go\n\n# Visualization (import only when needed)\nimport matplotlib\nmatplotlib.use('Agg')  # Non-interactive backend to save memory\nimport matplotlib.pyplot as plt\n\n# Suppress warnings for cleaner output\nwarnings.filterwarnings('ignore')\n\n# Set random seeds for reproducibility\nSEED = 42\nnp.random.seed(SEED)\n\n# MEMORY MANAGEMENT UTILITIES\ndef get_memory_usage():\n    \"\"\"Get current memory usage in MB.\"\"\"\n    import psutil\n    process = psutil.Process(os.getpid())\n    return process.memory_info().rss / 1024 / 1024\n\ndef clear_memory():\n    \"\"\"Aggressively clear memory.\"\"\"\n    gc.collect()\n\n\ndef optimize_dataframe(df: pd.DataFrame) -> pd.DataFrame:\n    \"\"\"Reduce DataFrame memory usage by optimizing dtypes.\"\"\"\n    for col in df.columns:\n        col_type = df[col].dtype\n        if col_type == 'float64':\n            df[col] = df[col].astype('float32')\n        elif col_type == 'int64':\n            df[col] = df[col].astype('int32')\n    return df\n\nprint(\" Libraries imported \")\n\n# Check if running on Kaggle\nIS_KAGGLE = os.path.exists('/kaggle/input')\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.348503Z","iopub.execute_input":"2026-02-26T08:33:00.349871Z","iopub.status.idle":"2026-02-26T08:33:00.363745Z","shell.execute_reply.started":"2026-02-26T08:33:00.349833Z","shell.execute_reply":"2026-02-26T08:33:00.36244Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# HELPER FUNCTIONS FOR FASTA PARSING\ndef parse_fasta(fasta_string: str) -> Dict[str, str]:\n    \"\"\"\n    Parse FASTA-formatted string into a dictionary of chain:sequence.\n    This is based on the competition's extra/parse_fasta_py.py helper.\n    \n    Args:\n        fasta_string: FASTA-formatted string\n        \n    Returns:\n        Dictionary mapping chain IDs to sequences\n    \"\"\"\n    sequences = {}\n    current_header = None\n    current_sequence = []\n    \n    for line in fasta_string.strip().split('\\n'):\n        if line.startswith('>'):\n            # Save previous sequence\n            if current_header and current_sequence:\n                # Extract chain ID from header\n                chain_id = current_header.split('|')[0].split('_')[-1] if '|' in current_header else current_header\n                sequences[chain_id] = ''.join(current_sequence)\n            \n            current_header = line[1:].strip()\n            current_sequence = []\n        else:\n            current_sequence.append(line.strip())\n    \n    # Save last sequence\n    if current_header and current_sequence:\n        chain_id = current_header.split('|')[0].split('_')[-1] if '|' in current_header else current_header\n        sequences[chain_id] = ''.join(current_sequence)\n    \n    return sequences\n\n\ndef parse_msa_file(msa_path: Path) -> Dict[str, List[str]]:\n    \"\"\"\n    Parse MSA FASTA file and return sequences grouped by chain.\n    \n    Args:\n        msa_path: Path to MSA FASTA file\n        \n    Returns:\n        Dictionary with chain IDs as keys and list of aligned sequences as values\n    \"\"\"\n    chain_sequences = defaultdict(list)\n    \n    if not msa_path.exists():\n        return chain_sequences\n    \n    with open(msa_path, 'r') as f:\n        current_header = None\n        current_sequence = []\n        \n        for line in f:\n            line = line.strip()\n            if line.startswith('>'):\n                if current_header and current_sequence:\n                    # Parse chain from header\n                    chain = 'A'  # default\n                    for part in current_header.split('|'):\n                        if part.startswith('chain='):\n                            chain = part.split('=')[1]\n                            break\n                    chain_sequences[chain].append(''.join(current_sequence))\n                \n                current_header = line[1:]\n                current_sequence = []\n            else:\n                current_sequence.append(line)\n        \n        # Last sequence\n        if current_header and current_sequence:\n            chain = 'A'\n            for part in current_header.split('|'):\n                if part.startswith('chain='):\n                    chain = part.split('=')[1]\n                    break\n            chain_sequences[chain].append(''.join(current_sequence))\n    \n    return dict(chain_sequences)\n\n\nprint(\"FASTA parsing functions defined\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.36516Z","iopub.execute_input":"2026-02-26T08:33:00.366427Z","iopub.status.idle":"2026-02-26T08:33:00.404299Z","shell.execute_reply.started":"2026-02-26T08:33:00.366384Z","shell.execute_reply":"2026-02-26T08:33:00.402681Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# MEMORY-EFFICIENT TEMPLATE DATABASE\nclass LightweightTemplateDB:\n    \"\"\"\n    Memory-efficient template database.\n    \n    Grandmaster Note: Instead of loading all coordinates into memory,\n    we only store sequences and load coordinates on-demand from disk.\n    \"\"\"\n    \n    def __init__(self, train_sequences: pd.DataFrame, labels_df: pd.DataFrame,\n                 max_templates: int = 500):\n        \"\"\"\n        Initialize lightweight template database.\n        \n        Args:\n            train_sequences: DataFrame with sequences (small)\n            labels_path: Path to labels CSV (loaded on demand)\n            max_templates: Maximum templates to keep (memory limit)\n        \"\"\"\n        self.labels_df = labels_df\n        \n        # Only keep a subset of templates (sorted by sequence length diversity)\n        train_sequences = train_sequences.copy()\n        train_sequences['seq_len'] = train_sequences['sequence'].str.len()\n        \n        # Sample diverse lengths\n        train_sequences = train_sequences.sort_values('seq_len')\n        step = max(1, len(train_sequences) // max_templates)\n        train_sequences = train_sequences.iloc[::step].head(max_templates)\n        \n        # Store only sequence index (minimal memory)\n        self.sequence_index = {}\n        for _, row in train_sequences.iterrows():\n            self.sequence_index[row['target_id']] = row['sequence']\n        \n        del train_sequences\n        clear_memory()\n        \n        print(f\"Lightweight template DB: {len(self.sequence_index)} templates\")\n    \n    def find_templates_fast(self, query_sequence: str, top_k: int = 3) -> List[Tuple[str, float]]:\n        \"\"\"\n        Fast template search using simple sequence similarity.\n        \n        Uses k-mer matching for speed instead of full alignment.\n        \"\"\"\n        query_len = len(query_sequence)\n        scores = []\n        \n        # Simple length-based filtering first\n        for target_id, template_seq in self.sequence_index.items():\n            template_len = len(template_seq)\n            \n            # Skip if length difference is too large\n            len_ratio = min(query_len, template_len) / max(query_len, template_len)\n            if len_ratio < 0.5:\n                continue\n            \n            # Fast k-mer similarity\n            k = 4\n            query_kmers = set(query_sequence[i:i+k] for i in range(len(query_sequence)-k+1))\n            template_kmers = set(template_seq[i:i+k] for i in range(len(template_seq)-k+1))\n            \n            if len(query_kmers) == 0 or len(template_kmers) == 0:\n                continue\n            \n            # Jaccard similarity\n            intersection = len(query_kmers & template_kmers)\n            union = len(query_kmers | template_kmers)\n            similarity = intersection / union if union > 0 else 0\n            \n            if similarity > 0.1:\n                scores.append((target_id, similarity))\n        \n        scores.sort(key=lambda x: x[1], reverse=True)\n        return scores[:top_k]\n    \n    def get_template_coordinates_chunked(self, template_id: str) -> Optional[np.ndarray]:\n        \"\"\"\n        Load coordinates for a single template from disk (memory efficient).\n        \"\"\"\n\n        try:\n            # Read in chunks to find the target\n            for chunk in self.labels_df:\n                chunk['target_id'] = chunk['ID'].apply(lambda x: '_'.join(x.split('_')[:-1]))\n                target_data = chunk[chunk['target_id'] == template_id]\n                \n                if len(target_data) > 0:\n                    if 'x_1' in target_data.columns:\n                        coords = target_data[['x_1', 'y_1', 'z_1']].values.astype(np.float32)\n                        del chunk, target_data\n                        clear_memory()\n                        return coords\n                \n                del chunk\n            \n        except Exception as e:\n            print(f\"   Warning: Could not load template {template_id}: {e}\")\n        \n        return None\n\n\nprint(\"LightweightTemplateDB class defined\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.405628Z","iopub.execute_input":"2026-02-26T08:33:00.406012Z","iopub.status.idle":"2026-02-26T08:33:00.447215Z","shell.execute_reply.started":"2026-02-26T08:33:00.405972Z","shell.execute_reply":"2026-02-26T08:33:00.445503Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 配列アライメントと構造転送","metadata":{}},{"cell_type":"markdown","source":"**ギャップの処理:**\n- テンプレートのギャップ: テンプレートを持たない残基の座標を「構築」する必要がある\n- クエリのギャップ: テンプレートの位置をスキップする\n- ギャップ領域は隣接する残基から座標を補完する","metadata":{}},{"cell_type":"code","source":"# MEMORY-EFFICIENT STRUCTURE PREDICTOR\nclass LightweightPredictor:\n    \"\"\"\n    Memory-efficient RNA structure predictor.\n    \n    Key optimizations:\n    - No large matrices in memory\n    - Simple coordinate generation\n    - Immediate garbage collection\n    \"\"\"\n    \n    def __init__(self, template_db: Optional['LightweightTemplateDB'] = None):\n        self.template_db = template_db\n        self.standard_c1_distance = 5.9  # Angstroms\n    \n    def predict_structure(self, sequence: str, target_id: str) -> Dict[int, np.ndarray]:\n        \"\"\"\n        Generate 5 structure predictions for a sequence.\n        \n        Returns dict mapping model number (1-5) to coordinates array.\n        \"\"\"\n        seq_len = len(sequence)\n        predictions = {}\n        \n        # Try template-based first\n        template_coords = None\n        if self.template_db is not None:\n            templates = self.template_db.find_templates_fast(sequence, top_k=1)\n            if templates:\n                template_id, score = templates[0]\n                if score > 0.15:\n                    template_coords = self.template_db.get_template_coordinates_chunked(template_id)\n        \n        # Generate 5 diverse predictions\n        for model_num in range(1, 6):\n            if template_coords is not None and len(template_coords) > 0:\n                # Transfer and adapt template coordinates\n                coords = self._adapt_template(template_coords, seq_len)\n                # Add increasing noise for diversity\n                noise_level = 0.2 * (model_num - 1)\n                coords = coords + np.random.randn(seq_len, 3).astype(np.float32) * noise_level\n            else:\n                # Generate de novo structure\n                coords = self._generate_helix(seq_len, variation=model_num)\n            \n            predictions[model_num] = coords\n        \n        return predictions\n    \n    def _adapt_template(self, template_coords: np.ndarray, target_len: int) -> np.ndarray:\n        \"\"\"Adapt template coordinates to target length.\"\"\"\n        template_len = len(template_coords)\n        \n        if template_len == target_len:\n            return template_coords.copy()\n        \n        # Interpolate or truncate\n        if template_len > target_len:\n            # Truncate\n            return template_coords[:target_len].copy()\n        else:\n            # Extend with helix\n            coords = np.zeros((target_len, 3), dtype=np.float32)\n            coords[:template_len] = template_coords\n            \n            # Extend remaining residues\n            if template_len > 1:\n                direction = template_coords[-1] - template_coords[-2]\n                direction = direction / (np.linalg.norm(direction) + 1e-6)\n            else:\n                direction = np.array([1, 0, 0], dtype=np.float32)\n            \n            for i in range(template_len, target_len):\n                coords[i] = coords[i-1] + direction * self.standard_c1_distance\n                # Add slight curve\n                direction = direction + np.random.randn(3).astype(np.float32) * 0.1\n                direction = direction / (np.linalg.norm(direction) + 1e-6)\n            \n            return coords\n    \n    def _generate_helix(self, seq_len: int, variation: int = 1) -> np.ndarray:\n        \"\"\"Generate idealized A-form RNA helix.\"\"\"\n        coords = np.zeros((seq_len, 3), dtype=np.float32)\n        \n        # A-form helix parameters with variation\n        rise = 2.81 + (variation - 3) * 0.1\n        rotation = 32.7 + (variation - 3) * 2\n        radius = 10.0 + (variation - 3) * 0.5\n        \n        for i in range(seq_len):\n            angle = np.radians(rotation * i)\n            coords[i, 0] = radius * np.cos(angle)\n            coords[i, 1] = radius * np.sin(angle)\n            coords[i, 2] = rise * i\n        \n        return coords\n\n\nprint(\"LightweightPredictor class defined\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.450949Z","iopub.execute_input":"2026-02-26T08:33:00.451744Z","iopub.status.idle":"2026-02-26T08:33:00.481686Z","shell.execute_reply.started":"2026-02-26T08:33:00.451695Z","shell.execute_reply":"2026-02-26T08:33:00.480498Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# MEMORY-EFFICIENT PIPELINE\nclass MemoryEfficientPipeline:\n    \"\"\"\n    Memory-efficient RNA prediction pipeline for Kaggle.\n    \n    Key features:\n    - Streaming predictions (don't accumulate in memory)\n    - Write directly to CSV in chunks\n    - Aggressive garbage collection\n    \"\"\"\n    \n    def __init__(self, train_seq: pd.DataFrame = None, labels_path: Path = None):\n        \"\"\"Initialize with minimal memory footprint.\"\"\"\n        self.template_db = None\n        \n        if train_seq is not None and labels_path is not None:\n            self.template_db = LightweightTemplateDB(\n                train_seq, labels_path, max_templates=300\n            )\n        \n        self.predictor = LightweightPredictor(self.template_db)\n        clear_memory()\n        print(\"Memory-efficient pipeline initialized\")\n    \n    def generate_submission_streaming(self, test_sequences: pd.DataFrame, \n                                       output_path: Path) -> None:\n        \"\"\"\n        Generate submission by streaming directly to file.\n        \n        This avoids accumulating all predictions in memory.\n        \"\"\"\n        print(\"Generating submission (streaming mode)...\")\n        start_time = time.time()\n        \n        # Prepare header\n        coord_cols = []\n        for i in range(1, 6):\n            coord_cols.extend([f'x_{i}', f'y_{i}', f'z_{i}'])\n        header = ['ID', 'resname', 'resid'] + coord_cols\n        \n        # Write header\n        with open(output_path, 'w') as f:\n            f.write(','.join(header) + '\\n')\n        \n        n_targets = len(test_sequences)\n        batch_size = 10  # Process and write in small batches\n        \n        for batch_start in tqdm(range(0, n_targets, batch_size), desc=\"Processing\"):\n            batch_end = min(batch_start + batch_size, n_targets)\n            batch_rows = []\n            \n            for idx in range(batch_start, batch_end):\n                row = test_sequences.iloc[idx]\n                target_id = row['target_id']\n                sequence = row['sequence']\n                \n                # Generate predictions\n                predictions = self.predictor.predict_structure(sequence, target_id)\n                \n                # Convert to rows\n                for resid, resname in enumerate(sequence, 1):\n                    row_data = [f\"{target_id}_{resid}\", resname, str(resid)]\n                    \n                    for model_num in range(1, 6):\n                        coords = predictions.get(model_num)\n                        if coords is not None and resid - 1 < len(coords):\n                            x, y, z = coords[resid - 1]\n                        else:\n                            x, y, z = 0.0, 0.0, 0.0\n                        \n                        # Clip and format\n                        x = max(-999.999, min(9999.999, float(x)))\n                        y = max(-999.999, min(9999.999, float(y)))\n                        z = max(-999.999, min(9999.999, float(z)))\n                        \n                        row_data.extend([f\"{x:.3f}\", f\"{y:.3f}\", f\"{z:.3f}\"])\n                    \n                    batch_rows.append(','.join(row_data))\n                \n                # Clear predictions from memory\n                del predictions\n            \n            # Write batch to file\n            with open(output_path, 'a') as f:\n                f.write('\\n'.join(batch_rows) + '\\n')\n            \n            del batch_rows\n            clear_memory()\n        \n        elapsed = time.time() - start_time\n        file_size = os.path.getsize(output_path) / 1024 / 1024\n        \n        print(f\"\\nSubmission complete!\")\n        print(f\"   File: {output_path}\")\n        print(f\"   Size: {file_size:.2f} MB\")\n        print(f\"   Time: {elapsed/60:.1f} minutes\")\n\n\nprint(\" MemoryEfficientPipeline class defined\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.483465Z","iopub.execute_input":"2026-02-26T08:33:00.484165Z","iopub.status.idle":"2026-02-26T08:33:00.515689Z","shell.execute_reply.started":"2026-02-26T08:33:00.484116Z","shell.execute_reply":"2026-02-26T08:33:00.514275Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### アンサンブル","metadata":{}},{"cell_type":"markdown","source":"**多様性のための戦略**\n1. 複数のテンプレート: 異なるテンプレート構造を使用する\n2. ノイズ注入: 座標に制御されたガウスノイズを追加する\n3. 異なるアライメントパラメータ: ギャップペナルティを変更する\n4. 構造変異: 柔軟な領域の異なるコンフォメーションのサンプル\n\n**ベストプラクティス**\n- モデル1: 摂動が最小限の最適なテンプレート\n- モデル2: 2番目/3番目にいいテンプレート \n- モデル3: 探査用のノイズを増やしたトップテンプレート","metadata":{}},{"cell_type":"code","source":"# 未完成","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.517455Z","iopub.execute_input":"2026-02-26T08:33:00.517929Z","iopub.status.idle":"2026-02-26T08:33:00.54757Z","shell.execute_reply.started":"2026-02-26T08:33:00.517889Z","shell.execute_reply":"2026-02-26T08:33:00.545942Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 提出物の生成","metadata":{}},{"cell_type":"markdown","source":"| Column | Description |\n|--------|-------------|\n| ID | `target_id_resid` のフォーマット(e.g., \"R1107_1\") |\n| resname | ヌクレオチド (A, C, G, U) |\n| resid | 残基番号 |\n| x_1, y_1, z_1 | 予測座標 1 |\n| ... | |\n| x_5, y_5, z_5 | 予測座標 5 |","metadata":{}},{"cell_type":"markdown","source":"**注意点**\n- 座標は-999.999から9999.999である必要がある\n- 残基ごとに正確に5つの座標セットを提出する必要がある","metadata":{}},{"cell_type":"markdown","source":"### パイプラインの実行","metadata":{}},{"cell_type":"code","source":"train_lbl_df = pd.read_csv(DATA_DIR + \"/train_labels.csv\", chunksize = 10000)\n\npipeline = MemoryEfficientPipeline(train_seq_df, train_lbl_df)\nclear_memory()\n\nprint(\"Pipeline ready for prediction\")\n\n# ============================================================================\n# FINAL PRODUCTION PIPELINE - MEMORY OPTIMIZED\n# ============================================================================\n\ndef run_final_submission_optimized():\n    \"\"\"\n    Memory-optimized submission pipeline for Kaggle.\n    \"\"\"\n    print(\" FINAL SUBMISSION - MEMORY OPTIMIZED\")\n    print(\"=\" * 70)\n    \n    clear_memory()\n    \n    # =========================================================================\n    # STEP 1: Load test sequences only\n    # =========================================================================\n    print(\"\\n Step 1: Loading test sequences...\")\n    \n    test_path = DATA_DIR + \"/test_sequences.csv\"\n\n    if test_path:\n        test_df = pd.read_csv(test_path)\n        print(f\"    Loaded {len(test_df)} test sequences\")\n    else:\n        print(\"   test_sequences.csv not found, using validation set\")\n        test_df = val_seq_df\n        if test_df is None:\n            raise FileNotFoundError(\"No test or validation sequences found!\")\n    \n    # =========================================================================\n    # STEP 2: Generate submission using streaming\n    # =========================================================================\n    print(\"\\nStep 2: Generating predictions (streaming)...\")\n    \n    output_path = Path(OUTPUT_DIR + \"/submission.csv\") \n    \n    pipeline.generate_submission_streaming(test_df, output_path)\n    \n    # =========================================================================\n    # STEP 3: Validate output\n    # =========================================================================\n    print(\"\\n Step 3: Validating submission...\")\n    \n    # Quick validation (read only first few rows)\n    sample = pd.read_csv(output_path, nrows=10)\n    print(f\"   Columns: {sample.columns.tolist()}\")\n    print(f\"   Sample rows: {len(sample)}\")\n    \n    # Count total rows\n    with open(output_path, 'r') as f:\n        total_rows = sum(1 for _ in f) - 1  # Subtract header\n    \n    print(f\"\\n SUBMISSION SUMMARY:\")\n    print(f\"   Total rows: {total_rows:,}\")\n    print(f\"   File: {output_path}\")\n    \n    del sample, test_df\n    clear_memory()\n    \n    print(\"\\n Submission complete!\")\n\n# Run the optimized pipeline\nrun_final_submission_optimized()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:00.549252Z","iopub.execute_input":"2026-02-26T08:33:00.549569Z","iopub.status.idle":"2026-02-26T08:33:16.53111Z","shell.execute_reply.started":"2026-02-26T08:33:00.549518Z","shell.execute_reply":"2026-02-26T08:33:16.529953Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 可視化","metadata":{}},{"cell_type":"code","source":"import plotly.io as pio\npio.renderers.default = \"iframe\"\n\ndef load_and_prep_data(filepath= OUTPUT_DIR + '/submission.csv'):\n    \"\"\"CSVを読み込み、カラム名を正規化し、ターゲットID列を作成する\"\"\"\n    try:\n        df = pd.read_csv(filepath)\n        print(f\"読み込み完了: {df.shape}\")\n    except FileNotFoundError:\n        print(f\"エラー: {filepath} が見つかりません。\")\n        return None\n\n    # カラム名の正規化（小文字化、空白削除）\n    df.columns = df.columns.str.strip().str.lower()\n    \n    # 'id'カラムの特定\n    id_col = 'id' if 'id' in df.columns else None\n    if not id_col:\n        # 'id'が含まれるカラムを探す\n        candidates = [c for c in df.columns if 'id' in c]\n        if candidates:\n            id_col = candidates[0]\n            df.rename(columns={id_col: 'id'}, inplace=True)\n        else:\n            print(\"エラー: 'id' カラムが見つかりません。\")\n            return None\n\n    # Target IDの抽出 (例: \"8ZNQ_1\" -> \"8ZNQ\")\n    # 最後のアンダースコアより前をターゲット名とする\n    df['target_name'] = df['id'].apply(lambda x: '_'.join(str(x).split('_')[:-1]))\n    \n    # 残基番号の抽出 (例: \"8ZNQ_1\" -> 1)\n    # ソートのために数値化しておく\n    try:\n        df['residue_number'] = df['id'].apply(lambda x: int(str(x).split('_')[-1]))\n    except:\n        df['residue_number'] = df.index # 失敗した場合は行番号\n\n    return df\n\ndef plot_protein_structure(df, target_name, model_num=1):\n    \"\"\"\n    指定されたターゲットとモデル番号の構造を3Dプロットする\n    \"\"\"\n    # ターゲットでフィルタリング\n    target_df = df[df['target_name'] == target_name].copy()\n    \n    if target_df.empty:\n        print(f\"警告: ターゲット '{target_name}' のデータが見つかりません。\")\n        return\n\n    # 残基順にソート（線がきれいに繋がるように）\n    target_df = target_df.sort_values('residue_number')\n\n    # 座標カラムの特定 (x_1, y_1, z_1 形式に対応)\n    x_col = f'x_{model_num}'\n    y_col = f'y_{model_num}'\n    z_col = f'z_{model_num}'\n\n    # カラムが存在するかチェック (x, y, z のみのケースも考慮)\n    if x_col not in target_df.columns:\n        if 'x' in target_df.columns:\n            x_col, y_col, z_col = 'x', 'y', 'z'\n        else:\n            print(f\"エラー: 座標カラム ({x_col}) が見つかりません。\")\n            return\n\n    # --- Plotlyによる描画 ---\n    fig = go.Figure()\n\n    fig.add_trace(go.Scatter3d(\n        x=target_df[x_col],\n        y=target_df[y_col],\n        z=target_df[z_col],\n        mode='lines+markers',\n        marker=dict(\n            size=4,\n            color=target_df['residue_number'], # 残基順に色を変える\n            colorscale='Viridis',\n            showscale=True,\n            colorbar=dict(title=\"Residue Index\")\n        ),\n        line=dict(\n            color='gray',\n            width=2\n        ),\n        text=[f\"Residue: {r}<br>AA: {n}\" for r, n in zip(target_df['residue_number'], target_df.get('resname', ['?']*len(target_df)))],\n        hoverinfo='text'\n    ))\n\n    fig.update_layout(\n        title=f\"Protein Structure: {target_name} (Model {model_num})\",\n        scene=dict(\n            xaxis_title=\"X\",\n            yaxis_title=\"Y\",\n            zaxis_title=\"Z\",\n            aspectmode='data' # アスペクト比をデータに合わせる（歪み防止）\n        ),\n        margin=dict(l=0, r=0, b=0, t=40),\n        height=700\n    )\n\n    fig.show()\n\n# ==========================================\n# 実行セクション\n# ==========================================\n\n# 1. データの読み込み\ndf = load_and_prep_data(OUTPUT_DIR + '/submission.csv')\n\nif df is not None:\n    # 2. 利用可能なターゲットの一覧を表示\n    unique_targets = df['target_name'].unique()\n    print(f\"\\n検出されたターゲット数: {len(unique_targets)}\")\n    print(f\"ターゲットID一覧 (最初の5件): {unique_targets[:5]}\")\n\n    # 3. 可視化の実行\n    # ここで表示したいターゲットIDを指定してください\n    if len(unique_targets) > 0:\n        target_to_plot = unique_targets[0] # 最初のターゲットを自動選択\n        \n        print(f\"\\n--- {target_to_plot} の Model 1 を描画します ---\")\n        plot_protein_structure(df, target_to_plot, model_num=1)\n        \n        # 必要であれば他のモデルも表示可能\n        # plot_protein_structure(df, target_to_plot, model_num=2)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-26T08:33:16.532738Z","iopub.execute_input":"2026-02-26T08:33:16.533171Z","iopub.status.idle":"2026-02-26T08:33:16.62612Z","shell.execute_reply.started":"2026-02-26T08:33:16.53313Z","shell.execute_reply":"2026-02-26T08:33:16.624645Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 再考察","metadata":{}},{"cell_type":"markdown","source":"1. テンプレートの選択は重要\n- PDBから最適なテンプレートを見つける\n- MSA情報を使用してホモログを識別する\n- 配列同一性が高いほど、構造の転移が優れている\n2. 5つのモデルの多様性\n- ランダムノイズを追加するだけではダメ\n- モデルごとに異なるテンプレートを使用する\n- 代替のコンフォメーションを検討する\n3. エッジケースの処理\n- 短い配列（< 30 nt）には特別なd0値が必要です\n- マルチチェーンターゲットには適切な処理が必要\n- 欠落残基は補間する必要がある\n4. 計算効率\n- 事前計算テンプレートデータベース\n- ベクトル化された演算を使用する\n- メモリ使用量を監視する\n5. 検証戦略\n- ローカルテストに検証セットを使用する\n- TMスコアの分布を追跡する\n- 弱点を特定する（アイデンティティの低いターゲット）","metadata":{}},{"cell_type":"markdown","source":"## コメント","metadata":{}},{"cell_type":"markdown","source":"### 最後に\nここまで読んでいただきありがとうございました。私はデータ分析の学習のためにkaggleのコンペティションに参加しています。何かアドバイスや疑問点があればお気軽にコメントしてください。日本語でも英語でもどちらでも対応しています。","metadata":{}},{"cell_type":"markdown","source":"参考\n\nStanford RNA 3D Folding Part2: \n[https://www.kaggle.com/code/adilshamim8/stanford-rna-3d-folding-part-2]\n\nStanford RNA Folding 2: Template-based Approach: \n[https://www.kaggle.com/code/nihilisticneuralnet/stanford-rna-folding-2-template-based-approach]\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}