{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.8.10"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":118765,"databundleVersionId":15231210,"sourceType":"competition"}],"dockerImageVersionId":31239,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"6fb2e63d-2e1b-4dc9-a8cf-3e6a2e3c963c","cell_type":"markdown","source":"# 🧬 Stanford RNA 3D Folding 2: Expanded EDA & Visualization 📊\n\nWelcome to this Exploratory Data Analysis (EDA) notebook! 🚀\n\nIn this notebook, we will:\n1.  **Analyze Sequence Lengths**: Understand the distribution of RNA lengths. 📏\n2.  **Check Character Composition**: See the balance of A, U, G, C bases. 🧬\n3.  **Ligand Analysis**: See what molecules bind to our RNAs. 💊\n4.  **Stoichiometry**: Analyze single vs. multi-chain complexes. 🧩\n5.  **WordCloud**: Visualize common terms in RNA descriptions. ☁️\n6.  **Visualize 3D Structures**: Plot the 3D coordinates. 🧊\n\nLet's dive in! 🌊","metadata":{}},{"id":"9ed9b56a-e800-41e9-81f1-7ec634a04605","cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nimport os\nfrom collections import Counter\nfrom wordcloud import WordCloud\n\n# 🎨 Set plot style for better aesthetics\nsns.set_theme(style=\"whitegrid\")\n\n# 📁 Define the base directory for Kaggle environment\nBASE_DIR = '/kaggle/input/stanford-rna-3d-folding-2'\n\nprint(\"✅ Setup complete!\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:52.624176Z","iopub.execute_input":"2026-01-08T17:07:52.624938Z","iopub.status.idle":"2026-01-08T17:07:52.632739Z","shell.execute_reply.started":"2026-01-08T17:07:52.624909Z","shell.execute_reply":"2026-01-08T17:07:52.631139Z"}},"outputs":[],"execution_count":null},{"id":"9480ab39-983b-4df0-8f72-e20c72c50397","cell_type":"markdown","source":"## 1. Loading Data 📜\nWe load the `train_sequences.csv`.","metadata":{}},{"id":"4444111c-4f82-4f50-996a-696ee520a369","cell_type":"code","source":"print(\"⏳ Loading train_sequences.csv...\")\ndf_seq = pd.read_csv(f'{BASE_DIR}/train_sequences.csv')\n\nprint(f\"📌 Sequences shape: {df_seq.shape}\")\ndf_seq.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:52.645432Z","iopub.execute_input":"2026-01-08T17:07:52.645929Z","iopub.status.idle":"2026-01-08T17:07:53.030266Z","shell.execute_reply.started":"2026-01-08T17:07:52.6459Z","shell.execute_reply":"2026-01-08T17:07:53.029348Z"}},"outputs":[],"execution_count":null},{"id":"28eac3be-05bb-4d6a-a013-8c2d6b515dfe","cell_type":"markdown","source":"## 2. Sequence Length Distribution 📏","metadata":{}},{"id":"e4849d9e-2bbc-47bc-ba8b-80d9bd80a7e8","cell_type":"code","source":"# Calculate length\ndf_seq['length'] = df_seq['sequence'].apply(len)\n\n# 📈 Plot histogram\nplt.figure(figsize=(12, 6))\nsns.histplot(df_seq['length'], bins=50, kde=True, color='skyblue')\nplt.title('Distribution of RNA Sequence Lengths ', fontsize=16)\nplt.xlabel('Length (nucleotides)', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:53.031715Z","iopub.execute_input":"2026-01-08T17:07:53.032165Z","iopub.status.idle":"2026-01-08T17:07:53.622451Z","shell.execute_reply.started":"2026-01-08T17:07:53.03214Z","shell.execute_reply":"2026-01-08T17:07:53.621443Z"}},"outputs":[],"execution_count":null},{"id":"c634cd4a-2037-4a42-9337-8641041ad9bd","cell_type":"markdown","source":"## 3. Character Distribution 🧬","metadata":{}},{"id":"3d9092e0-d7b3-4cba-8988-d05fcab48c20","cell_type":"code","source":"all_chars = ''.join(df_seq['sequence'].tolist())\nchar_counts = Counter(all_chars)\n\nplt.figure(figsize=(6, 6))\nplt.pie(char_counts.values(), labels=char_counts.keys(), autopct='%1.1f%%', colors=sns.color_palette('pastel'))\nplt.title('Nucleotide Composition')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:53.623551Z","iopub.execute_input":"2026-01-08T17:07:53.624021Z","iopub.status.idle":"2026-01-08T17:07:54.234587Z","shell.execute_reply.started":"2026-01-08T17:07:53.623991Z","shell.execute_reply":"2026-01-08T17:07:54.233414Z"}},"outputs":[],"execution_count":null},{"id":"cc239e79-a8b7-4119-baf0-e8a527cf0fd3","cell_type":"markdown","source":"## 4. Ligand Analysis 💊\nUnderstanding what binds to the RNA structure (ions like Mg2+, drugs, etc.) is critical for 3D folding.\nLet's see the most common ligands in the dataset.","metadata":{}},{"id":"bd4d463d-e1f6-4c4c-8266-e816ddffa6fc","cell_type":"code","source":"# Filter out NaNs and split multiple ligands (separated by ';')\nligands_series = df_seq['ligand_ids'].dropna().astype(str)\nall_ligands = []\nfor s in ligands_series:\n    all_ligands.extend(s.split(';'))\n\nligand_counts = Counter(all_ligands)\nprint(f\"Total unique ligands found: {len(ligand_counts)}\")\n\n# Plot top 15 ligands\ntop_ligands = dict(ligand_counts.most_common(15))\n\nplt.figure(figsize=(12, 6))\nsns.barplot(x=list(top_ligands.values()), y=list(top_ligands.keys()), palette='viridis')\nplt.title('Top 15 Most Common Ligands Found in Structures ', fontsize=16)\nplt.xlabel('Count')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:54.237033Z","iopub.execute_input":"2026-01-08T17:07:54.237352Z","iopub.status.idle":"2026-01-08T17:07:54.568163Z","shell.execute_reply.started":"2026-01-08T17:07:54.237328Z","shell.execute_reply":"2026-01-08T17:07:54.567212Z"}},"outputs":[],"execution_count":null},{"id":"4ffc44aa-f60e-4359-9d9c-c5d8d8f57bf1","cell_type":"markdown","source":"## 5. Stoichiometry (Complexes) 🧩\nPart 2 of the competition introduces complexes (multiple chains). Let's look at the `stoichiometry` column to understand the variety.\nExamples:\n- `A:1` : Single chain, one copy.\n- `A:1,B:1` : Heterodimer (Two different chains).\n- `A:2` : Homodimer (Same chain, 2 copies).","metadata":{}},{"id":"7ce89bde-3273-4d0e-a9d0-019f4fa708a6","cell_type":"code","source":"stoich_counts = df_seq['stoichiometry'].value_counts().head(10)\n\nplt.figure(figsize=(12, 6))\nsns.barplot(x=stoich_counts.values, y=stoich_counts.index, palette='magma')\nplt.title('Top 10 Stoichiometry Configurations ', fontsize=16)\nplt.xlabel('Count')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:54.5691Z","iopub.execute_input":"2026-01-08T17:07:54.569376Z","iopub.status.idle":"2026-01-08T17:07:54.870397Z","shell.execute_reply.started":"2026-01-08T17:07:54.569355Z","shell.execute_reply":"2026-01-08T17:07:54.868789Z"}},"outputs":[],"execution_count":null},{"id":"47206716-35fb-4492-b76e-30779179b6f0","cell_type":"markdown","source":"## 6. WordCloud of Descriptions ☁️\nWhat kind of RNAs are we dealing with? Ribosomes? Riboswitches? Let's check the descriptions.","metadata":{}},{"id":"d4a191e1-03b0-4b06-a2e1-70dc2d4f8d97","cell_type":"code","source":"text = \" \".join(title for title in df_seq.description.dropna())\n\n# Create WordCloud\nwordcloud = WordCloud(width=800, height=400, background_color='white', colormap='ocean').generate(text)\n\nplt.figure(figsize=(15, 7))\nplt.imshow(wordcloud, interpolation='bilinear')\nplt.axis(\"off\")\nplt.title('Word Cloud of RNA Descriptions ', fontsize=20)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:54.872641Z","iopub.execute_input":"2026-01-08T17:07:54.873023Z","iopub.status.idle":"2026-01-08T17:07:56.344191Z","shell.execute_reply.started":"2026-01-08T17:07:54.873Z","shell.execute_reply":"2026-01-08T17:07:56.342999Z"}},"outputs":[],"execution_count":null},{"id":"ec49ed49-6e3d-49e1-9d50-1ad8ca5de183","cell_type":"markdown","source":"## 7. Visualizing 3D Structure 🧊\nFinally, visualization of the C1' coordinates.","metadata":{}},{"id":"ba420880-bcd6-43a1-992c-d5161f1eddc9","cell_type":"code","source":"print(\"⏳ Loading train_labels.csv...\")\ndf_labels = pd.read_csv(f'{BASE_DIR}/train_labels.csv')\n\ntarget_id_sample = \"157D\"\nsample_coords = df_labels[df_labels['ID'].astype(str).str.startswith(target_id_sample)]\n\nif not sample_coords.empty:\n    fig = plt.figure(figsize=(10, 10))\n    ax = fig.add_subplot(111, projection='3d')\n    \n    # Use x_1, y_1, z_1\n    ax.plot(sample_coords['x_1'], sample_coords['y_1'], sample_coords['z_1'], \n            marker='o', linestyle='-', markersize=5, alpha=0.6, color='purple')\n    \n    ax.set_title(f'3D Structure of {target_id_sample} ', fontsize=15)\n    plt.show()\nelse:\n    print(\"❌ Target not found.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-08T17:07:56.345389Z","iopub.execute_input":"2026-01-08T17:07:56.345824Z","iopub.status.idle":"2026-01-08T17:08:08.335376Z","shell.execute_reply.started":"2026-01-08T17:07:56.345743Z","shell.execute_reply":"2026-01-08T17:08:08.334273Z"}},"outputs":[],"execution_count":null},{"id":"a220f074-e4e5-4841-b242-bdc1a61c4be7","cell_type":"markdown","source":"## Conclusion 🏆\nThis expanded EDA gives us a much better understanding of the dataset diversity (ligands, complexes, types). Good luck with your modeling!","metadata":{}}]}