{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":41875,"databundleVersionId":5521661,"sourceType":"competition"},{"sourceId":116062,"databundleVersionId":14084779,"sourceType":"competition"}],"dockerImageVersionId":31153,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!pip install biopython\n!pip install goatools","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:29:03.967019Z","iopub.execute_input":"2025-11-01T01:29:03.967343Z","iopub.status.idle":"2025-11-01T01:29:50.553325Z","shell.execute_reply.started":"2025-11-01T01:29:03.96732Z","shell.execute_reply":"2025-11-01T01:29:50.552164Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# EDA on GO Terms Differences Between CAFA-5 and CAFA-6\n\nThis notebook performs exploratory data analysis on the differences between GO terms assigned to overlapping protein sequences in CAFA-5 and CAFA-6 competitions.\n\n## Table of Contents\n1. [Introduction](#Introduction)\n2. [Imports](#Imports)\n3. [Helper Functions](#Helper-Functions)\n4. [Data Loading](#Data-Loading)\n5. [Sequence Overlap Analysis](#Sequence-Overlap-Analysis)\n6. [Data Analysis and Visualization](#Data-Analysis-and-Visualization)\n7. [Conclusion](#Conclusion)","metadata":{}},{"cell_type":"markdown","source":"## Introduction\n\nThe Critical Assessment of protein Function Annotation (CAFA) challenges evaluate automated methods for protein function prediction. This analysis focuses on comparing GO (Gene Ontology) term annotations between CAFA-5 and CAFA-6 datasets for overlapping protein sequences.\n\nKey questions addressed:\n- How do GO term annotations differ between CAFA-5 and CAFA-6?\n- What patterns emerge in the annotation differences?\n- How can we quantify the similarity between annotations?","metadata":{}},{"cell_type":"markdown","source":"## Imports\n\nImport necessary libraries for data analysis and visualization.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom collections import Counter\nimport warnings\nimport os\nfrom Bio import SeqIO\nfrom collections import defaultdict\nfrom tqdm import tqdm\n\nwarnings.filterwarnings('ignore')\n\n# Set style\nplt.style.use('default')\nsns.set_palette(\"husl\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:29:50.554367Z","iopub.execute_input":"2025-11-01T01:29:50.554812Z","iopub.status.idle":"2025-11-01T01:29:52.086369Z","shell.execute_reply.started":"2025-11-01T01:29:50.554777Z","shell.execute_reply":"2025-11-01T01:29:52.085565Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Helper Functions\n\nDefine utility functions for loading sequences and analyzing FASTA files.","metadata":{}},{"cell_type":"code","source":"def load_sequences(file_path):\n    \"\"\"Load sequences into a dict {id: sequence_str}.\"\"\"\n    if not os.path.exists(file_path):\n        return {}\n\n    seq_dict = {}\n    for record in SeqIO.parse(file_path, 'fasta'):\n        seq_dict[record.id] = str(record.seq)\n    return seq_dict\n\n\ndef analyze_fasta(file_path):\n    \"\"\"Analyze a FASTA file and return statistics.\"\"\"\n    if not os.path.exists(file_path):\n        return None\n\n    sequences = list(SeqIO.parse(file_path, 'fasta'))\n    lengths = [len(seq.seq) for seq in sequences]\n\n    stats = {\n        'num_sequences': len(sequences),\n        'total_length': sum(lengths),\n        'avg_length': np.mean(lengths),\n        'min_length': min(lengths),\n        'max_length': max(lengths)\n    }\n\n    return stats","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:29:52.088519Z","iopub.execute_input":"2025-11-01T01:29:52.089127Z","iopub.status.idle":"2025-11-01T01:29:52.095494Z","shell.execute_reply.started":"2025-11-01T01:29:52.089103Z","shell.execute_reply":"2025-11-01T01:29:52.094744Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Loading\n\nLoad sequences and labels from CAFA-5 and CAFA-6 datasets.","metadata":{}},{"cell_type":"code","source":"# Define file paths\ncafa5_train_fasta = '/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta'\ncafa5_test_fasta = '/kaggle/input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset.fasta'\ncafa6_train_fasta = '/kaggle/input/cafa-6-protein-function-prediction/Train/train_sequences.fasta'\ncafa6_test_fasta = '/kaggle/input/cafa-6-protein-function-prediction/Test/testsuperset.fasta'\n\n# Label files\ncafa5_terms = '/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv'\ncafa6_terms = '/kaggle/input/cafa-6-protein-function-prediction/Train/train_terms.tsv'\n\n# Load sequences\ncafa5_train_sequences = load_sequences(cafa5_train_fasta)\ncafa5_test_sequences = load_sequences(cafa5_test_fasta)\ncafa6_train_sequences = load_sequences(cafa6_train_fasta)\ncafa6_test_sequences = load_sequences(cafa6_test_fasta)\n\n# Load labels\ncafa5_terms_df = pd.read_csv(cafa5_terms, sep='\\t')\ncafa6_terms_df = pd.read_csv(cafa6_terms, sep='\\t')\n\n# Create label dictionaries\ncafa5_labels = defaultdict(set)\nfor _, row in cafa5_terms_df.iterrows():\n    cafa5_labels[row['EntryID']].add(row['term'])\n\ncafa6_labels = defaultdict(set)\nfor _, row in cafa6_terms_df.iterrows():\n    cafa6_labels[row['EntryID']].add(row['term'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:29:52.096553Z","iopub.execute_input":"2025-11-01T01:29:52.096964Z","iopub.status.idle":"2025-11-01T01:34:16.314665Z","shell.execute_reply.started":"2025-11-01T01:29:52.096933Z","shell.execute_reply":"2025-11-01T01:34:16.313555Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Sequence Overlap Analysis\n\nAnalyze the overlap between sequences in CAFA-5 and CAFA-6 datasets.","metadata":{}},{"cell_type":"code","source":"# Create sets of sequences\ncafa5_train_seqs = set(cafa5_train_sequences.values())\ncafa5_test_seqs = set(cafa5_test_sequences.values())\ncafa5_all_seqs = cafa5_train_seqs.union(cafa5_test_seqs)\n\ncafa6_train_seqs = set(cafa6_train_sequences.values())\ncafa6_test_seqs = set(cafa6_test_sequences.values())\ncafa6_all_seqs = cafa6_train_seqs.union(cafa6_test_seqs)\n\nprint(\"Dataset Sizes:\")\nprint(\"=\" * 50)\nprint(f\"CAFA-5 Train: {len(cafa5_train_sequences)} sequences\")\nprint(f\"CAFA-5 Test:  {len(cafa5_test_sequences)} sequences\")\nprint(f\"CAFA-6 Train: {len(cafa6_train_sequences)} sequences\")\nprint(f\"CAFA-6 Test:  {len(cafa6_test_sequences)} sequences\")\n\nprint(\"\\n\" + \"=\" * 50)\nprint(\"Sequence Overlap Analysis:\")\nprint(\"=\" * 50)\n\nprint(f\"\\nCAFA-5 total unique sequences: {len(cafa5_all_seqs)}\")\nprint(f\"CAFA-6 total unique sequences: {len(cafa6_all_seqs)}\")\n\nseq_intersection = cafa5_all_seqs.intersection(cafa6_all_seqs)\nprint(f\"Identical sequences in both CAFA-5 and CAFA-6: {len(seq_intersection)}\")\n\n# Detailed sequence overlaps\ntrain_train_seq_overlap = cafa5_train_seqs.intersection(cafa6_train_seqs)\ntrain_test_seq_overlap = cafa5_train_seqs.intersection(cafa6_test_seqs)\ntest_train_seq_overlap = cafa5_test_seqs.intersection(cafa6_train_seqs)\ntest_test_seq_overlap = cafa5_test_seqs.intersection(cafa6_test_seqs)\n\nprint(f\"\\nDetailed sequence overlaps:\")\nprint(f\"CAFA-5 Train ∩ CAFA-6 Train: {len(train_train_seq_overlap)} ({len(train_train_seq_overlap)/len(cafa5_train_seqs)*100:.2f}% of CAFA-5 Train, {len(train_train_seq_overlap)/len(cafa6_train_seqs)*100:.2f}% of CAFA-6 Train)\")\nprint(f\"CAFA-5 Train ∩ CAFA-6 Test: {len(train_test_seq_overlap)} ({len(train_test_seq_overlap)/len(cafa5_train_seqs)*100:.2f}% of CAFA-5 Train, {len(train_test_seq_overlap)/len(cafa6_test_seqs)*100:.2f}% of CAFA-6 Test)\")\nprint(f\"CAFA-5 Test ∩ CAFA-6 Train: {len(test_train_seq_overlap)} ({len(test_train_seq_overlap)/len(cafa5_test_seqs)*100:.2f}% of CAFA-5 Test, {len(test_train_seq_overlap)/len(cafa6_train_seqs)*100:.2f}% of CAFA-6 Train)\")\nprint(f\"CAFA-5 Test ∩ CAFA-6 Test: {len(test_test_seq_overlap)} ({len(test_test_seq_overlap)/len(cafa5_test_seqs)*100:.2f}% of CAFA-5 Test, {len(test_test_seq_overlap)/len(cafa6_test_seqs)*100:.2f}% of CAFA-6 Test)\")\n\n# Whole dataset overlaps\nprint(f\"\\nWhole dataset overlaps:\")\ncafa5_train_vs_cafa6_all = cafa5_train_seqs.intersection(cafa6_all_seqs)\ncafa5_test_vs_cafa6_all = cafa5_test_seqs.intersection(cafa6_all_seqs)\ncafa6_train_vs_cafa5_all = cafa6_train_seqs.intersection(cafa5_all_seqs)\ncafa6_test_vs_cafa5_all = cafa6_test_seqs.intersection(cafa5_all_seqs)\n\nprint(f\"CAFA-5 Train ∩ CAFA-6 (all): {len(cafa5_train_vs_cafa6_all)} ({len(cafa5_train_vs_cafa6_all)/len(cafa5_train_seqs)*100:.2f}% of CAFA-5 Train)\")\nprint(f\"CAFA-5 Test ∩ CAFA-6 (all): {len(cafa5_test_vs_cafa6_all)} ({len(cafa5_test_vs_cafa6_all)/len(cafa5_test_seqs)*100:.2f}% of CAFA-5 Test)\")\nprint(f\"CAFA-6 Train ∩ CAFA-5 (all): {len(cafa6_train_vs_cafa5_all)} ({len(cafa6_train_vs_cafa5_all)/len(cafa6_train_seqs)*100:.2f}% of CAFA-6 Train)\")\nprint(f\"CAFA-6 Test ∩ CAFA-5 (all): {len(cafa6_test_vs_cafa5_all)} ({len(cafa6_test_vs_cafa5_all)/len(cafa6_test_seqs)*100:.2f}% of CAFA-6 Test)\")\n\n# Combined train+test overlaps\nprint(f\"\\nCombined Train+Test overlaps:\")\ncafa5_all_vs_cafa6_all = cafa5_all_seqs.intersection(cafa6_all_seqs)\nprint(f\"CAFA-5 (Train+Test) ∩ CAFA-6 (Train+Test): {len(cafa5_all_vs_cafa6_all)} ({len(cafa5_all_vs_cafa6_all)/len(cafa5_all_seqs)*100:.2f}% of CAFA-5 all, {len(cafa5_all_vs_cafa6_all)/len(cafa6_all_seqs)*100:.2f}% of CAFA-6 all)\")\n\n# New sequences in CAFA-6\nnew_in_cafa6 = cafa6_all_seqs - cafa5_all_seqs\nprint(f\"New sequences in CAFA-6 (not in CAFA-5): {len(new_in_cafa6)} ({len(new_in_cafa6)/len(cafa6_all_seqs)*100:.2f}% of CAFA-6)\")\n\n# Sequences only in CAFA-5\nonly_in_cafa5 = cafa5_all_seqs - cafa6_all_seqs\nprint(f\"Sequences only in CAFA-5 (not in CAFA-6): {len(only_in_cafa5)} ({len(only_in_cafa5)/len(cafa5_all_seqs)*100:.2f}% of CAFA-5)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:16.316328Z","iopub.execute_input":"2025-11-01T01:34:16.316735Z","iopub.status.idle":"2025-11-01T01:34:17.292602Z","shell.execute_reply.started":"2025-11-01T01:34:16.316703Z","shell.execute_reply":"2025-11-01T01:34:17.291745Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create DataFrame for overlapping sequences with labels\nprint(\"\\n\" + \"=\" * 50)\nprint(\"Creating DataFrame for Overlapping Sequences:\")\nprint(\"=\" * 50)\n\n# Create reverse mapping: sequence -> list of IDs\nseq_to_ids_cafa5 = defaultdict(list)\nfor id_, seq in cafa5_train_sequences.items():\n    seq_to_ids_cafa5[seq].append(id_)\nfor id_, seq in cafa5_test_sequences.items():\n    seq_to_ids_cafa5[seq].append(id_)\n\nseq_to_ids_cafa6 = defaultdict(list)\nfor id_, seq in cafa6_train_sequences.items():\n    seq_to_ids_cafa6[seq].append(id_)\nfor id_, seq in cafa6_test_sequences.items():\n    seq_to_ids_cafa6[seq].append(id_)\n\n# Create DataFrame\noverlapping_data = []\nfor seq in tqdm(cafa5_all_vs_cafa6_all):\n    cafa5_ids = seq_to_ids_cafa5.get(seq, [])\n    cafa6_ids = seq_to_ids_cafa6.get(seq, [])\n\n    # Get labels for these IDs\n    cafa5_go_terms = set()\n    cafa6_go_terms = set()\n\n    for id_ in cafa5_ids:\n        cafa5_go_terms.update(cafa5_labels.get(id_, set()))\n    for id_ in cafa6_ids:\n        cafa6_go_terms.update(cafa6_labels.get(id_, set()))\n\n    overlapping_data.append({\n        'sequence': seq,\n        'cafa5_ids': cafa5_ids,\n        'cafa6_ids': cafa6_ids,\n        'cafa5_go_terms': list(cafa5_go_terms),\n        'cafa6_go_terms': list(cafa6_go_terms),\n        'num_cafa5_terms': len(cafa5_go_terms),\n        'num_cafa6_terms': len(cafa6_go_terms)\n    })\n\ndf_overlapping = pd.DataFrame(overlapping_data)\nprint(f\"Created DataFrame with {len(df_overlapping)} overlapping sequences\")\nprint(f\"Columns: {list(df_overlapping.columns)}\")\n\n# Summary statistics\nprint(f\"\\nSummary:\")\nprint(f\"Total overlapping sequences: {len(df_overlapping)}\")\nprint(f\"Average CAFA-5 GO terms per sequence: {df_overlapping['num_cafa5_terms'].mean():.2f}\")\nprint(f\"Average CAFA-6 GO terms per sequence: {df_overlapping['num_cafa6_terms'].mean():.2f}\")\nprint(f\"Max CAFA-5 GO terms: {df_overlapping['num_cafa5_terms'].max()}\")\nprint(f\"Max CAFA-6 GO terms: {df_overlapping['num_cafa6_terms'].max()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:17.293596Z","iopub.execute_input":"2025-11-01T01:34:17.293919Z","iopub.status.idle":"2025-11-01T01:34:23.227603Z","shell.execute_reply.started":"2025-11-01T01:34:17.293888Z","shell.execute_reply":"2025-11-01T01:34:23.226546Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ID-based overlaps (as before)\nprint(\"\\n\" + \"=\" * 50)\nprint(\"ID Overlap Analysis:\")\nprint(\"=\" * 50)\n\ncafa5_train_ids = set(cafa5_train_sequences.keys())\ncafa5_test_ids = set(cafa5_test_sequences.keys())\ncafa5_all_ids = cafa5_train_ids.union(cafa5_test_ids)\n\ncafa6_train_ids = set(cafa6_train_sequences.keys())\ncafa6_test_ids = set(cafa6_test_sequences.keys())\ncafa6_all_ids = cafa6_train_ids.union(cafa6_test_ids)\n\nprint(f\"\\nCAFA-5 total unique IDs: {len(cafa5_all_ids)}\")\nprint(f\"CAFA-6 total unique IDs: {len(cafa6_all_ids)}\")\n\nid_intersection = cafa5_all_ids.intersection(cafa6_all_ids)\nprint(f\"IDs present in both CAFA-5 and CAFA-6: {len(id_intersection)}\")\n\n# Detailed ID overlaps\nid_train_train_overlap = cafa5_train_ids.intersection(cafa6_train_ids)\nid_train_test_overlap = cafa5_train_ids.intersection(cafa6_test_ids)\nid_test_train_overlap = cafa5_test_ids.intersection(cafa6_train_ids)\nid_test_test_overlap = cafa5_test_ids.intersection(cafa6_test_ids)\n\nprint(f\"\\nDetailed ID overlaps:\")\nprint(f\"CAFA-5 Train ∩ CAFA-6 Train: {len(id_train_train_overlap)}\")\nprint(f\"CAFA-5 Train ∩ CAFA-6 Test: {len(id_train_test_overlap)}\")\nprint(f\"CAFA-5 Test ∩ CAFA-6 Train: {len(id_test_train_overlap)}\")\nprint(f\"CAFA-5 Test ∩ CAFA-6 Test: {len(id_test_test_overlap)}\")\n\ndf_overlapping['cafa5_go_terms_len'] = df_overlapping['cafa5_go_terms'].apply(len)\ndf_overlapping['cafa6_go_terms_len'] = df_overlapping['cafa6_go_terms'].apply(len)\ndf = df_overlapping.loc[(df_overlapping['cafa5_go_terms_len'] > 0) & (df_overlapping['cafa6_go_terms_len'] > 0)]\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:23.22871Z","iopub.execute_input":"2025-11-01T01:34:23.229011Z","iopub.status.idle":"2025-11-01T01:34:23.688549Z","shell.execute_reply.started":"2025-11-01T01:34:23.228987Z","shell.execute_reply":"2025-11-01T01:34:23.687211Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Analysis and Visualization\n\nThis section analyzes the GO terms differences and creates visualizations.","metadata":{}},{"cell_type":"markdown","source":"### Distribution of GO terms per sequence","metadata":{}},{"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(15, 6))\n\n# CAFA-5 GO terms distribution\naxes[0].hist(df['num_cafa5_terms'], bins=50, alpha=0.7, edgecolor='black')\naxes[0].set_title('Distribution of CAFA-5 GO Terms per Sequence')\naxes[0].set_xlabel('Number of GO Terms')\naxes[0].set_ylabel('Frequency')\naxes[0].axvline(df['num_cafa5_terms'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"num_cafa5_terms\"].mean():.1f}')\naxes[0].legend()\n\n# CAFA-6 GO terms distribution\naxes[1].hist(df['num_cafa6_terms'], bins=50, alpha=0.7, edgecolor='black')\naxes[1].set_title('Distribution of CAFA-6 GO Terms per Sequence')\naxes[1].set_xlabel('Number of GO Terms')\naxes[1].set_ylabel('Frequency')\naxes[1].axvline(df['num_cafa6_terms'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"num_cafa6_terms\"].mean():.1f}')\naxes[1].legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:23.689997Z","iopub.execute_input":"2025-11-01T01:34:23.690339Z","iopub.status.idle":"2025-11-01T01:34:24.511545Z","shell.execute_reply.started":"2025-11-01T01:34:23.690307Z","shell.execute_reply":"2025-11-01T01:34:24.510555Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Scatter plot of CAFA-5 vs CAFA-6 GO terms","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 8))\nplt.scatter(df['num_cafa5_terms'], df['num_cafa6_terms'], alpha=0.6, s=30)\nplt.xlabel('Number of CAFA-5 GO Terms')\nplt.ylabel('Number of CAFA-6 GO Terms')\nplt.title('CAFA-5 vs CAFA-6 GO Terms per Sequence')\nplt.grid(True, alpha=0.3)\n\n# Add diagonal line\nmax_terms = max(df['num_cafa5_terms'].max(), df['num_cafa6_terms'].max())\nplt.plot([0, max_terms], [0, max_terms], 'r--', alpha=0.7, label='Equal terms')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:24.514292Z","iopub.execute_input":"2025-11-01T01:34:24.514594Z","iopub.status.idle":"2025-11-01T01:34:25.793295Z","shell.execute_reply.started":"2025-11-01T01:34:24.514571Z","shell.execute_reply":"2025-11-01T01:34:25.792359Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### GO Terms Difference Analysis","metadata":{}},{"cell_type":"code","source":"# Calculate differences\ndf['go_terms_diff'] = df['num_cafa6_terms'] - df['num_cafa5_terms']\ndf['go_terms_ratio'] = df.apply(lambda row: \n                               row['num_cafa6_terms'] / row['num_cafa5_terms'] if row['num_cafa5_terms'] > 0 \n                               else (float('inf') if row['num_cafa6_terms'] > 0 else 0), axis=1)\n\nprint(\"GO Terms Difference Statistics:\")\nprint(f\"Mean difference (CAFA-6 - CAFA-5): {df['go_terms_diff'].mean():.2f}\")\nprint(f\"Median difference: {df['go_terms_diff'].median():.2f}\")\nprint(f\"Sequences with more CAFA-6 terms: {(df['go_terms_diff'] > 0).sum()}\")\nprint(f\"Sequences with equal terms: {(df['go_terms_diff'] == 0).sum()}\")\nprint(f\"Sequences with more CAFA-5 terms: {(df['go_terms_diff'] < 0).sum()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:25.794229Z","iopub.execute_input":"2025-11-01T01:34:25.794528Z","iopub.status.idle":"2025-11-01T01:34:26.520401Z","shell.execute_reply.started":"2025-11-01T01:34:25.794505Z","shell.execute_reply":"2025-11-01T01:34:26.519305Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Distribution of differences","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nplt.hist(df['go_terms_diff'], bins=100, alpha=0.7, edgecolor='black')\nplt.xlabel('Difference (CAFA-6 GO Terms - CAFA-5 GO Terms)')\nplt.ylabel('Frequency')\nplt.title('Distribution of GO Terms Differences')\nplt.axvline(df['go_terms_diff'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"go_terms_diff\"].mean():.2f}')\nplt.axvline(df['go_terms_diff'].median(), color='green', linestyle='--', label=f'Median: {df[\"go_terms_diff\"].median():.2f}')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:26.521347Z","iopub.execute_input":"2025-11-01T01:34:26.522094Z","iopub.status.idle":"2025-11-01T01:34:26.884291Z","shell.execute_reply.started":"2025-11-01T01:34:26.522062Z","shell.execute_reply":"2025-11-01T01:34:26.883392Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### GO Terms Uniqueness Analysis","metadata":{}},{"cell_type":"code","source":"# Analyze GO terms that were skipped in CAFA-6\ndef flatten_go_terms(column):\n    all_terms = []\n    for terms in df[column]:\n        if isinstance(terms, list):\n            all_terms.extend(terms)\n    return all_terms\n\ncafa5_all_terms = flatten_go_terms('cafa5_go_terms')\ncafa6_all_terms = flatten_go_terms('cafa6_go_terms')\n\ncafa5_term_counts = Counter(cafa5_all_terms)\ncafa6_term_counts = Counter(cafa6_all_terms)\n\n# GO terms only in CAFA-5\nonly_cafa5_terms = set(cafa5_term_counts.keys()) - set(cafa6_term_counts.keys())\n# GO terms only in CAFA-6\nonly_cafa6_terms = set(cafa6_term_counts.keys()) - set(cafa5_term_counts.keys())\n# Common GO terms\ncommon_terms = set(cafa5_term_counts.keys()) & set(cafa6_term_counts.keys())\n\nprint(f\"GO terms only in CAFA-5: {len(only_cafa5_terms)}\")\nprint(f\"GO terms only in CAFA-6: {len(only_cafa6_terms)}\")\nprint(f\"Common GO terms: {len(common_terms)}\")\nprint(f\"Total unique GO terms in CAFA-5: {len(set(cafa5_term_counts.keys()))}\")\nprint(f\"Total unique GO terms in CAFA-6: {len(set(cafa6_term_counts.keys()))}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:26.885618Z","iopub.execute_input":"2025-11-01T01:34:26.885968Z","iopub.status.idle":"2025-11-01T01:34:27.54782Z","shell.execute_reply.started":"2025-11-01T01:34:26.885938Z","shell.execute_reply":"2025-11-01T01:34:27.546977Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Sequence Length Analysis","metadata":{}},{"cell_type":"code","source":"# Sequence length analysis\ndf['sequence_length'] = df['sequence'].str.len()\n\nplt.figure(figsize=(12, 6))\nplt.scatter(df['sequence_length'], df['go_terms_diff'], alpha=0.6, s=30)\nplt.xlabel('Sequence Length')\nplt.ylabel('GO Terms Difference (CAFA-6 - CAFA-5)')\nplt.title('Sequence Length vs GO Terms Difference')\nplt.grid(True, alpha=0.3)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:27.54878Z","iopub.execute_input":"2025-11-01T01:34:27.5491Z","iopub.status.idle":"2025-11-01T01:34:28.035618Z","shell.execute_reply.started":"2025-11-01T01:34:27.549073Z","shell.execute_reply":"2025-11-01T01:34:28.034676Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Correlation Analysis","metadata":{}},{"cell_type":"code","source":"# Correlation analysis\ncorrelation_matrix = df[['sequence_length', 'num_cafa5_terms', 'num_cafa6_terms', 'go_terms_diff']].corr()\n\nplt.figure(figsize=(8, 6))\nsns.heatmap(correlation_matrix, annot=True, cmap='coolwarm', center=0, fmt='.2f')\nplt.title('Correlation Matrix')\nplt.show()\n\nprint(\"Correlation with sequence length:\")\nprint(f\"CAFA-5 GO terms: {df['sequence_length'].corr(df['num_cafa5_terms']):.3f}\")\nprint(f\"CAFA-6 GO terms: {df['sequence_length'].corr(df['num_cafa6_terms']):.3f}\")\nprint(f\"GO terms difference: {df['sequence_length'].corr(df['go_terms_diff']):.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:28.036605Z","iopub.execute_input":"2025-11-01T01:34:28.036873Z","iopub.status.idle":"2025-11-01T01:34:28.315108Z","shell.execute_reply.started":"2025-11-01T01:34:28.036853Z","shell.execute_reply":"2025-11-01T01:34:28.313986Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Similarity Metrics Calculation","metadata":{}},{"cell_type":"code","source":"# Calculate recall and precision metrics for GO terms\ndef calculate_go_recall_precision(row):\n    cafa5_terms = set(row['cafa5_go_terms']) if isinstance(row['cafa5_go_terms'], list) else set()\n    cafa6_terms = set(row['cafa6_go_terms']) if isinstance(row['cafa6_go_terms'], list) else set()\n    \n    if not cafa6_terms:\n        return {'recall': 0, 'precision': 0, 'f1': 0, 'jaccard': 0}\n    \n    # Recall: how many CAFA-6 terms are in CAFA-5 (what fraction of CAFA-6 predictions are correct)\n    recall = len(cafa6_terms & cafa5_terms) / len(cafa6_terms) if cafa6_terms else 0\n    \n    # Precision: how many CAFA-5 terms are in CAFA-6 (what fraction of CAFA-5 terms are predicted)\n    precision = len(cafa5_terms & cafa6_terms) / len(cafa5_terms) if cafa5_terms else 0\n    \n    # F1 score\n    f1 = 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0\n    \n    # Jaccard similarity\n    jaccard = len(cafa5_terms & cafa6_terms) / len(cafa5_terms | cafa6_terms) if (cafa5_terms | cafa6_terms) else 0\n    \n    return {'recall': recall, 'precision': precision, 'f1': f1, 'jaccard': jaccard}\n\n# Apply metrics calculation\nmetrics_df = df.apply(calculate_go_recall_precision, axis=1, result_type='expand')\ndf = pd.concat([df, metrics_df], axis=1)\n\nprint(\"GO Terms Similarity Metrics:\")\nprint(f\"Mean Recall (CAFA-6 terms found in CAFA-5): {df['recall'].mean():.3f}\")\nprint(f\"Mean Precision (CAFA-5 terms found in CAFA-6): {df['precision'].mean():.3f}\")\nprint(f\"Mean F1 Score: {df['f1'].mean():.3f}\")\nprint(f\"Mean Jaccard Similarity: {df['jaccard'].mean():.3f}\")\n\n# Sequences with perfect recall (all CAFA-6 terms are in CAFA-5)\nperfect_recall = (df['recall'] == 1.0).sum()\nprint(f\"Sequences with perfect recall: {perfect_recall} ({perfect_recall/len(df)*100:.1f}%)\")\n\n# Sequences with perfect precision (all CAFA-5 terms are in CAFA-6)\nperfect_precision = (df['precision'] == 1.0).sum()\nprint(f\"Sequences with perfect precision: {perfect_precision} ({perfect_precision/len(df)*100:.1f}%)\")\n\n# Sequences with no overlap\nno_overlap = (df['jaccard'] == 0).sum()\nprint(f\"Sequences with no GO term overlap: {no_overlap} ({no_overlap/len(df)*100:.1f}%)\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:28.316236Z","iopub.execute_input":"2025-11-01T01:34:28.316605Z","iopub.status.idle":"2025-11-01T01:34:31.34466Z","shell.execute_reply.started":"2025-11-01T01:34:28.316576Z","shell.execute_reply":"2025-11-01T01:34:31.343761Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Distribution of Similarity Metrics","metadata":{}},{"cell_type":"code","source":"# Distribution of recall and precision\nfig, axes = plt.subplots(2, 2, figsize=(15, 12))\n\n# Recall distribution\naxes[0,0].hist(df['recall'], bins=50, alpha=0.7, edgecolor='black')\naxes[0,0].set_title('Distribution of Recall')\naxes[0,0].set_xlabel('Recall')\naxes[0,0].set_ylabel('Frequency')\naxes[0,0].axvline(df['recall'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"recall\"].mean():.3f}')\naxes[0,0].legend()\n\n# Precision distribution\naxes[0,1].hist(df['precision'], bins=50, alpha=0.7, edgecolor='black')\naxes[0,1].set_title('Distribution of Precision')\naxes[0,1].set_xlabel('Precision')\naxes[0,1].set_ylabel('Frequency')\naxes[0,1].axvline(df['precision'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"precision\"].mean():.3f}')\naxes[0,1].legend()\n\n# F1 distribution\naxes[1,0].hist(df['f1'], bins=50, alpha=0.7, edgecolor='black')\naxes[1,0].set_title('Distribution of F1 Score')\naxes[1,0].set_xlabel('F1 Score')\naxes[1,0].set_ylabel('Frequency')\naxes[1,0].axvline(df['f1'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"f1\"].mean():.3f}')\naxes[1,0].legend()\n\n# Jaccard distribution\naxes[1,1].hist(df['jaccard'], bins=50, alpha=0.7, edgecolor='black')\naxes[1,1].set_title('Distribution of Jaccard Similarity')\naxes[1,1].set_xlabel('Jaccard Similarity')\naxes[1,1].set_ylabel('Frequency')\naxes[1,1].axvline(df['jaccard'].mean(), color='red', linestyle='--', label=f'Mean: {df[\"jaccard\"].mean():.3f}')\naxes[1,1].legend()\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:31.345527Z","iopub.execute_input":"2025-11-01T01:34:31.345771Z","iopub.status.idle":"2025-11-01T01:34:32.606871Z","shell.execute_reply.started":"2025-11-01T01:34:31.345751Z","shell.execute_reply":"2025-11-01T01:34:32.605746Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Recall vs Precision Scatter Plot","metadata":{}},{"cell_type":"code","source":"# Scatter plot: Recall vs Precision\nplt.figure(figsize=(10, 8))\nplt.scatter(df['recall'], df['precision'], alpha=0.6, s=30, c=df['sequence_length'], cmap='viridis')\nplt.xlabel('Recall')\nplt.ylabel('Precision')\nplt.title('Recall vs Precision (colored by sequence length)')\nplt.colorbar(label='Sequence Length')\nplt.grid(True, alpha=0.3)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:32.607957Z","iopub.execute_input":"2025-11-01T01:34:32.608288Z","iopub.status.idle":"2025-11-01T01:34:34.296017Z","shell.execute_reply.started":"2025-11-01T01:34:32.608258Z","shell.execute_reply":"2025-11-01T01:34:34.295008Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Fake CAFA-6 Predictions Analysis\n\nCreate fake CAFA-6 predictions by removing GO terms only in CAFA-5 from CAFA-5 annotations.","metadata":{}},{"cell_type":"code","source":"print(\"=\" * 60)\nprint(\"ANALYSIS: Creating Fake CAFA-6 Predictions from CAFA-5\")\nprint(\"=\" * 60)\n\nprint(f\"GO terms to remove from CAFA-5: {len(only_cafa5_terms)}\")\n\n# Create fake CAFA-6 predictions\ndef create_fake_cafa6(row):\n    cafa5_terms = set(row['cafa5_go_terms']) if isinstance(row['cafa5_go_terms'], list) else set()\n    # Remove terms that don't exist in CAFA-6\n    fake_terms = cafa5_terms - only_cafa5_terms\n    return list(fake_terms)\n\ndf['fake_cafa6_go_terms'] = df.apply(create_fake_cafa6, axis=1)\ndf['num_fake_cafa6_terms'] = df['fake_cafa6_go_terms'].apply(len)\n\n# Calculate metrics between fake CAFA-6 and real CAFA-6\ndef calculate_fake_vs_real_metrics(row):\n    fake_terms = set(row['fake_cafa6_go_terms'])\n    real_terms = set(row['cafa6_go_terms']) if isinstance(row['cafa6_go_terms'], list) else set()\n    \n    if not real_terms:\n        return {'fake_recall': 0, 'fake_precision': 0, 'fake_f1': 0, 'fake_jaccard': 0}\n    \n    # Recall: how many real CAFA-6 terms are in fake predictions\n    recall = len(real_terms & fake_terms) / len(real_terms) if real_terms else 0\n    \n    # Precision: how many fake predictions are correct (in real CAFA-6)\n    precision = len(fake_terms & real_terms) / len(fake_terms) if fake_terms else 0\n    \n    # F1 score\n    f1 = 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0\n    \n    # Jaccard similarity\n    jaccard = len(fake_terms & real_terms) / len(fake_terms | real_terms) if (fake_terms | real_terms) else 0\n    \n    return {'fake_recall': recall, 'fake_precision': precision, 'fake_f1': f1, 'fake_jaccard': jaccard}\n\nfake_metrics_df = df.apply(calculate_fake_vs_real_metrics, axis=1, result_type='expand')\ndf = pd.concat([df, fake_metrics_df], axis=1)\n\nprint(f\"\\nFake CAFA-6 predictions from CAFA-5:\")\nprint(f\"Average terms per sequence: {df['num_fake_cafa6_terms'].mean():.2f}\")\nprint(f\"Sequences with fake predictions: {(df['num_fake_cafa6_terms'] > 0).sum()}\")\nprint(f\"Sequences with no fake predictions: {(df['num_fake_cafa6_terms'] == 0).sum()}\")\n\nprint(f\"\\nMetrics between Fake CAFA-6 and Real CAFA-6:\")\nprint(f\"Mean Recall (real terms found in fake): {df['fake_recall'].mean():.3f}\")\nprint(f\"Mean Precision (fake terms that are real): {df['fake_precision'].mean():.3f}\")\nprint(f\"Mean F1 Score: {df['fake_f1'].mean():.3f}\")\nprint(f\"Mean Jaccard Similarity: {df['fake_jaccard'].mean():.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:34.297161Z","iopub.execute_input":"2025-11-01T01:34:34.297474Z","iopub.status.idle":"2025-11-01T01:34:39.35537Z","shell.execute_reply.started":"2025-11-01T01:34:34.297427Z","shell.execute_reply":"2025-11-01T01:34:39.354536Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate original metrics first\ndef calculate_original_metrics(row):\n    cafa5_terms = set(row['cafa5_go_terms']) if isinstance(row['cafa5_go_terms'], list) else set()\n    cafa6_terms = set(row['cafa6_go_terms']) if isinstance(row['cafa6_go_terms'], list) else set()\n    \n    if not cafa6_terms:\n        return {'orig_recall': 0, 'orig_precision': 0, 'orig_f1': 0}\n    \n    # Recall: how many CAFA-6 terms are in CAFA-5\n    recall = len(cafa6_terms & cafa5_terms) / len(cafa6_terms)\n    \n    # Precision: how many CAFA-5 terms are in CAFA-6\n    precision = len(cafa5_terms & cafa6_terms) / len(cafa5_terms) if cafa5_terms else 0\n    \n    # F1 score\n    f1 = 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0\n    \n    return {'orig_recall': recall, 'orig_precision': precision, 'orig_f1': f1}\n\norig_metrics_df = df.apply(calculate_original_metrics, axis=1, result_type='expand')\ndf = pd.concat([df, orig_metrics_df], axis=1)\n\n# Compare with original CAFA-5 performance\nprint(f\"\\nComparison with original CAFA-5 vs CAFA-6:\")\nprint(f\"Original CAFA-5 Recall: {df['orig_recall'].mean():.3f} → Fake CAFA-6 Recall: {df['fake_recall'].mean():.3f}\")\nprint(f\"Original CAFA-5 Precision: {df['orig_precision'].mean():.3f} → Fake CAFA-6 Precision: {df['fake_precision'].mean():.3f}\")\nprint(f\"Original CAFA-5 F1: {df['orig_f1'].mean():.3f} → Fake CAFA-6 F1: {df['fake_f1'].mean():.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-01T01:34:39.356648Z","iopub.execute_input":"2025-11-01T01:34:39.356978Z","iopub.status.idle":"2025-11-01T01:34:42.976294Z","shell.execute_reply.started":"2025-11-01T01:34:39.356948Z","shell.execute_reply":"2025-11-01T01:34:42.975509Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Conclusion\n\nThis analysis provides insights into the differences between GO term annotations in CAFA-5 and CAFA-6 datasets for overlapping protein sequences.\n\n### Key Findings:\n\n1. **Sequence Overlap**: CAFA-5 Train ∩ CAFA-6 Train: 77781 (55.99% of CAFA-5 Train, 95.81% of CAFA-6 Train)\n\n2. **GO Term Differences**: \n    - GO terms only in CAFA-5: 5027\n    - GO terms only in CAFA-6: 854\n    - Common GO terms: 25090\n    - Total unique GO terms in CAFA-5: 30117\n    - Total unique GO terms in CAFA-6: 25944\n      \n3. **GO Term Differences Statistics**:\n    - Mean difference (CAFA-6 - CAFA-5): -40.69\n    - Median difference: -30.00\n    - Sequences with more CAFA-6 terms: 259\n    - Sequences with equal terms: 119\n    - Sequences with more CAFA-5 terms: 77403\n\n4. **Similarity Metrics**:\n    - Mean Recall (CAFA-6 terms found in CAFA-5): 0.942\n    - Mean Precision (CAFA-5 terms found in CAFA-6): 0.151\n    - Mean F1 Score: 0.251\n    - Mean Jaccard Similarity: 0.148\n    - Sequences with perfect recall: 61225 (78.7%)\n    - Sequences with perfect precision: 0 (0.0%)\n    - Sequences with no GO term overlap: 142 (0.2%)\n\n5. **Fake CAFA-6 Predictions**:\n   - Fake CAFA-6 Predictions - are simply taken predictions from CAFA-5 and removed GO terms which CAFA-6 dosen't have.\n\n6. **Comparison with original CAFA-5 vs CAFA-6**:\n   - Original CAFA-5 Recall: 0.942 → Fake CAFA-6 Recall: 0.942\n   - Original CAFA-5 Precision: 0.151 → Fake CAFA-6 Precision: 0.241\n   - Original CAFA-5 F1: 0.251 → Fake CAFA-6 F1: 0.365\n\n### Implications:\n\n- The analysis shows how GO annotations have evolved between CAFA challenges\n- Understanding these differences can help improve protein function prediction methods\n- The fake CAFA-6 predictions analysis demonstrates the impact of GO term curation changes\n\n### Future work:\n\nThere appears to be substantial potential to augment the dataset using CAFA-5, thereby facilitating progressive enhancement in precision, as recall metrics have already reached satisfactory levels.","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}