{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Exploring Protein Sequence Taxonomy and Analysis\n\n## Introduction\n\nIn the realm of bioinformatics, understanding the taxonomic classification and properties of protein sequences plays a crucial role in unraveling the complexities of molecular biology. This notebook delves into the world of protein sequences, employing the Biopython library and the Entrez module to extract taxonomic information and perform a diverse range of protein sequence analyses. From taxonomy exploration to structural and physicochemical insights, this notebook demonstrates how to harness the power of computational tools to glean valuable insights from protein sequences.\n\n## Objective\n\nThe primary objective of this notebook is to showcase an end-to-end analysis pipeline for protein sequence data. We will explore the taxonomic hierarchy of protein sequences by leveraging the Entrez module to fetch taxonomic details associated with unique taxonomy IDs. Additionally, we will delve into various protein sequence analysis techniques, ranging from amino acid composition and hydropathy scoring to physicochemical properties such as molecular weight, isoelectric point, and more. By the end of this notebook, you will have gained a comprehensive understanding of how to handle, analyze, and extract valuable information from protein sequence data.\n\n## Contents\n\n1. **Importing Libraries and Data**\n   - Loading the required libraries for the analysis\n   - Reading protein sequence data from input FASTA files\n   - Preparing the dataset for analysis\n   \n2. **Exploring Taxonomic Information**\n   - Utilizing the Entrez module to fetch taxonomic data\n   - Creating a DataFrame containing taxonomic details\n   \n3. **Protein Sequence Analysis**\n   - Calculating amino acid counts and percentages\n   - Estimating molecular weight and isoelectric point\n   - Evaluating hydropathy scores and other physicochemical properties\n   \n\n\nThis notebook serves as a comprehensive guide for both beginners and intermediate users interested in exploring protein sequence data using Python. By following along, readers will acquire the skills and knowledge needed to extract valuable information from protein sequences and gain a deeper understanding of their taxonomic and functional implications.\n","metadata":{"execution":{"iopub.status.busy":"2023-07-20T10:04:45.370005Z","iopub.execute_input":"2023-07-20T10:04:45.370981Z","iopub.status.idle":"2023-07-20T10:05:30.772971Z","shell.execute_reply.started":"2023-07-20T10:04:45.370922Z","shell.execute_reply":"2023-07-20T10:05:30.771574Z"}}},{"cell_type":"code","source":"!pip install obonet -q\n!pip install pyvis -q","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:36:22.480076Z","iopub.execute_input":"2023-08-15T04:36:22.480635Z","iopub.status.idle":"2023-08-15T04:36:53.942114Z","shell.execute_reply.started":"2023-08-15T04:36:22.480583Z","shell.execute_reply":"2023-08-15T04:36:53.940759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from Bio.SeqUtils.ProtParam import ProteinAnalysis\nimport os\nimport json\nfrom collections import Counter\nfrom Bio import SeqIO\nimport re\nimport random\nimport obonet\nimport networkx\nimport pandas as pd\nimport numpy as np\nfrom Bio import SeqIO\nfrom pyvis.network import Network\nfrom Bio.SeqUtils.ProtParam import ProteinAnalysis\nfrom tqdm import tqdm\nfrom Bio import Entrez","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:36:53.944218Z","iopub.execute_input":"2023-08-15T04:36:53.944608Z","iopub.status.idle":"2023-08-15T04:36:54.567539Z","shell.execute_reply.started":"2023-08-15T04:36:53.944569Z","shell.execute_reply":"2023-08-15T04:36:54.566243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time() \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-15T04:36:54.569441Z","iopub.execute_input":"2023-08-15T04:36:54.569891Z","iopub.status.idle":"2023-08-15T04:36:54.595116Z","shell.execute_reply.started":"2023-08-15T04:36:54.569843Z","shell.execute_reply":"2023-08-15T04:36:54.594068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"amino_acids=['A','R','N','D','C','E','Q','G','H','O','I','L','K','M','F','P','U','S','T','W','Y','V']","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:36:54.598893Z","iopub.execute_input":"2023-08-15T04:36:54.599307Z","iopub.status.idle":"2023-08-15T04:36:54.606429Z","shell.execute_reply.started":"2023-08-15T04:36:54.599257Z","shell.execute_reply":"2023-08-15T04:36:54.604838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# thanks https://www.kaggle.com/code/joonyoungjang/multi-input-model-with-t5-taxonomy\n\n# Define the path to the input FASTA file\nfasta_file = '/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta'\n\n# Initialize empty lists to store extracted information\nids = []\ndescriptions = []\nspecies = []\ngenes = []\nlengths = []\nsequences = []\nspecies_ID = []\n\n# Read fasta file using SeqIO.parse() method\nrecords = list(SeqIO.parse(fasta_file, 'fasta'))\n\n# Defining Regular Expression Patterns\npattern = re.compile(r'(?:OS=)([^\\s]+)')\n# Define another regular expression pattern to match species taxonomy IDs\npattern2 = re.compile(r'(?:OX=)([^\\s]+)')\n\n# Extracting values using regular expression patterns\nfor record in records:\n    # id\n    id_ = record.id\n    \n    # description\n    description = record.description\n    \n    # species\n    species_match = pattern.search(description)\n    if species_match:\n        species_ = species_match.group(1)\n    else:\n        species_ = ''\n        \n        \n    # species ID\n    species_id_match = pattern2.search(description)\n    if species_id_match:\n        species_id = species_id_match.group(1)\n    else:\n        species_id = ''    \n    \n    # gene\n    gene_match = re.search(r'(?:GN=)([^\\s]+)', description)\n    if gene_match:\n        gene_ = gene_match.group(1)\n    else:\n        gene_ = ''\n    \n    # length\n    length_ = len(record.seq)\n    \n    # sequence\n    sequence = str(record.seq)\n    \n    # Add extracted values to the list\n    ids.append(id_)\n    descriptions.append(description)\n    species.append(species_)\n    genes.append(gene_)\n    lengths.append(length_)\n    sequences.append(sequence)\n    species_ID.append(species_id)\n\n# Convert to Data Frame\ntrain_df = pd.DataFrame({\n    'EntryID': ids,\n    'description': descriptions,\n    'Species': species,\n    'gene': genes,\n    'length': lengths,\n    'sequence': sequences,\n    'taxonomyID':species_ID\n})\n\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:36:54.608027Z","iopub.execute_input":"2023-08-15T04:36:54.608394Z","iopub.status.idle":"2023-08-15T04:37:00.661281Z","shell.execute_reply.started":"2023-08-15T04:36:54.608358Z","shell.execute_reply":"2023-08-15T04:37:00.659881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.663432Z","iopub.execute_input":"2023-08-15T04:37:00.663971Z","iopub.status.idle":"2023-08-15T04:37:00.687972Z","shell.execute_reply.started":"2023-08-15T04:37:00.663916Z","shell.execute_reply":"2023-08-15T04:37:00.686503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Test Data","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.689506Z","iopub.execute_input":"2023-08-15T04:37:00.689895Z","iopub.status.idle":"2023-08-15T04:37:00.701461Z","shell.execute_reply.started":"2023-08-15T04:37:00.689856Z","shell.execute_reply":"2023-08-15T04:37:00.699844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Read the taxon list file into a pandas DataFrame\ntest_taxon = pd.read_csv('/kaggle/input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset-taxon-list.tsv', sep='\\t', encoding='ISO-8859-1')","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.706759Z","iopub.execute_input":"2023-08-15T04:37:00.707255Z","iopub.status.idle":"2023-08-15T04:37:00.7291Z","shell.execute_reply.started":"2023-08-15T04:37:00.707196Z","shell.execute_reply":"2023-08-15T04:37:00.727606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_taxon","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.731402Z","iopub.execute_input":"2023-08-15T04:37:00.731974Z","iopub.status.idle":"2023-08-15T04:37:00.747868Z","shell.execute_reply.started":"2023-08-15T04:37:00.731919Z","shell.execute_reply":"2023-08-15T04:37:00.74648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Rename the 'ID' column to 'taxonomyID' in the test_taxon DataFrame\ntest_taxon.rename(columns={'ID': 'taxonomyID'}, inplace=True)\ntest_taxon.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.753914Z","iopub.execute_input":"2023-08-15T04:37:00.754598Z","iopub.status.idle":"2023-08-15T04:37:00.772124Z","shell.execute_reply.started":"2023-08-15T04:37:00.754554Z","shell.execute_reply":"2023-08-15T04:37:00.770795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the values for the new row\nlist_row = [8667, \"Oxyuranus\"]\n# Add the new row to the test_taxon DataFrame\ntest_taxon.loc[len(test_taxon)] = list_row","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.775524Z","iopub.execute_input":"2023-08-15T04:37:00.775936Z","iopub.status.idle":"2023-08-15T04:37:00.784635Z","shell.execute_reply.started":"2023-08-15T04:37:00.775899Z","shell.execute_reply":"2023-08-15T04:37:00.783702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_taxon.to_csv('test_taxon.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.786316Z","iopub.execute_input":"2023-08-15T04:37:00.786705Z","iopub.status.idle":"2023-08-15T04:37:00.809072Z","shell.execute_reply.started":"2023-08-15T04:37:00.786666Z","shell.execute_reply":"2023-08-15T04:37:00.807727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# thanks https://www.kaggle.com/code/joonyoungjang/multi-input-model-with-t5-taxonomy\n# Define the path to the input FASTA file\nfasta_file = '/kaggle/input/cafa-5-protein-function-prediction/Test (Targets)/testsuperset.fasta'\n# Initialize a dictionary to store data\ndata = {'EntryID': [], 'taxonomyID': [], 'sequence': []}\n# Initialize a variable to keep track of the current entry\ncurrent_entry = None\n\nwith open(fasta_file, 'r') as f:\n    for line in f:\n        if line.startswith('>'):\n            # Extract entry ID and taxonomy ID from the line\n            entry_id, taxonomy_id = line.strip().lstrip('>').split('\\t')\n            current_entry = entry_id\n            data['EntryID'].append(entry_id)\n            data['taxonomyID'].append(taxonomy_id)\n            data['sequence'].append('')\n            \n        else:\n            data['sequence'][-1] += line.strip()\n           \n# Create a pandas DataFrame from the collected data\ntest_seq_df = pd.DataFrame(data)\ntest_seq_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:00.811031Z","iopub.execute_input":"2023-08-15T04:37:00.811834Z","iopub.status.idle":"2023-08-15T04:37:02.789232Z","shell.execute_reply.started":"2023-08-15T04:37:00.811777Z","shell.execute_reply":"2023-08-15T04:37:02.787814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_seq_df['taxonomyID'] = test_seq_df['taxonomyID'].astype(int)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:02.791311Z","iopub.execute_input":"2023-08-15T04:37:02.794325Z","iopub.status.idle":"2023-08-15T04:37:02.836947Z","shell.execute_reply.started":"2023-08-15T04:37:02.794275Z","shell.execute_reply":"2023-08-15T04:37:02.835282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Convert the 'taxonomyID' column in the test_seq_df DataFrame to integer type\ntest_df = pd.merge(test_seq_df,test_taxon,how=\"left\", on='taxonomyID')\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:02.838099Z","iopub.execute_input":"2023-08-15T04:37:02.838531Z","iopub.status.idle":"2023-08-15T04:37:02.925124Z","shell.execute_reply.started":"2023-08-15T04:37:02.83849Z","shell.execute_reply":"2023-08-15T04:37:02.923817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:02.92638Z","iopub.execute_input":"2023-08-15T04:37:02.927019Z","iopub.status.idle":"2023-08-15T04:37:02.944605Z","shell.execute_reply.started":"2023-08-15T04:37:02.92698Z","shell.execute_reply":"2023-08-15T04:37:02.943401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Split the 'Species' column in the test_df DataFrame based on whitespace and keep the first part\ntest_df['Species'] = test_df['Species'].str.split(' ', expand=True)[0]","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:02.946215Z","iopub.execute_input":"2023-08-15T04:37:02.947324Z","iopub.status.idle":"2023-08-15T04:37:03.746056Z","shell.execute_reply.started":"2023-08-15T04:37:02.947229Z","shell.execute_reply":"2023-08-15T04:37:03.744566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:03.74751Z","iopub.execute_input":"2023-08-15T04:37:03.747932Z","iopub.status.idle":"2023-08-15T04:37:03.767149Z","shell.execute_reply.started":"2023-08-15T04:37:03.747884Z","shell.execute_reply":"2023-08-15T04:37:03.765743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Capitalize the first letter of each word in the 'Species' column of the test_df DataFrame\ntest_df['Species'] = test_df['Species'].str.capitalize()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:03.768829Z","iopub.execute_input":"2023-08-15T04:37:03.769182Z","iopub.status.idle":"2023-08-15T04:37:03.880751Z","shell.execute_reply.started":"2023-08-15T04:37:03.769148Z","shell.execute_reply":"2023-08-15T04:37:03.879103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the length of each sequence in the 'sequence' column and store it in the 'length' column\ntest_df['length'] = test_df['sequence'].str.len()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:03.882325Z","iopub.execute_input":"2023-08-15T04:37:03.882748Z","iopub.status.idle":"2023-08-15T04:37:04.043785Z","shell.execute_reply.started":"2023-08-15T04:37:03.88268Z","shell.execute_reply":"2023-08-15T04:37:04.042596Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.to_csv('test_Species.csv', columns=['Species'])\ntrain_df.to_csv('train_Species.csv', columns=['Species'])","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:04.045977Z","iopub.execute_input":"2023-08-15T04:37:04.046513Z","iopub.status.idle":"2023-08-15T04:37:04.809795Z","shell.execute_reply.started":"2023-08-15T04:37:04.046458Z","shell.execute_reply":"2023-08-15T04:37:04.808367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to count the occurrence of a specific amino acid in a protein sequence\ndef amino_acids_count(row,acid):\n    try:\n        return ProteinAnalysis(row['sequence']).count_amino_acids()[acid]\n    except:\n        return 0\n# Function to calculate the percentage of a specific amino acid in a protein sequence    \ndef amino_acids_percent(row,acid):\n    try:\n        return ProteinAnalysis(row['sequence']).get_amino_acids_percent()[acid]\n    except:\n        return 0  \n# Function to calculate the molecular weight of a protein sequence    \ndef get_molecular_weight(row):\n    try:\n        return ProteinAnalysis(row['sequence']).molecular_weight()\n    except:\n        return 0 \n# Function to calculate the aromaticity of a protein sequence    \ndef get_aromaticity(row):\n    try:\n        return ProteinAnalysis(row['sequence']).aromaticity()\n    except:\n        return 0 \n# Function to calculate the instability index of a protein sequence    \ndef instability_index(row):\n    try:\n        return ProteinAnalysis(row['sequence']).instability_index()\n    except:\n        return 0 \n# Function to calculate the GRAVY score of a protein sequence    \ndef gravy(row):\n    try:\n        return ProteinAnalysis(row['sequence']).gravy()\n    except:\n        return 0\n# Function to calculate the isoelectric point of a protein sequence    \ndef isoelectric_point(row):\n    try:\n        return ProteinAnalysis(row['sequence']).isoelectric_point()\n    except:\n        return 0\n# Function to calculate the charge of a protein sequence at a specific pH    \ndef charge_at_pH(row,i):\n    try:\n        return ProteinAnalysis(row['sequence']).charge_at_pH(i)\n    except:\n        return 0 \n# Function to calculate the fraction of a specific secondary structure in a protein sequence    \ndef secondary_structure_fraction_0(row):\n    try:\n        return ProteinAnalysis(row['sequence']).secondary_structure_fraction()[0]\n    except:\n        return 0\ndef secondary_structure_fraction_1(row):\n    try:\n        return ProteinAnalysis(row['sequence']).secondary_structure_fraction()[1]\n    except:\n        return 0     \ndef secondary_structure_fraction_2(row):\n    try:\n        return ProteinAnalysis(row['sequence']).secondary_structure_fraction()[2]\n    except:\n        return 0 \n# Function to calculate the molar extinction coefficient of a protein sequence      \ndef molar_extinction_coefficient_0(row):\n    try:\n        return ProteinAnalysis(row['sequence']).molar_extinction_coefficient()[0]\n    except:\n        return 0\ndef molar_extinction_coefficient_1(row):\n    try:\n        return ProteinAnalysis(row['sequence']).molar_extinction_coefficient()[1]\n    except:\n        return 0\n      \ndef instability_index(row):\n    try:\n        return ProteinAnalysis(row['sequence']).instability_index()\n    except:\n        return 0\ndef isoelectric_point(row):\n    try:\n        return ProteinAnalysis(row['sequence']).isoelectric_point()\n    except:\n        return 0  \n    \n    \n    \n # Hydropathy scale for amino acids (Kyte-Doolittle)\nhydropathy_scale = {\n    'A': 1.8, 'R': -4.5, 'N': -3.5, 'D': -3.5, 'C': 2.5,\n    'Q': -3.5, 'E': -3.5, 'G': -0.4, 'H': -3.2, 'I': 4.5,\n    'L': 3.8, 'K': -3.9, 'M': 1.9, 'F': 2.8, 'P': -1.6,\n    'S': -0.8, 'T': -0.7, 'W': -0.9, 'Y': -1.3, 'V': 4.2\n}\n\n\n# #Function to calculate the hydropathy value for a protein sequence using the Kyte-Doolittle hydropathy scale\ndef calculate_hydropathy_value(row, hydropathy_scale):\n    hydropathy_scores = [hydropathy_scale[aa] if aa in hydropathy_scale else 0 for aa in row['sequence']]\n    non_zero_length = len([item for item in hydropathy_scores if item != 0])\n    hydropathy_value = sum(hydropathy_scores)/ non_zero_length\n    return hydropathy_value\n\n\n\n\ndef groups_amino_acids(row):\n    translation_dict = {\n        'F': 'I',  \n        'I': 'I',  \n        'L': 'I',  \n        'M': 'I',  \n        'V': 'I',\n        'D': 'E',  \n        'E': 'E',  \n        'H': 'E',  \n        'K': 'E',  \n        'N': 'E',\n        'Q': 'E',\n        'R': 'E',    \n        'S': 'A',  \n        'T': 'A',  \n        'Y': 'A',  \n        'C': 'A',  \n        'W': 'A',\n        'G': 'A',  \n        'P': 'A',  \n        'A': 'A'            \n    }\n\n    translated_sequence = ''.join(translation_dict[aa] if aa in translation_dict else aa for aa in row['sequence'])\n    return translated_sequence","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:04.812035Z","iopub.execute_input":"2023-08-15T04:37:04.812445Z","iopub.status.idle":"2023-08-15T04:37:04.842255Z","shell.execute_reply.started":"2023-08-15T04:37:04.812405Z","shell.execute_reply":"2023-08-15T04:37:04.840936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_features(df):\n    for acid in amino_acids:\n        df[acid+'_count'] = df.apply(amino_acids_count,acid=acid,axis=1)\n    for acid in amino_acids:\n        df[acid+'_percent'] = df.apply(amino_acids_percent,acid=acid,axis=1)\n    for i in range(15):\n        df['charge_at_pH_'+str(i)] = train_df.apply(charge_at_pH,i=i,axis=1)    \n    df['molecular_weight']=df.apply(get_molecular_weight,axis=1) \n    df['get_aromaticity']=df.apply(get_aromaticity,axis=1)\n    df['instability_index']=df.apply(instability_index,axis=1)\n    df['gravy']=df.apply(gravy,axis=1)\n    df['isoelectric_point']=df.apply(isoelectric_point,axis=1)\n    df['secondary_structure_fraction_0']=df.apply(secondary_structure_fraction_0,axis=1) \n    df['secondary_structure_fraction_1']=df.apply(secondary_structure_fraction_1,axis=1)  \n    df['secondary_structure_fraction_2']=df.apply(secondary_structure_fraction_2,axis=1)\n    df['molar_extinction_coefficient_0']=df.apply(molar_extinction_coefficient_0,axis=1) \n    df['molar_extinction_coefficient_1']=df.apply(molar_extinction_coefficient_1,axis=1)\n    df['instability_index']=df.apply(instability_index,axis=1)\n    df['isoelectric_point']=df.apply(isoelectric_point,axis=1)\n    df['hydropathy_value']=df.apply(calculate_hydropathy_value,hydropathy_scale=hydropathy_scale,axis=1)\n    df['groups_amino_acids']=df.apply(groups_amino_acids,axis=1)\n    ","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:04.84385Z","iopub.execute_input":"2023-08-15T04:37:04.844187Z","iopub.status.idle":"2023-08-15T04:37:04.862056Z","shell.execute_reply.started":"2023-08-15T04:37:04.844154Z","shell.execute_reply":"2023-08-15T04:37:04.860741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_features(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:04.863812Z","iopub.execute_input":"2023-08-15T04:37:04.86432Z","iopub.status.idle":"2023-08-15T04:37:52.323806Z","shell.execute_reply.started":"2023-08-15T04:37:04.864267Z","shell.execute_reply":"2023-08-15T04:37:52.320015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.324639Z","iopub.status.idle":"2023-08-15T04:37:52.325097Z","shell.execute_reply.started":"2023-08-15T04:37:52.324877Z","shell.execute_reply":"2023-08-15T04:37:52.3249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_features(test_df)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.327408Z","iopub.status.idle":"2023-08-15T04:37:52.328517Z","shell.execute_reply.started":"2023-08-15T04:37:52.328062Z","shell.execute_reply":"2023-08-15T04:37:52.328102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.33019Z","iopub.status.idle":"2023-08-15T04:37:52.331142Z","shell.execute_reply.started":"2023-08-15T04:37:52.330804Z","shell.execute_reply":"2023-08-15T04:37:52.330839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(test_df['taxonomyID'].unique())","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.334081Z","iopub.status.idle":"2023-08-15T04:37:52.335116Z","shell.execute_reply.started":"2023-08-15T04:37:52.334728Z","shell.execute_reply":"2023-08-15T04:37:52.334775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#change missing value in train\ntrain_df.loc[ train_df[\"EntryID\"] == \"Q5QP01\",\"Species\"] = \"Homo\"\ntrain_df.loc[ train_df[\"EntryID\"] == \"B0QYD3\",\"Species\"] = \"Homo\"\ntrain_df.loc[ train_df[\"EntryID\"] == \"A0A087WZI3\",\"Species\"] = \"Homo\"","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.336781Z","iopub.status.idle":"2023-08-15T04:37:52.338565Z","shell.execute_reply.started":"2023-08-15T04:37:52.338326Z","shell.execute_reply":"2023-08-15T04:37:52.338353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenate the 'length' columns from train_df and test_df into a single DataFrame\nconcatenated_data=pd.concat([train_df['length'],test_df['length']]).to_frame()","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.339863Z","iopub.status.idle":"2023-08-15T04:37:52.341008Z","shell.execute_reply.started":"2023-08-15T04:37:52.340765Z","shell.execute_reply":"2023-08-15T04:37:52.340797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create an array of evenly spaced bin edges from 0 to 2817 with 20 bins\nbins = np.linspace(0, 2817, 20)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.342089Z","iopub.status.idle":"2023-08-15T04:37:52.342587Z","shell.execute_reply.started":"2023-08-15T04:37:52.342308Z","shell.execute_reply":"2023-08-15T04:37:52.342332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a new column 'length_cut' in the concatenated_data DataFrame by binning the 'length' values\n# using the specified bins and assigning labels as False\nconcatenated_data[\"length_cut\"] = pd.cut(concatenated_data[\"length\"], bins=bins, labels=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.344285Z","iopub.status.idle":"2023-08-15T04:37:52.344697Z","shell.execute_reply.started":"2023-08-15T04:37:52.34449Z","shell.execute_reply":"2023-08-15T04:37:52.344511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Assign bin index 19 to values in the 'length_cut' column where the 'length' is greater than 2816\nconcatenated_data.loc[concatenated_data.length>2816,'length_cut']=19","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.346127Z","iopub.status.idle":"2023-08-15T04:37:52.346538Z","shell.execute_reply.started":"2023-08-15T04:37:52.346324Z","shell.execute_reply":"2023-08-15T04:37:52.346352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"concatenated_data.drop_duplicates(subset=['length'],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.347786Z","iopub.status.idle":"2023-08-15T04:37:52.348202Z","shell.execute_reply.started":"2023-08-15T04:37:52.347981Z","shell.execute_reply":"2023-08-15T04:37:52.348002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = pd.merge(test_df,concatenated_data,how=\"left\",on=\"length\")\ntrain_df = pd.merge(train_df,concatenated_data,how=\"left\",on=\"length\")","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.349796Z","iopub.status.idle":"2023-08-15T04:37:52.350203Z","shell.execute_reply.started":"2023-08-15T04:37:52.349994Z","shell.execute_reply":"2023-08-15T04:37:52.350015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.351391Z","iopub.status.idle":"2023-08-15T04:37:52.352069Z","shell.execute_reply.started":"2023-08-15T04:37:52.351844Z","shell.execute_reply":"2023-08-15T04:37:52.351868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"Extract Tazonamy data ueing API","metadata":{}},{"cell_type":"code","source":"# Loop through unique taxonomy IDs in train_df and retrieve taxonomic information\n#Train Data\nimport tqdm\nfrom tqdm import tqdm as progress_bar\nEntrez.email = 'haryan80@gmail.com'\nrow_list = []\ndict1={}\nfor i in progress_bar(train_df['taxonomyID'].unique()):\n    if i=='':continue\n    counter = 0\n    while counter < 10:\n        try:\n            handle = Entrez.efetch(db='taxonomy', id=int(i), retmode='xml')\n            record = Entrez.read(handle)\n            dict1['taxonomyID']=int(i)\n            dict1['GeneticCode']=record[0]['GeneticCode']['GCId']\n            dict1['MitoGeneticCode']=record[0]['MitoGeneticCode']['MGCId']\n            for data in record[0]['LineageEx']:\n                  if data['Rank'] in ['superkingdom','kingdom','phylum','class','order','family','genus']:\n                      dict1[data['Rank']] = data['ScientificName']\n            break        \n        except:\n            counter += 1\n            print(f\"Exception raised, retrying...{i}\")\n            if counter == 10:\n            # Stop retrying\n                print(\"Giving up\")    \n    row_list.append(dict1)\n    dict1={}\n    \ntaxonomy_train_df = pd.DataFrame(row_list)\ntaxonomy_train_df.head()  ","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.353327Z","iopub.status.idle":"2023-08-15T04:37:52.353783Z","shell.execute_reply.started":"2023-08-15T04:37:52.353538Z","shell.execute_reply":"2023-08-15T04:37:52.35356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"taxonomy_train_df.to_csv('taxonomy_train_df.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.354845Z","iopub.status.idle":"2023-08-15T04:37:52.355249Z","shell.execute_reply.started":"2023-08-15T04:37:52.355039Z","shell.execute_reply":"2023-08-15T04:37:52.35506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"taxonomy_train_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.356829Z","iopub.status.idle":"2023-08-15T04:37:52.357598Z","shell.execute_reply.started":"2023-08-15T04:37:52.357371Z","shell.execute_reply":"2023-08-15T04:37:52.357397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Extract Tazonamy data ueing API","metadata":{}},{"cell_type":"code","source":"# Loop through unique taxonomy IDs in train_df and retrieve taxonomic information\n#Test Data\nrow_list = []\ndict1={}\nfor i in progress_bar(test_df['taxonomyID'].unique()):\n    if i=='':continue\n    counter = 0\n    while counter < 10:\n        try:\n            handle = Entrez.efetch(db='taxonomy', id=int(i), retmode='xml')\n            record = Entrez.read(handle)\n            dict1['taxonomyID']=int(i)\n            dict1['GeneticCode']=record[0]['GeneticCode']['GCId']\n            dict1['MitoGeneticCode']=record[0]['MitoGeneticCode']['MGCId']\n            for data in record[0]['LineageEx']:\n                  if data['Rank'] in ['superkingdom','kingdom','phylum','class','order','family','genus']:\n                      dict1[data['Rank']] = data['ScientificName']\n            break        \n        except:\n            counter += 1\n            print(f\"Exception raised, retrying...{i}\")\n            if counter == 10:\n            # Stop retrying\n                print(\"Giving up\")    \n    row_list.append(dict1)\n    dict1={}\n    \ntaxonomy_test_df = pd.DataFrame(row_list)\ntaxonomy_test_df.head()   ","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.358637Z","iopub.status.idle":"2023-08-15T04:37:52.359386Z","shell.execute_reply.started":"2023-08-15T04:37:52.359147Z","shell.execute_reply":"2023-08-15T04:37:52.359172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"taxonomy_test_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.360544Z","iopub.status.idle":"2023-08-15T04:37:52.360973Z","shell.execute_reply.started":"2023-08-15T04:37:52.360757Z","shell.execute_reply":"2023-08-15T04:37:52.360779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"taxonomy_test_df.to_csv('taxonomy_test_df.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.362189Z","iopub.status.idle":"2023-08-15T04:37:52.362879Z","shell.execute_reply.started":"2023-08-15T04:37:52.362641Z","shell.execute_reply":"2023-08-15T04:37:52.362667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.364094Z","iopub.status.idle":"2023-08-15T04:37:52.364503Z","shell.execute_reply.started":"2023-08-15T04:37:52.364288Z","shell.execute_reply":"2023-08-15T04:37:52.364309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['taxonomyID'] = train_df['taxonomyID'].astype(str)\ntest_df['taxonomyID'] = test_df['taxonomyID'].astype(str)\ntaxonomy_train_df['taxonomyID'] = taxonomy_train_df['taxonomyID'].astype(str)\ntaxonomy_test_df['taxonomyID'] = taxonomy_test_df['taxonomyID'].astype(str)","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.366262Z","iopub.status.idle":"2023-08-15T04:37:52.366724Z","shell.execute_reply.started":"2023-08-15T04:37:52.366482Z","shell.execute_reply":"2023-08-15T04:37:52.366504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.merge(train_df,taxonomy_train_df,how=\"left\", on='taxonomyID')\ntest_df = pd.merge(test_df,taxonomy_test_df,how=\"left\", on='taxonomyID')","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.36808Z","iopub.status.idle":"2023-08-15T04:37:52.368495Z","shell.execute_reply.started":"2023-08-15T04:37:52.36828Z","shell.execute_reply":"2023-08-15T04:37:52.368302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.to_csv('train_df.csv')\ntest_df.to_csv('test_df.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-15T04:37:52.370451Z","iopub.status.idle":"2023-08-15T04:37:52.370935Z","shell.execute_reply.started":"2023-08-15T04:37:52.370679Z","shell.execute_reply":"2023-08-15T04:37:52.370704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}