{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Transfer Learning \n\nDeep Learning and Pre-trained Models: Train a deep learning model (e.g., neural network) on the homo sapiens data using a large dataset and then fine-tune the model on the other species' data. Pre-trained models, such as those trained on large protein databases, can provide valuable features that can be fine-tuned for specific tasks.\n\nDomain Adaptation: Techniques from domain adaptation, which aim to align the feature distributions across domains, can be useful in transfer learning settings when dealing with data from different species. These methods help in reducing the domain shift and improving the model's generalization to the target species.\n\n**Few-shot Learning: As we have a limited amount of data for the target species, few-shot learning techniques might be beneficial. These methods aim to learn from a small amount of data, making them suitable for the transfer learning scenario.**\n\nRegularization Techniques: Use regularization methods to prevent overfitting on the limited target species data while transferring knowledge from the source species.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib as plt","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:29:04.673151Z","iopub.execute_input":"2023-08-03T09:29:04.673673Z","iopub.status.idle":"2023-08-03T09:29:04.680416Z","shell.execute_reply.started":"2023-08-03T09:29:04.673629Z","shell.execute_reply":"2023-08-03T09:29:04.679096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms_df = pd.read_csv('/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv', sep='\\t')\ntrain_terms_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:29:06.347784Z","iopub.execute_input":"2023-08-03T09:29:06.348292Z","iopub.status.idle":"2023-08-03T09:29:10.460545Z","shell.execute_reply.started":"2023-08-03T09:29:06.348254Z","shell.execute_reply":"2023-08-03T09:29:10.458466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_tax_df = pd.read_csv('/kaggle/input/cafa-5-protein-function-prediction/Train/train_taxonomy.tsv', sep='\\t')\ntrain_tax_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:29:13.127207Z","iopub.execute_input":"2023-08-03T09:29:13.127693Z","iopub.status.idle":"2023-08-03T09:29:13.292011Z","shell.execute_reply.started":"2023-08-03T09:29:13.127654Z","shell.execute_reply":"2023-08-03T09:29:13.290146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# number of unique entries for taxonomy\nunique_tax_ids = train_tax_df['taxonomyID'].nunique()\nprint(f\"Unique TaxIDs: {unique_tax_ids}\")\n\n# top 10 species based on occurence count\ntop_ten_species = train_tax_df[\"taxonomyID\"].value_counts().nlargest(10)\nprint(f\"Top 10 species: {top_ten_species}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:29:15.371652Z","iopub.execute_input":"2023-08-03T09:29:15.372093Z","iopub.status.idle":"2023-08-03T09:29:15.405027Z","shell.execute_reply.started":"2023-08-03T09:29:15.37206Z","shell.execute_reply":"2023-08-03T09:29:15.403792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## PART 1: Create baseline model using homo sapiens data\n\n9606 is homo sapiens. It has the highest count in our dataset. We will create a new dataframe with it and use it to train a deep learning model. We can then fine-tune the model on the top 4 most common species. In order to do this, we need to:\n1. Merge train_terms_df and train_tax_df\n2. Create a new dataframe homo_sapiens_df\n3. Create homo_sapiens.fasta protein sequence using 'ox= homo sapiens' from train_sequence.fasta","metadata":{}},{"cell_type":"code","source":"combined_df = pd.merge(train_terms_df, train_tax_df, on='EntryID', how='inner')\ncombined_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:29:19.705611Z","iopub.execute_input":"2023-08-03T09:29:19.706248Z","iopub.status.idle":"2023-08-03T09:29:21.774671Z","shell.execute_reply.started":"2023-08-03T09:29:19.706193Z","shell.execute_reply":"2023-08-03T09:29:21.773393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filter_id = 9606\nhomo_sapiens_df = combined_df[combined_df['taxonomyID'] == filter_id]\n\nhomo_sapiens_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:29:23.65956Z","iopub.execute_input":"2023-08-03T09:29:23.661253Z","iopub.status.idle":"2023-08-03T09:29:23.789184Z","shell.execute_reply.started":"2023-08-03T09:29:23.661186Z","shell.execute_reply":"2023-08-03T09:29:23.787806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# save the new dataframe as a tsv file for download and reupload\n\nhomo_sapiens_tsv = '/kaggle/working/homo_sapiens.tsv'\nhomo_sapiens_df.to_csv(homo_sapiens_tsv, sep='\\t', index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T05:49:36.32651Z","iopub.execute_input":"2023-08-03T05:49:36.327157Z","iopub.status.idle":"2023-08-03T05:49:38.049147Z","shell.execute_reply.started":"2023-08-03T05:49:36.327123Z","shell.execute_reply":"2023-08-03T05:49:38.047558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_entries = homo_sapiens_df['EntryID'].nunique()\nprint(f\"NO. unique EntryID: {unique_entries}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:30:57.676806Z","iopub.execute_input":"2023-08-03T09:30:57.677299Z","iopub.status.idle":"2023-08-03T09:30:57.813801Z","shell.execute_reply.started":"2023-08-03T09:30:57.677265Z","shell.execute_reply":"2023-08-03T09:30:57.812128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## FASTA FILE","metadata":{}},{"cell_type":"code","source":"from Bio import SeqIO\n\nfasta_file = '/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta'\n\nnum_records = 5\nfasta_records = []\n\nwith open(fasta_file, \"r\") as handle:\n    for i, record in enumerate(SeqIO.parse(handle, 'fasta')):\n        fasta_records.append(record)\n        if i + 1 == num_records:\n            break\nprint(fasta_records)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:31:04.842761Z","iopub.execute_input":"2023-08-03T09:31:04.843345Z","iopub.status.idle":"2023-08-03T09:31:05.074Z","shell.execute_reply.started":"2023-08-03T09:31:04.843289Z","shell.execute_reply":"2023-08-03T09:31:05.072015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_taxonomy_ids_from_fasta(fasta_file):\n    taxonomy_ids = set()\n    for record in SeqIO.parse(fasta_file, \"fasta\"):\n        description = record.description\n        # Extract the taxonomyID (OX) from the description\n        start_index = description.find('OX=')\n        end_index = description.find(' ', start_index)\n        if start_index != -1 and end_index != -1:\n            taxonomy_id = description[start_index + 3:end_index]\n            taxonomy_ids.add(taxonomy_id)\n    return taxonomy_ids\n\ndef get_top_ten_species(fasta_file):\n    taxonomy_ids = get_taxonomy_ids_from_fasta(fasta_file)\n    num_unique_taxonomy_ids = len(taxonomy_ids)\n    print(f\"Number of unique taxonomyIDs: {num_unique_taxonomy_ids}\")\n\n    # Count occurrences of each taxonomyID and get the top ten species\n    top_ten_species = {}\n    for record in SeqIO.parse(fasta_file, \"fasta\"):\n        description = record.description\n        start_index = description.find('OX=')\n        end_index = description.find(' ', start_index)\n        if start_index != -1 and end_index != -1:\n            taxonomy_id = description[start_index + 3:end_index]\n            if taxonomy_id in taxonomy_ids:\n                top_ten_species[taxonomy_id] = top_ten_species.get(taxonomy_id, 0) + 1\n\n    sorted_species = sorted(top_ten_species.items(), key=lambda x: x[1], reverse=True)\n    print(\"Top ten species:\")\n    for species, count in sorted_species[:10]:\n        print(f\"TaxonomyID: {species}, Count: {count}\")\n\n# Replace 'train_sequences.fasta' with the path to your actual fasta file.\nfasta_file_path = '/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta'\nget_top_ten_species(fasta_file_path)\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:31:18.402861Z","iopub.execute_input":"2023-08-03T09:31:18.403287Z","iopub.status.idle":"2023-08-03T09:31:23.845058Z","shell.execute_reply.started":"2023-08-03T09:31:18.403255Z","shell.execute_reply":"2023-08-03T09:31:23.843723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From a preliminary look at the fasta file, we seem to have more homo sapiens entries in homo_sapiens_df (25125). We will need to drop the extra entries before we pass the sequence through the ProtBert embeddings.","metadata":{}},{"cell_type":"code","source":"# Create homo_sapiens.fasta and reupload so we don't have to run the code every single time\n\ndef get_entry_ids_from_fasta(fasta_file):\n    entry_ids = set()\n    for record in SeqIO.parse(fasta_file, \"fasta\"):\n        entry_ids.add(record.id)\n    return entry_ids\n\ndef create_new_fasta_with_homo_sapiens_sequences(input_fasta, output_fasta, homo_sapiens_entry_ids):\n    with open(output_fasta, \"w\") as output_handle:\n        for record in SeqIO.parse(input_fasta, \"fasta\"):\n            if record.id in homo_sapiens_entry_ids:\n                SeqIO.write(record, output_handle, \"fasta\")\n\n# Assuming you already have the 'homo_sapiens_df' DataFrame and 'train_sequences.fasta' file.\n\n# Step 1: Get the EntryIDs from 'homo_sapiens_df'\nhomo_sapiens_entry_ids = set(homo_sapiens_df['EntryID'])\n\n# Step 2: Create a new fasta file containing only Homo sapiens protein sequences\noutput_fasta_path = 'output_homo_sapiens.fasta'\ncreate_new_fasta_with_homo_sapiens_sequences('/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta', output_fasta_path, homo_sapiens_entry_ids)\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T03:45:23.836466Z","iopub.status.idle":"2023-08-03T03:45:23.836932Z","shell.execute_reply.started":"2023-08-03T03:45:23.836707Z","shell.execute_reply":"2023-08-03T03:45:23.836728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# I've saved and reuploaded the new fasta file\n\nfasta_file = '/kaggle/input/homo-sapiens-sequence/homo_sapiens_sequence.fasta'\n\nnum_records = 5\nfasta_records = []\n\nwith open(fasta_file, \"r\") as handle:\n    for i, record in enumerate(SeqIO.parse(handle, 'fasta')):\n        fasta_records.append(record)\n        if i + 1 == num_records:\n            break\nprint(fasta_records)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:32:17.866532Z","iopub.execute_input":"2023-08-03T09:32:17.867004Z","iopub.status.idle":"2023-08-03T09:32:17.882248Z","shell.execute_reply.started":"2023-08-03T09:32:17.86697Z","shell.execute_reply":"2023-08-03T09:32:17.881015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Step 1: Get the unique EntryIDs from '/kaggle/input/homo-sapiens-sequence/homo_sapiens_sequence.fasta'\n\ndef get_unique_entry_ids_from_fasta(fasta_file):\n    entry_ids = set()\n    for record in SeqIO.parse(fasta_file, \"fasta\"):\n        entry_ids.add(record.id)\n    return entry_ids\n\noutput_fasta_path = '/kaggle/input/homo-sapiens-sequence/homo_sapiens_sequence.fasta'\nunique_entry_ids = get_unique_entry_ids_from_fasta(output_fasta_path)\n\n# Step 2: Count the number of unique EntryIDs\nnum_unique_entries = len(unique_entry_ids)\nprint(f\"Number of unique EntryIDs in 'output_homo_sapiens.fasta': {num_unique_entries}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:33:34.772925Z","iopub.execute_input":"2023-08-03T09:33:34.773451Z","iopub.status.idle":"2023-08-03T09:33:35.267147Z","shell.execute_reply.started":"2023-08-03T09:33:34.773412Z","shell.execute_reply":"2023-08-03T09:33:35.266096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Weird. Now we do have matching EntryIds. Let's move ahead with [ProtBert embeddings.](https://www.kaggle.com/code/suzzystranger/cafa-subset-exploration-part-1)","metadata":{}}]}