{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# CAFA 5 protein function Prediction with TensorFlow + GPU\n​\n**This is the original starter notebook engineered for GPU**\n\nThis notebook walks you through how to train a DNN model using TensorFlow on the CAFA 5 protein function Prediction dataset made available for this competition. \n​\nThe objective of the model is to predict the function(aka **GO term ID**) of a set of proteins based on their amino acid sequences and other data.\n​\n​\n**Note** : This notebook runs without any GPU. This is because enabling GPUs leaves less RAM memory on the VM and the submission step needs a lot of memory. One point where this would impact is when training the model. With CPU it will take around 2 minutes while on GPU it would take around 30 seconds.","metadata":{"execution":{"iopub.status.busy":"2023-07-24T19:11:49.693604Z","iopub.execute_input":"2023-07-24T19:11:49.694647Z","iopub.status.idle":"2023-07-24T19:11:57.85132Z","shell.execute_reply.started":"2023-07-24T19:11:49.694593Z","shell.execute_reply":"2023-07-24T19:11:57.850304Z"}}},{"cell_type":"markdown","source":"## About the Data\n\n### Protein Sequence\n\nEach protein is composed of dozens or hundreds of amino acids that are linked sequentially. Each amino acid in the sequence may be represented by a one-letter or three-letter code. Thus the sequence of a protein is often notated as a string of letters. \n\n<img src=\"https://cityu-bioinformatics.netlify.app/img/tools/protein/pro_seq.png\" alt =\"Sequence.png\" style='width: 800px;' >\n\nImage source - [https://cityu-bioinformatics.netlify.app/](https://cityu-bioinformatics.netlify.app/too2/new_proteo/pro_seq/)\n\nThe `train_sequences.fasta` made available for this competitions, contains the sequences for proteins with annotations (labelled proteins).","metadata":{}},{"cell_type":"markdown","source":"# Gene Ontology\n\nWe can define the functional properties of a proteins using Gene Ontology(GO). Gene Ontology (GO) describes our understanding of the biological domain with respect to three aspects:\n1. Molecular Function (MF)\n2. Biological Process (BP)\n3. Cellular Component (CC)\n\nRead more about Gene Ontology [here](http://geneontology.org/docs/ontology-documentation).\n\nFile `train_terms.tsv` contains the list of annotated terms (ground truth) for the proteins in `train_sequences.fasta`. In `train_terms.tsv` the first column indicates the protein's UniProt accession ID (unique protein id), the second is the `GO Term ID`, and the third indicates in which ontology the term appears. ","metadata":{}},{"cell_type":"markdown","source":"# Labels of the dataset\n\nThe objective of our model is to predict the terms (functions) of a protein sequence. One protein sequence can have many functions and can thus be classified into any number of terms. Each term is uniquely identified by a `GO Term ID`. Thus our model has to predict all the `GO Term ID`s for a protein sequence. This means that the task at hand is a multi-label classification problem. ","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# Protein embeddings for train and test data\n\nTo train a machine learning model we cannot use the alphabetical protein sequences in`train_sequences.fasta` directly. They have to be converted into a vector format. In this notebook, we will use embeddings of the protein sequences to train the model. You can think of protein embeddings to be similar to word embeddings used to train NLP models.\n<!-- Instead, to make calculations and data preparation easier we will use precalculated protein embeddings.\n -->\nProtein embeddings are a machine-friendly method of capturing the protein's structural and functional characteristics, mainly through its sequence. One approach is to train a custom ML model to learn the protein embeddings of the protein sequences in the dataset being used in this notebook. Since this dataset represents proteins using amino-acid sequences which is a standard approach, we can use any publicly available pre-trained protein embedding models to generate the embeddings.\n\nThere are a variety of protein embedding models. To make data preparation easier, we have used the precalculated protein embeddings created by [Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model in this notebook. The precalculated protein embeddings can be found [here](https://www.kaggle.com/datasets/sergeifironov/t5embeds). We have added this dataset to the notebook along with the dataset made available for the competition.\n\nTo add this to your enviroment, on the right side panel, click on `Add Data` and search for `t5embeds` (make sure that it's the correct [one](https://www.kaggle.com/datasets/sergeifironov/t5embeds)) and then click on the `+` beside it.\n","metadata":{}},{"cell_type":"markdown","source":"# Import the Required Libraries","metadata":{}},{"cell_type":"code","source":"import os\nos.environ['TF_CPP_MIN_LOG_LEVEL'] = '2'\nimport tensorflow as tf, tensorflow.keras as keras\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\n%matplotlib inline\nimport matplotlib.pyplot as plt\nimport plotly.express as px\n\nimport gc\nimport time\n# Required for progressbar widget\nimport progressbar\n%load_ext memory_profiler\n\nimport sys\n\nimport cupy as cp\nimport pandas as pd\n\n\nimport cudf\n\ncp.random.seed(0)\n#https://www.tensorflow.org/api_docs/python/tf/keras/utils/set_random_seed\ntf.keras.utils.set_random_seed(0)\n\n\n# This helps with GPU memory utilization.\nconfig = tf.compat.v1.ConfigProto(allow_soft_placement=True, log_device_placement=True)\nconfig.gpu_options.allow_growth = True\nsess = tf.compat.v1.Session(config=config)\n\nclass Timer():\n    '''\n    Time the notebook runtime.\n    '''\n    def __init__(self, lim:'RunTimeLimit'=1000): self.t0, self.lim, _ = time.time(), lim, print(f'started. You have {lim} sec. Good luck!')\n    def ShowTime(self):\n        msg = f'Runtime is {time.time()-self.t0:.0f} sec'\n        print(f'\\033[91m\\033[1m' + msg + f' > {self.lim} sec limit!!!\\033[0m' if (time.time()-self.t0-1) > self.lim else msg)\n\n    \n    \n\ndef sizeof_fmt(num, suffix='B'):\n    ''' by Fred Cirera,  https://stackoverflow.com/a/1094933/1870254, modified'''\n    for unit in ['','Ki','Mi','Gi','Ti','Pi','Ei','Zi']:\n        if abs(num) < 1024.0:\n            return \"%3.1f %s%s\" % (num, unit, suffix)\n        num /= 1024.0\n    return \"%.1f %s%s\" % (num, 'Yi', suffix)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:41:58.103768Z","iopub.execute_input":"2023-07-26T16:41:58.104148Z","iopub.status.idle":"2023-07-26T16:42:16.137237Z","shell.execute_reply.started":"2023-07-26T16:41:58.104118Z","shell.execute_reply":"2023-07-26T16:42:16.136215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:16.139415Z","iopub.execute_input":"2023-07-26T16:42:16.140051Z","iopub.status.idle":"2023-07-26T16:42:17.171794Z","shell.execute_reply.started":"2023-07-26T16:42:16.139996Z","shell.execute_reply":"2023-07-26T16:42:17.170654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('GPU name: ', tf.config.experimental.list_physical_devices('GPU')[0])","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:17.174081Z","iopub.execute_input":"2023-07-26T16:42:17.174514Z","iopub.status.idle":"2023-07-26T16:42:17.182888Z","shell.execute_reply.started":"2023-07-26T16:42:17.174474Z","shell.execute_reply":"2023-07-26T16:42:17.181983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"TensorFlow v\" + tf.__version__)\nprint(\"Numpy v\" + np.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:17.185782Z","iopub.execute_input":"2023-07-26T16:42:17.186374Z","iopub.status.idle":"2023-07-26T16:42:17.195279Z","shell.execute_reply.started":"2023-07-26T16:42:17.186341Z","shell.execute_reply":"2023-07-26T16:42:17.194266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmr = Timer()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:17.19703Z","iopub.execute_input":"2023-07-26T16:42:17.197774Z","iopub.status.idle":"2023-07-26T16:42:17.20534Z","shell.execute_reply.started":"2023-07-26T16:42:17.197741Z","shell.execute_reply":"2023-07-26T16:42:17.204328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the Dataset","metadata":{}},{"cell_type":"markdown","source":"First we will load the file `train_terms.tsv` which contains the list of annotated terms (functions) for the proteins. We will extract the labels aka `GO term ID` and create a label dataframe for the protein embeddings.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n# https://docs.rapids.ai/api/cudf/nightly/user_guide/10min/\ntrain_terms = cudf.read_csv(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv\",sep=\"\\t\")","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:17.20689Z","iopub.execute_input":"2023-07-26T16:42:17.207626Z","iopub.status.idle":"2023-07-26T16:42:19.554958Z","shell.execute_reply.started":"2023-07-26T16:42:17.207551Z","shell.execute_reply":"2023-07-26T16:42:19.553855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%memit\n# Get the shape of the DataFrame\nshape = train_terms.shape\n\n# Extract the number of rows and columns\nnum_rows = shape[0]\nnum_columns = shape[1]\n\n# Print the shape information\nprint(\"Number of rows:\", num_rows)\nprint(\"Number of columns:\", num_columns)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:19.557464Z","iopub.execute_input":"2023-07-26T16:42:19.55825Z","iopub.status.idle":"2023-07-26T16:42:20.039595Z","shell.execute_reply.started":"2023-07-26T16:42:19.558213Z","shell.execute_reply":"2023-07-26T16:42:20.038435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"`train_terms` dataframe is composed of 3 columns and 5363863 entries. We can see all 3 dimensions of our dataset by printing out the first 5 entries using the following code:","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\ntrain_terms.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:20.041584Z","iopub.execute_input":"2023-07-26T16:42:20.04196Z","iopub.status.idle":"2023-07-26T16:42:20.532434Z","shell.execute_reply.started":"2023-07-26T16:42:20.041914Z","shell.execute_reply":"2023-07-26T16:42:20.531185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we look at the first entry of `train_terms.tsv`, we can see that it contains protein id(`A0A009IHW8`), the GO term(`GO:0008152`) and its aspect(`BPO`). ","metadata":{}},{"cell_type":"markdown","source":"# Loading the protein embeddings\n\n\nWe will now load the pre calculated protein embeddings created by [Sergei Fironov](https://www.kaggle.com/sergeifironov) using the Rost Lab's T5 protein language model.\n\nIf the `tfembeds` is not yet on the input data of the notebook, you can add it to your enviromentby clicking on `Add Data` and search for `t5embeds` (make sure that it's the correct [one](https://www.kaggle.com/datasets/sergeifironov/t5embeds) ) and then click on the `+` beside it.\n\nThe protein embeddings to be used for training are recorded in `train_embeds.npy` and the corresponding protein ids are available in `train_ids.npy`.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\ntrain_protein_ids = np.load('/kaggle/input/t5embeds/t5embeds//train_ids.npy')\ntrain_embeddings = np.load('/kaggle/input/t5embeds/t5embeds/train_embeds.npy')\ndisplay(train_protein_ids[:5])","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:20.534489Z","iopub.execute_input":"2023-07-26T16:42:20.534861Z","iopub.status.idle":"2023-07-26T16:42:31.490477Z","shell.execute_reply.started":"2023-07-26T16:42:20.534823Z","shell.execute_reply":"2023-07-26T16:42:31.489278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<!-- Now, we will load`train_embeds.py` (did that above) which contains the pre-calculated embeddings of the proteins in the train dataset. with protein_ids (`id`s we loaded previously from the **train_ids.npy**) into a numpy array. This array now contains the precalculated embeddings for the protein_ids( Ids we loaded above from **train_ids.npy**) needed for training. -->\n\nAfter loading the files as numpy arrays, we will convert them into Pandas dataframe.\n\nEach protein embedding is a vector of length 1024. We create the resulting dataframe such that there are 1024 columns to represent the values in each of the 1024 places in the vector.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\ncolumn_num = train_embeddings.shape[1]\n\ntrain_df =     cudf.DataFrame.from_pandas(\\\n                                          pd.DataFrame(train_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\\\n                                          .astype('float32'))","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:31.495771Z","iopub.execute_input":"2023-07-26T16:42:31.496104Z","iopub.status.idle":"2023-07-26T16:42:36.734795Z","shell.execute_reply.started":"2023-07-26T16:42:31.496076Z","shell.execute_reply":"2023-07-26T16:42:36.733648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `train_df` dataframe which contains the embeddings is composed of 1024 columns and 142246 entries. We can see all 1024 dimensions(results will be truncated since column length is too long)  of our dataset by printing out the first 5 entries using the following code:","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:36.736681Z","iopub.execute_input":"2023-07-26T16:42:36.737121Z","iopub.status.idle":"2023-07-26T16:42:37.39326Z","shell.execute_reply.started":"2023-07-26T16:42:36.737073Z","shell.execute_reply":"2023-07-26T16:42:37.392076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare the dataset\n\nReference: https://www.kaggle.com/code/alexandervc/baseline-multilabel-to-multitarget-binary\n\nFirst we will extract all the needed labels(`GO term ID`) from `train_terms.tsv` file. There are more than 40,000 labels. In order to simplify our model, we will choose the most frequent 1500 `GO term ID`s as labels.\n\nLet's plot the most frequent 100 `GO Term ID`s in `train_terms.tsv`.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n# Select first 100 values for plotting\nplot_df = train_terms['term'].value_counts()[:100].to_pandas()\nfigure, axis = plt.subplots(1, 1, figsize=(12, 6))\nbp = sns.barplot(ax=axis, x=np.array(plot_df.index), y=plot_df.values)\nbp.set_xticklabels(bp.get_xticklabels(), rotation=90, size = 6)\naxis.set_title('Top 100 frequent GO term IDs')\nbp.set_xlabel(\"GO term IDs\", fontsize = 12)\nbp.set_ylabel(\"Count\", fontsize = 12)\nplt.show()\ndel plot_df\n\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:37.395339Z","iopub.execute_input":"2023-07-26T16:42:37.395755Z","iopub.status.idle":"2023-07-26T16:42:39.641945Z","shell.execute_reply.started":"2023-07-26T16:42:37.395718Z","shell.execute_reply":"2023-07-26T16:42:39.640765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will now save the first 1500 most frequent GO term Ids into a list.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n# Set the limit for label\nnum_of_labels = 1500\n\n# Take value counts in descending order and fetch first 1500 'GO term ID' as labels\nlabels = train_terms['term'].value_counts().to_frame().index[:num_of_labels].to_arrow().to_pylist()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:39.643524Z","iopub.execute_input":"2023-07-26T16:42:39.644124Z","iopub.status.idle":"2023-07-26T16:42:40.144809Z","shell.execute_reply.started":"2023-07-26T16:42:39.644087Z","shell.execute_reply":"2023-07-26T16:42:40.142902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, we will create a new dataframe by filtering the train terms with the selected `GO Term ID`s.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n# Fetch the train_terms data for the relevant labels only\ntrain_terms_updated = train_terms.loc[train_terms['term'].isin(labels)]","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:40.146744Z","iopub.execute_input":"2023-07-26T16:42:40.147155Z","iopub.status.idle":"2023-07-26T16:42:40.783957Z","shell.execute_reply.started":"2023-07-26T16:42:40.147116Z","shell.execute_reply":"2023-07-26T16:42:40.782721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms.info()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:40.785861Z","iopub.execute_input":"2023-07-26T16:42:40.786247Z","iopub.status.idle":"2023-07-26T16:42:40.802604Z","shell.execute_reply.started":"2023-07-26T16:42:40.786211Z","shell.execute_reply":"2023-07-26T16:42:40.801139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us plot the aspect values in the new train_terms_updated dataframe using a pie chart.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\npie_df = train_terms_updated['aspect'].value_counts().to_pandas()\npalette_color = sns.color_palette('bright')\nplt.pie(pie_df.values, labels=np.array(pie_df.index), colors=palette_color, autopct='%.0f%%')\nplt.show()\ndel pie_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:40.805187Z","iopub.execute_input":"2023-07-26T16:42:40.805785Z","iopub.status.idle":"2023-07-26T16:42:41.987842Z","shell.execute_reply.started":"2023-07-26T16:42:40.805751Z","shell.execute_reply":"2023-07-26T16:42:41.985884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see, majority of the `GO term Id`s have BPO(Biological Process Ontology) as their aspect.\n\nSince this is a multi label classification problem, in the labels array we will denote the presence or absence of each Go Term Id for a protein id using a 1 or 0.\nFirst, we will create a numpy array `train_labels` of required size for the labels. To update the `train_labels` array with the appropriate values, we will loop through the label list.","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n# Setup progressbar settings.\n# This is strictly for aesthetic.\nbar = progressbar.ProgressBar(maxval=num_of_labels, \\\n    widgets=[progressbar.Bar('=', '[', ']'), ' ', progressbar.Percentage()])\nbar.start()\n# Create an empty dataframe of required size for storing the labels,\n# i.e, train_size x num_of_labels (142246 x 1500)\ntrain_size = train_protein_ids.shape[0] # len(X)\ntrain_labels = cp.zeros((train_size ,num_of_labels))\n\n# Convert from numpy to pandas series for better handling\nseries_train_protein_ids = cudf.Series(train_protein_ids)\n\n# Loop through each label\nfor i in range(num_of_labels):\n    # For each label, fetch the corresponding train_terms data\n    n_train_terms = train_terms_updated[train_terms_updated['term'] ==  labels[i]]\n    \n    # Fetch all the unique EntryId aka proteins related to the current label(GO term ID)\n    label_related_proteins = n_train_terms['EntryID'].unique()\n    \n    # In the series_train_protein_ids pandas series, if a protein is related\n    # to the current label, then mark it as 1, else 0.\n    # Replace the ith column of train_Y with with that pandas series.\n    train_labels[:,i] =  series_train_protein_ids.isin(label_related_proteins).to_cupy().astype('float32')\n    \n    # Progress bar percentage increase\n    bar.update(i+1)\n\n# Notify the end of progress bar \nbar.finish()\n\n# Convert train_Y numpy into pandas dataframe\nlabels_df = cudf.DataFrame(data = train_labels, columns = labels).astype('float32')\nprint(labels_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:42:41.989406Z","iopub.execute_input":"2023-07-26T16:42:41.990597Z","iopub.status.idle":"2023-07-26T16:43:05.938244Z","shell.execute_reply.started":"2023-07-26T16:42:41.990561Z","shell.execute_reply":"2023-07-26T16:43:05.936896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The final labels dataframe (label_df) is composed of 1500 columns and 142246 entries. We can see all 1500 dimensions(results will be truncated since the number of columns is big) of our dataset by printing out the first 5 entries using the following code:","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:43:05.939921Z","iopub.execute_input":"2023-07-26T16:43:05.940632Z","iopub.status.idle":"2023-07-26T16:43:05.948601Z","shell.execute_reply.started":"2023-07-26T16:43:05.940593Z","shell.execute_reply":"2023-07-26T16:43:05.94768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:43:05.950078Z","iopub.execute_input":"2023-07-26T16:43:05.950663Z","iopub.status.idle":"2023-07-26T16:43:05.962896Z","shell.execute_reply.started":"2023-07-26T16:43:05.95063Z","shell.execute_reply":"2023-07-26T16:43:05.961494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<class 'cudf.core.dataframe.DataFrame'>\nRangeIndex: 142246 entries, 0 to 142245\nColumns: 1500 entries, GO:0005575 to GO:0070828\ndtypes: float64(1500)\nmemory usage: 1.6 GB","metadata":{}},{"cell_type":"code","source":"# Need to free up memory\ndel series_train_protein_ids\ndel train_protein_ids\ndel train_embeddings\ndel train_labels\ndel train_terms\ndel train_terms_updated\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:43:05.963988Z","iopub.execute_input":"2023-07-26T16:43:05.964586Z","iopub.status.idle":"2023-07-26T16:43:06.366627Z","shell.execute_reply.started":"2023-07-26T16:43:05.964547Z","shell.execute_reply":"2023-07-26T16:43:06.36542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training\n\nNext, we will use Tensorflow to train a Deep Neural Network with the protein embeddings.","metadata":{}},{"cell_type":"markdown","source":"## Insert Your algorithm Here.\n\nTheoretically we should be able to insert our algorithm here. ","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\nINPUT_SHAPE = [train_df.shape[1]]\nBATCH_SIZE = 5120\n\nmodel = tf.keras.Sequential([\n    tf.keras.layers.BatchNormalization(input_shape=INPUT_SHAPE),    \n    tf.keras.layers.Dense(units=512, activation='relu'),\n    tf.keras.layers.Dense(units=512, activation='relu'),\n    tf.keras.layers.Dense(units=512, activation='relu'),\n    tf.keras.layers.Dense(units=num_of_labels,activation='sigmoid')\n])\n\n\n# Compile model\nmodel.compile(\n    optimizer=tf.keras.optimizers.Adam(learning_rate=0.001),\n    loss='binary_crossentropy',\n    metrics=['binary_accuracy', tf.keras.metrics.AUC()],\n)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:43:06.368649Z","iopub.execute_input":"2023-07-26T16:43:06.36906Z","iopub.status.idle":"2023-07-26T16:43:07.077844Z","shell.execute_reply.started":"2023-07-26T16:43:06.369021Z","shell.execute_reply":"2023-07-26T16:43:07.076709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Transform from Cuda Dataframes to Numpy arrays for fitting the model","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n# Transform cuda dataframes to numpy arrays for processing.\ntrain_df2 = train_df.to_numpy()\nlabels_df2 = labels_df.to_numpy()\ndel train_df\ndel labels_df\n#\ngc.collect()\n\n# There are actually 212 million + training observations. The way this is formatted makes it \n# Difficult for other algorithms. \n\n## Can labels_df2 be combined into unique Y values to indicate one value targets?\n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:43:07.079601Z","iopub.execute_input":"2023-07-26T16:43:07.079965Z","iopub.status.idle":"2023-07-26T16:43:09.582673Z","shell.execute_reply.started":"2023-07-26T16:43:07.079924Z","shell.execute_reply":"2023-07-26T16:43:09.581147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now fit the model","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\nhistory = model.fit(\n    train_df, labels_df,\n    batch_size=BATCH_SIZE,\n    epochs=5\n)\n## Free up memory \ndel labels_df2\ndel train_df2\ngc.collect()\ntime.sleep(5)","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:43:09.584301Z","iopub.execute_input":"2023-07-26T16:43:09.584968Z","iopub.status.idle":"2023-07-26T16:44:01.523267Z","shell.execute_reply.started":"2023-07-26T16:43:09.584932Z","shell.execute_reply":"2023-07-26T16:44:01.522268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot the model's loss and accuracy for each epoch","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\nhistory_df = pd.DataFrame(history.history)\nhistory_df.loc[:, ['loss']].plot(title=\"Cross-entropy\")\nhistory_df.loc[:, ['binary_accuracy']].plot(title=\"Accuracy\")\ndel history_df\ndel history\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:44:01.525201Z","iopub.execute_input":"2023-07-26T16:44:01.526308Z","iopub.status.idle":"2023-07-26T16:44:03.220029Z","shell.execute_reply.started":"2023-07-26T16:44:01.526273Z","shell.execute_reply":"2023-07-26T16:44:03.218936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submissions","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\n\ntest_df = pd.DataFrame(np.load('/kaggle/input/t5embeds/t5embeds/test_embeds.npy'), columns = [\"Column_\" + str(i) for i in range(1, 1024+1)]).astype('float32')\ngc.collect()\ntime.sleep(2) # pausing for a breath -- slow down - \nprint(test_df.shape)\ndisplay(test_df.info())","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:44:03.221746Z","iopub.execute_input":"2023-07-26T16:44:03.222424Z","iopub.status.idle":"2023-07-26T16:44:16.967397Z","shell.execute_reply.started":"2023-07-26T16:44:03.222388Z","shell.execute_reply":"2023-07-26T16:44:16.966165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `test_df` is composed of 1024 columns and 141865 entries. We can see all 1024 dimensions(results will be truncated since column length is too long) of our dataset by printing out the first 5 entries using the following code:","metadata":{}},{"cell_type":"code","source":"test_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:44:16.9693Z","iopub.execute_input":"2023-07-26T16:44:16.969693Z","iopub.status.idle":"2023-07-26T16:44:16.976586Z","shell.execute_reply.started":"2023-07-26T16:44:16.969657Z","shell.execute_reply":"2023-07-26T16:44:16.975507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will now use the model to make predictions on the test embeddings. ","metadata":{}},{"cell_type":"code","source":"%%time\n%%memit\npredictions =  cp.array(model.predict(test_df))","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:44:16.978303Z","iopub.execute_input":"2023-07-26T16:44:16.97895Z","iopub.status.idle":"2023-07-26T16:44:40.004376Z","shell.execute_reply.started":"2023-07-26T16:44:16.978914Z","shell.execute_reply":"2023-07-26T16:44:40.003162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the predictions we will create the submission data frame.\n\n**Note**: This will take atleast **15 to 20** minutes to finish.\n\n\n**Note**: With GPU total notebook execution time is about 4 miutes -- even writing the CSV to disk.","metadata":{}},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:47:02.402423Z","iopub.execute_input":"2023-07-26T16:47:02.402797Z","iopub.status.idle":"2023-07-26T16:47:03.456178Z","shell.execute_reply.started":"2023-07-26T16:47:02.402764Z","shell.execute_reply":"2023-07-26T16:47:03.455071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%memit\n\ntest_protein_ids = np.load('/kaggle/input/t5embeds/t5embeds/test_ids.npy')\nl = []\nfor k in list(test_protein_ids):\n    l += [ k] * predictions.shape[1]   \n","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:46:36.520829Z","iopub.execute_input":"2023-07-26T16:46:36.521841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(len(l))\ndisplay(l[:10])","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:49:21.946574Z","iopub.execute_input":"2023-07-26T16:49:21.946961Z","iopub.status.idle":"2023-07-26T16:49:21.956048Z","shell.execute_reply.started":"2023-07-26T16:49:21.94693Z","shell.execute_reply":"2023-07-26T16:49:21.955016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleanup\nprint(len(l))\ndel test_df\ndel model\ndel test_protein_ids\ngc.collect()\ntime.sleep(2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%memit\ndf_submission = cudf.DataFrame({'Protein Id':l})\ndel l\ngc.collect()\ntime.sleep(2)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions.shape\n141865*1500","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:55:32.398996Z","iopub.execute_input":"2023-07-26T16:55:32.39943Z","iopub.status.idle":"2023-07-26T16:55:32.40853Z","shell.execute_reply.started":"2023-07-26T16:55:32.399401Z","shell.execute_reply":"2023-07-26T16:55:32.407443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%memit\ndf_submission['GO Term Id']= labels * predictions.shape[0]\ndel labels\ngc.collect()\ntime.sleep(2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%memit\ndf_submission['Prediction'] = predictions.ravel()\ndel predictions\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"1024*1500","metadata":{"execution":{"iopub.status.busy":"2023-07-26T16:54:39.626855Z","iopub.execute_input":"2023-07-26T16:54:39.628093Z","iopub.status.idle":"2023-07-26T16:54:39.63873Z","shell.execute_reply.started":"2023-07-26T16:54:39.628028Z","shell.execute_reply":"2023-07-26T16:54:39.636805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n%%memit\n# This can take 15 minutes. :-/ but only 26 seconds when a cuda dataframe.\ndf_submission.to_csv(\"submission.tsv\",header=False, index=False, sep=\"\\t\",chunksize=10000)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tmr.ShowTime() ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df_submission\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Thanks!","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}