{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Input data files are available in the read-only \"../input/\" directory\n\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-07-18T14:19:47.010186Z","iopub.execute_input":"2023-07-18T14:19:47.010527Z","iopub.status.idle":"2023-07-18T14:19:47.016144Z","shell.execute_reply.started":"2023-07-18T14:19:47.0105Z","shell.execute_reply":"2023-07-18T14:19:47.014557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Loading Dataset 'train_terms.tsv' which contains the list of annotated terms (functions) for the proteins. We will extract the labels aka GO term ID and create a label dataframe for the protein embeddings.","metadata":{}},{"cell_type":"markdown","source":"Gene Ontology(GO)is a controlled vocabulary that describes the functions of genes and gene products. It is constantly being updated as new information becomes available.","metadata":{}},{"cell_type":"code","source":"train_terms = pd.read_csv(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv\",sep=\"\\t\")\nprint(train_terms.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:20:06.94012Z","iopub.execute_input":"2023-07-18T14:20:06.940458Z","iopub.status.idle":"2023-07-18T14:20:09.528628Z","shell.execute_reply.started":"2023-07-18T14:20:06.940435Z","shell.execute_reply":"2023-07-18T14:20:09.527554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"train_terms dataframe is composed of 3 columns and 5363863 entries. We can see all 3 dimensions of our dataset below","metadata":{}},{"cell_type":"code","source":"train_terms.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:20:13.552528Z","iopub.execute_input":"2023-07-18T14:20:13.553423Z","iopub.status.idle":"2023-07-18T14:20:13.575904Z","shell.execute_reply.started":"2023-07-18T14:20:13.553338Z","shell.execute_reply":"2023-07-18T14:20:13.574769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we look at the first entry of train_terms.tsv, we can see that it contains protein id(A0A009IHW8), the GO term(GO:0008152) and its aspect(BPO)","metadata":{}},{"cell_type":"markdown","source":"**Overview of train_sequence.fasta**","metadata":{}},{"cell_type":"markdown","source":"In bioinformatics and biochemistry, the FASTA format is a text-based format for representing either nucleotide sequences or amino acid (protein) sequences, in which nucleotides or amino acids are represented using single-letter codes.","metadata":{}},{"cell_type":"code","source":"with open(\"/kaggle/input/cafa-5-protein-function-prediction/Train/train_sequences.fasta\", \"r\") as file:\n    fasta_100 = file.readlines()[:100]","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:20:18.469141Z","iopub.execute_input":"2023-07-18T14:20:18.469492Z","iopub.status.idle":"2023-07-18T14:20:19.774941Z","shell.execute_reply.started":"2023-07-18T14:20:18.469466Z","shell.execute_reply":"2023-07-18T14:20:19.773745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fasta_100[:10]","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:26:59.513472Z","iopub.execute_input":"2023-07-18T14:26:59.513845Z","iopub.status.idle":"2023-07-18T14:26:59.522106Z","shell.execute_reply.started":"2023-07-18T14:26:59.513816Z","shell.execute_reply":"2023-07-18T14:26:59.520498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Understand the fasta header: https://www.uniprot.org/help/fasta-headers","metadata":{}},{"cell_type":"markdown","source":"**Loading the protein embeddings**","metadata":{}},{"cell_type":"markdown","source":"We will now load the pre calculated protein embeddings created by Sergei Fironov using the Rost Lab's T5 protein language model.\n\n(https://www.kaggle.com/datasets/sergeifironov/t5embeds)\n\nThe protein embeddings to be used for training are recorded in train_embeds.npy and the corresponding protein ids are available in train_ids.npy.\n\nFirst, we will load the protein ids of the protein embeddings in the train dataset contained in train_ids.npy into a numpy array.","metadata":{}},{"cell_type":"markdown","source":"T5 (Text-To-Text Transfer Transformer) is a versatile transformer-based model capable of various natural language processing tasks. In the context of protein embeddings, T5 can be fine-tuned to learn protein representations by treating the amino acid sequences as text inputs.","metadata":{}},{"cell_type":"markdown","source":"Each protein's amino acid sequence is considered as a \"sentence\" or \"text,\" and the corresponding protein function(s) serve as the \"labels\" or \"categories.\" By fine-tuning the T5 model with labeled protein sequences, we can train it to predict the function(s) of unseen proteins based on their sequences.","metadata":{}},{"cell_type":"code","source":"train_protein_ids = np.load('/kaggle/input/t5embeds/train_ids.npy')\nprint(train_protein_ids.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:27:03.322164Z","iopub.execute_input":"2023-07-18T14:27:03.322489Z","iopub.status.idle":"2023-07-18T14:27:03.337178Z","shell.execute_reply.started":"2023-07-18T14:27:03.322463Z","shell.execute_reply":"2023-07-18T14:27:03.335775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_protein_ids[:5]","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:27:05.153116Z","iopub.execute_input":"2023-07-18T14:27:05.153469Z","iopub.status.idle":"2023-07-18T14:27:05.160992Z","shell.execute_reply.started":"2023-07-18T14:27:05.153437Z","shell.execute_reply":"2023-07-18T14:27:05.159412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After loading the files as numpy arrays, we will convert them into Pandas dataframe.\n\nEach protein embedding is a vector of length 1024. We create the resulting dataframe such that there are 1024 columns to represent the values in each of the 1024 places in the vector.","metadata":{}},{"cell_type":"code","source":"train_embeddings = np.load('/kaggle/input/t5embeds/train_embeds.npy')\n\n# Now lets convert embeddings numpy array(train_embeddings) into pandas dataframe.\ncolumn_num = train_embeddings.shape[1]\ntrain_df = pd.DataFrame(train_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\nprint(train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:27:08.085502Z","iopub.execute_input":"2023-07-18T14:27:08.085873Z","iopub.status.idle":"2023-07-18T14:27:08.471467Z","shell.execute_reply.started":"2023-07-18T14:27:08.085842Z","shell.execute_reply":"2023-07-18T14:27:08.47007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:27:12.577728Z","iopub.execute_input":"2023-07-18T14:27:12.578081Z","iopub.status.idle":"2023-07-18T14:27:12.601898Z","shell.execute_reply.started":"2023-07-18T14:27:12.578055Z","shell.execute_reply":"2023-07-18T14:27:12.600855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms.shape","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:27:18.071482Z","iopub.execute_input":"2023-07-18T14:27:18.071855Z","iopub.status.idle":"2023-07-18T14:27:18.078861Z","shell.execute_reply.started":"2023-07-18T14:27:18.071827Z","shell.execute_reply":"2023-07-18T14:27:18.077597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms['term'].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:27:19.991033Z","iopub.execute_input":"2023-07-18T14:27:19.991819Z","iopub.status.idle":"2023-07-18T14:27:20.373023Z","shell.execute_reply.started":"2023-07-18T14:27:19.991788Z","shell.execute_reply":"2023-07-18T14:27:20.372125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Plotting the most frequent terms","metadata":{}},{"cell_type":"code","source":"# Select first 50 values for plotting\nplot_df = train_terms['term'].value_counts().iloc[:50]\n\nfigure, axis = plt.subplots(1, 1, figsize=(12, 6))\n\nbp = sns.barplot(ax=axis, x=np.array(plot_df.index), y=plot_df.values)\nbp.set_xticklabels(bp.get_xticklabels(), rotation=90, size = 6)\naxis.set_title('Top 50 frequent GO term IDs')\nbp.set_xlabel(\"GO term IDs\", fontsize = 12)\nbp.set_ylabel(\"Count\", fontsize = 12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:04.102296Z","iopub.execute_input":"2023-07-18T14:28:04.102649Z","iopub.status.idle":"2023-07-18T14:28:05.048575Z","shell.execute_reply.started":"2023-07-18T14:28:04.102621Z","shell.execute_reply":"2023-07-18T14:28:05.047539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_terms.term.value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:19.445902Z","iopub.execute_input":"2023-07-18T14:28:19.44624Z","iopub.status.idle":"2023-07-18T14:28:19.917445Z","shell.execute_reply.started":"2023-07-18T14:28:19.446212Z","shell.execute_reply":"2023-07-18T14:28:19.915863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Using the most common GO term ID we will create a new dataframe by filtering the train terms","metadata":{}},{"cell_type":"markdown","source":"**Classifing proteins using random forest**","metadata":{}},{"cell_type":"code","source":"train_terms.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:24.147692Z","iopub.execute_input":"2023-07-18T14:28:24.148051Z","iopub.status.idle":"2023-07-18T14:28:24.15733Z","shell.execute_reply.started":"2023-07-18T14:28:24.148024Z","shell.execute_reply":"2023-07-18T14:28:24.156077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# list of 10 most common proteins in the dataset\nno_labels = 10\ntop_10_label_list = train_terms.value_counts('term', ascending=False).index[:no_labels].tolist()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:29.999195Z","iopub.execute_input":"2023-07-18T14:28:29.999569Z","iopub.status.idle":"2023-07-18T14:28:30.496067Z","shell.execute_reply.started":"2023-07-18T14:28:29.999539Z","shell.execute_reply":"2023-07-18T14:28:30.495024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_10_label_list","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:37.406597Z","iopub.execute_input":"2023-07-18T14:28:37.407766Z","iopub.status.idle":"2023-07-18T14:28:37.414446Z","shell.execute_reply.started":"2023-07-18T14:28:37.407692Z","shell.execute_reply":"2023-07-18T14:28:37.413342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if 'term' in 'train_terms' df is not in list then assign value '11' (for other)\ntrain_terms['label'] = train_terms['term'].apply(lambda x: top_10_label_list.index(x) if x in top_10_label_list else 11)\ntrain_terms.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:40.61546Z","iopub.execute_input":"2023-07-18T14:28:40.61585Z","iopub.status.idle":"2023-07-18T14:28:43.568722Z","shell.execute_reply.started":"2023-07-18T14:28:40.615813Z","shell.execute_reply":"2023-07-18T14:28:43.567968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create an empty series of required size for storing the labels\n# train_labels = np.zeros((train_protein_ids.shape[0], 1))\ntrain_protein_ids_df = pd.DataFrame(train_protein_ids, columns=['EntryID'])\ntrain_terms_label_agg = train_terms.groupby('EntryID')['label'].agg(lambda x: list(set(x))).reset_index()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:46.175137Z","iopub.execute_input":"2023-07-18T14:28:46.175608Z","iopub.status.idle":"2023-07-18T14:28:50.267928Z","shell.execute_reply.started":"2023-07-18T14:28:46.175585Z","shell.execute_reply":"2023-07-18T14:28:50.266921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"label_df = pd.merge(train_protein_ids_df, train_terms_label_agg, on='EntryID', how='inner')\n# label_df = list(label_df.loc[:,'label'])","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:54.559475Z","iopub.execute_input":"2023-07-18T14:28:54.559849Z","iopub.status.idle":"2023-07-18T14:28:54.708957Z","shell.execute_reply.started":"2023-07-18T14:28:54.559815Z","shell.execute_reply":"2023-07-18T14:28:54.707975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Proteins often have multiple functions or can be involved in various biological processes. In this case, we can use multi-label classification, where a protein may be associated with multiple function labels. The T5 model can be fine-tuned to handle multi-label tasks, predicting several functions simultaneously for a given protein sequence.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MultiLabelBinarizer\nmlb = MultiLabelBinarizer()\nmlb.fit(label_df['label'])","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:28:58.25254Z","iopub.execute_input":"2023-07-18T14:28:58.252959Z","iopub.status.idle":"2023-07-18T14:28:58.398332Z","shell.execute_reply.started":"2023-07-18T14:28:58.252924Z","shell.execute_reply":"2023-07-18T14:28:58.397314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# transform label df into multilabel binary dataframe\nbinary_matrix = mlb.transform(label_df['label'])\nbinary_df = pd.DataFrame(binary_matrix, columns=mlb.classes_)\nbinary_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:32:21.316075Z","iopub.execute_input":"2023-07-18T14:32:21.31646Z","iopub.status.idle":"2023-07-18T14:32:21.599579Z","shell.execute_reply.started":"2023-07-18T14:32:21.316427Z","shell.execute_reply":"2023-07-18T14:32:21.598433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.multioutput import MultiOutputClassifier","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:32:25.268487Z","iopub.execute_input":"2023-07-18T14:32:25.268834Z","iopub.status.idle":"2023-07-18T14:32:25.273185Z","shell.execute_reply.started":"2023-07-18T14:32:25.268807Z","shell.execute_reply":"2023-07-18T14:32:25.272159Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_classifier = RandomForestClassifier(n_estimators=3)\nmulti_target_classifier = MultiOutputClassifier(rf_classifier)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:32:30.965782Z","iopub.execute_input":"2023-07-18T14:32:30.966104Z","iopub.status.idle":"2023-07-18T14:32:30.972851Z","shell.execute_reply.started":"2023-07-18T14:32:30.966081Z","shell.execute_reply":"2023-07-18T14:32:30.97099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"multi_target_classifier.fit(train_df, binary_df)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:32:36.572359Z","iopub.execute_input":"2023-07-18T14:32:36.572687Z","iopub.status.idle":"2023-07-18T14:36:49.043535Z","shell.execute_reply.started":"2023-07-18T14:32:36.572663Z","shell.execute_reply":"2023-07-18T14:36:49.042217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\nwith open('multi_target_classifier_model.pkl', 'wb') as file:\n    pickle.dump(multi_target_classifier, file)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:38:05.909385Z","iopub.execute_input":"2023-07-18T14:38:05.909783Z","iopub.status.idle":"2023-07-18T14:38:05.970736Z","shell.execute_reply.started":"2023-07-18T14:38:05.909746Z","shell.execute_reply":"2023-07-18T14:38:05.969919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:38:12.343642Z","iopub.execute_input":"2023-07-18T14:38:12.344025Z","iopub.status.idle":"2023-07-18T14:38:12.35138Z","shell.execute_reply.started":"2023-07-18T14:38:12.343996Z","shell.execute_reply":"2023-07-18T14:38:12.350235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"binary_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:38:14.271339Z","iopub.execute_input":"2023-07-18T14:38:14.271717Z","iopub.status.idle":"2023-07-18T14:38:14.279483Z","shell.execute_reply.started":"2023-07-18T14:38:14.271669Z","shell.execute_reply":"2023-07-18T14:38:14.278441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Submission**","metadata":{}},{"cell_type":"markdown","source":"For submission we will use the protein embeddings of the test data created by Sergei Fironov using the Rost Lab's T5 protein language model.\n\nconvert to submission: 141865 x 1500\n\nwhy 1500?? last step output binary is 1500?","metadata":{}},{"cell_type":"code","source":"test_embeddings = np.load('/kaggle/input/t5embeds/test_embeds.npy')\n\n# Convert test_embeddings to dataframe\ncolumn_num = test_embeddings.shape[1]\ntest_df = pd.DataFrame(test_embeddings, columns = [\"Column_\" + str(i) for i in range(1, column_num+1)])\nprint(test_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:38:18.400189Z","iopub.execute_input":"2023-07-18T14:38:18.400588Z","iopub.status.idle":"2023-07-18T14:38:28.466898Z","shell.execute_reply.started":"2023-07-18T14:38:18.400559Z","shell.execute_reply":"2023-07-18T14:38:28.464982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test_df is composed of 1024 columns and 141865 entries. We can see all 1024 dimensions(results will be truncated since column length is too long) of our dataset by printing out the first 5 entries using the following code:","metadata":{}},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:41:07.145696Z","iopub.execute_input":"2023-07-18T14:41:07.146085Z","iopub.status.idle":"2023-07-18T14:41:07.169013Z","shell.execute_reply.started":"2023-07-18T14:41:07.146056Z","shell.execute_reply":"2023-07-18T14:41:07.168114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will now use the model to make predictions on the test embeddings.","metadata":{}},{"cell_type":"code","source":"predictions =  multi_target_classifier.predict_proba(test_df)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:42:54.5439Z","iopub.execute_input":"2023-07-18T14:42:54.544247Z","iopub.status.idle":"2023-07-18T14:42:58.340206Z","shell.execute_reply.started":"2023-07-18T14:42:54.54422Z","shell.execute_reply":"2023-07-18T14:42:58.339185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:43:01.566769Z","iopub.execute_input":"2023-07-18T14:43:01.567151Z","iopub.status.idle":"2023-07-18T14:43:01.578852Z","shell.execute_reply.started":"2023-07-18T14:43:01.567124Z","shell.execute_reply":"2023-07-18T14:43:01.577764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(predictions)","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:43:06.24549Z","iopub.execute_input":"2023-07-18T14:43:06.245887Z","iopub.status.idle":"2023-07-18T14:43:06.251985Z","shell.execute_reply.started":"2023-07-18T14:43:06.245856Z","shell.execute_reply":"2023-07-18T14:43:06.251071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission = pd.DataFrame(columns = ['Protein Id', 'GO Term Id','Prediction'])\ntest_protein_ids = np.load('/kaggle/input/t5embeds/test_ids.npy')","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:43:12.081356Z","iopub.execute_input":"2023-07-18T14:43:12.081693Z","iopub.status.idle":"2023-07-18T14:43:12.091695Z","shell.execute_reply.started":"2023-07-18T14:43:12.081668Z","shell.execute_reply":"2023-07-18T14:43:12.090642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:41:32.517833Z","iopub.execute_input":"2023-07-18T14:41:32.518186Z","iopub.status.idle":"2023-07-18T14:41:32.52798Z","shell.execute_reply.started":"2023-07-18T14:41:32.518159Z","shell.execute_reply":"2023-07-18T14:41:32.526247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission['Protein Id'] = test_protein_ids\ndf_submission['GO Term Id'] = 'GO:0005575'\ndf_submission['Prediction'] = predictions[:,1]","metadata":{"execution":{"iopub.status.busy":"2023-07-18T14:43:15.131555Z","iopub.execute_input":"2023-07-18T14:43:15.131918Z","iopub.status.idle":"2023-07-18T14:43:15.195343Z","shell.execute_reply.started":"2023-07-18T14:43:15.131891Z","shell.execute_reply":"2023-07-18T14:43:15.194272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission","metadata":{"execution":{"iopub.status.busy":"2023-07-11T12:03:26.182992Z","iopub.execute_input":"2023-07-11T12:03:26.183387Z","iopub.status.idle":"2023-07-11T12:03:26.193558Z","shell.execute_reply.started":"2023-07-11T12:03:26.183355Z","shell.execute_reply":"2023-07-11T12:03:26.192339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission.to_csv(\"submission.tsv\",header=False, index=False, sep=\"\\t\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}