{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ProteiNet v2 🧬 [Inference Notebook]","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://structuralbioinformatician.files.wordpress.com/2013/03/1ece.gif\">","metadata":{}},{"cell_type":"markdown","source":"# 1. Go beyond with Aspects Experts Predictions!","metadata":{}},{"cell_type":"markdown","source":"In the previous episode, we took a look at the competition context and developed a first model for this 5th edition of the CAFA competition named **ProteiNet** : https://www.kaggle.com/code/henriupton/proteinet-pytorch-ems2-t5-protbert-embeddings\n\nFollowing this, we implemented various additions to build its big brother: **ProteiNet v2**! This new version actually aims to train not one, not two, but 3 models, all three specialized in predicting a group of GOs of a particular aspect among the three sets presented for CAFA5: Molecular Function (MF), Biological Process (BP), and Cellular Component (CC).\n\nThis notebook is the second section of proteiNet v2 and it is dedicated to the inference from the models. If you want to have a look into the training section, follow this link : https://www.kaggle.com/code/henriupton/proteinet-v2-training-notebook\n\nGitHub version of proteiNet v2 is also available : https://github.com/henriupton99/proteinet-cafa5\n\nFeel free to give feedback for improvement, and drop an upvote to support our investment to the project. \n\nLet's take a look at what's new in this model!","metadata":{}},{"cell_type":"markdown","source":"# 2. New features of proteiNet v2","metadata":{}},{"cell_type":"markdown","source":"Thanks to the great interest shown in the notebook dedicated to ProteiNet, a large number of bugs and defects have been corrected in this new version. On the other hand, my team (M. Sato, F. Lin and myself) have tried to innovate as much as possible and incorporate various topics of discussion from the competition for ProteiNet v2. Here is an exhaustive list of the most important innovations:\n\n- **Rather than training a single model to predict the scores of all GOs for each protein, train three separate models capable of predicting the scores of each protein for GOs of a specific aspect (BPO, MFO, CCO)**. This is why we call these three models \"experts\". This practice has several theoretical virtues, such as the fact that each model aims to perform multilabel classification on a smaller number of classes. In addition, they are trained to work on GO embeddings that are highly likely to be parent/child, as they come from the same aspect. Once the models have been trained, the predictions of each model will be concatenated to form the final submission.\n\n- **The GO classes to be predicted are no longer naively the top K of the most frequent GO classes in the database.** It has been discussed time and again that various other methods can be used to select GO classes more strategically. In addition, as we train experts based on aspect groups, we define as classes for each model the top K most frequent GOs filtered on aspects. The number of classes per aspect is defined on the basis of these observations: https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/431491 \n\n- **We restrict the GO classes to be predicted to those whose evidence code is Inferred from Experiment (EXP) (and its subgroups)**.(https://wiki.geneontology.org/index.php/Inferred_from_Experiment_(EXP)) This choice stems from the desire to be as close as possible to the explanations given in the Background section of the Competition Evaluation page: https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/overview/evaluation\n\n- **We incorporate the weights given by IA.txt into our Cross Entropy loss function when training the model.** Again in an effort to keep up with the competition, the weights enable us to place greater emphasis on infrequent GOs (at the root of the graph), which are consequently the most important ones. (https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/405237)(https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html)\n\n- **We have implemented a cross-validation (CV) process to prevent from overfitting.** There has long been a consensus that incorporating this method into the pipeline of one's work reduces the risk of overfitting on the public LB and by consequence get a good result on the private LB. The actual CV is composed of 5 folds (so 5 different classic 80-20 train-test splits). (https://github.com/christianversloot/machine-learning-articles/blob/main/how-to-use-k-fold-cross-validation-with-pytorch.md) Once the models perform well for the 5 folds, we train the final models on the full train set thanks to the config hyperparameter VALIDATION_MODE = False.\n\n- **Once the predictions were formed for the three expert models, we defined a minimal score threshold in order to filter the predictions.** For each prediction row, if it does not exceed the threshold, it is deleted from the final submission. (https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/431652) This hyperparameter can be tuned thanks to the constant PROB_THRESHOLD in the config class. It allows to benefit at maximum from the propagation process of GOs predictions. (https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/overview/evaluation)","metadata":{}},{"cell_type":"markdown","source":"# 3. Configuration of the working environment","metadata":{}},{"cell_type":"code","source":"import torch\nclass CONFIG:\n    \n    # CONSTANTS FOR DATA PATHS\n    MAIN_DIR = \"/kaggle/input/cafa-5-protein-function-prediction/\"\n    GO_OBO_FILE = MAIN_DIR + \"Train/go-basic.obo\"\n    TRAIN_SEQUENCES_FASTA = MAIN_DIR  + \"Train/train_sequences.fasta\"\n    TRAIN_LABELS = MAIN_DIR + \"Train/train_terms.tsv\"\n    TRAIN_IDS = \"/kaggle/input/protbert-embeddings-for-cafa5/train_ids.npy\"\n    IA_WEIGHTS = MAIN_DIR + \"IA.txt\"\n    TEST_SEQUENCES_FASTA = MAIN_DIR + \"/Test (Targets)/testsuperset.fasta\"\n    TARGETS_PATH = \"/kaggle/working/train-labels-targets/\"\n    EVIDENCE_CODES = \"/kaggle/input/enhanced-train-terms/propagated_evidenceCode.parquet\"\n    \n    # CONSTANTS FOR ASPECTS :\n    ASPECTS = [\"BPO\", \"CCO\", \"MFO\"]\n    ASPECTS_LABELS = {\"BPO\" : 1100, \"CCO\" : 300, \"MFO\" : 450}\n\n    # CONSTANTS FOR TRAINING : \n    EMBEDDINGS_SOURCE = \"ESM2\"\n    K_FOLDS = 5\n    N_EPOCHS = 15\n    BATCHS_SIZE = 256\n    LEARNING_RATE = 0.001\n    DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n    ASPECTS_HIDDEN_SIZE = {\"BPO\" : 1256, \"CCO\" : 512, \"MFO\" : 850}\n    VALIDATION_MODE = False\n    \n    # CONSTANTS FOR POSTPROCESSING :\n    PROB_THRESHOLD = 0.10","metadata":{"execution":{"iopub.status.busy":"2023-08-21T19:52:53.88748Z","iopub.execute_input":"2023-08-21T19:52:53.887841Z","iopub.status.idle":"2023-08-21T19:52:56.805966Z","shell.execute_reply.started":"2023-08-21T19:52:53.88781Z","shell.execute_reply":"2023-08-21T19:52:56.804914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Preprocess the data : build the targets","metadata":{}},{"cell_type":"code","source":"\"\"\"FUNCTIONS FOR GENERATING THE LABELS TAGRETS FOR ALL ASPECTS\n\"\"\"\nimport pandas as pd\nimport numpy as np\nfrom tqdm import tqdm\nimport re\nimport gc\n\ndef extract_go_terms_and_branches(\n    file_path : str\n    ) -> dict:\n    \"\"\"Utilitary function to construct a mapping {GO TERM : ASPECT} for each GO TERM in input OBO file\n\n    Args:\n        file_path (str): file path for input GO terms\n\n    Returns:\n        go_terms_dict (dict): mapping dictionnary\n    \"\"\"\n    with open(file_path, 'r') as file:\n        content = file.read()\n        stanzas = re.findall(r'\\[Term\\][\\s\\S]*?(?=\\n\\[|$)', content)\n    go_terms_dict = {}\n    for stanza in stanzas:\n        go_id = re.search(r'^id: (GO:\\d+)', stanza, re.MULTILINE)\n        if go_id:\n            go_id = go_id.group(1)\n        namespace = re.search(r'^namespace: (\\w+)', stanza, re.MULTILINE)\n        if namespace:\n            namespace = namespace.group(1)\n        if go_id and namespace:\n            branch_abbr = {'biological_process': 'BPO', 'cellular_component': 'CCO', 'molecular_function': 'MFO'}\n            go_terms_dict[go_id] = branch_abbr[namespace]\n\n    return go_terms_dict\n\ndef generate_labels_matrix(\n    ids : np.ndarray,\n    labels_names : list[str],\n    id_labels : dict,\n    go_terms_map : dict\n    ):\n    \"\"\"Utilitary function to generate labels target matrix given :\n    - protein ids, labels_names (GO terms names)\n    - id labels : id of labels\n    - go terms map generated by function extract_go_terms_and_branches\n\n    \"\"\"\n    labels_matrix = np.zeros((len(ids), len(labels_names)))\n    \n    for index, id in tqdm(enumerate(ids)):\n        try :\n            id_gos_list = id_labels[id]\n            temp = [go_terms_map[go] for go in labels_names if go in id_gos_list]\n            labels_matrix[index, temp] = 1\n        except:\n            pass\n        \n    return labels_matrix\n\ndef generate_targets(\n    ids_path : str,\n    labels_path : str,\n    weights_path : str,\n    go_obo_path : str,\n    evidence_codes_path : str,\n    targets_path : str,\n    aspects_list : str,\n    go_terms_per_aspects : dict[int]\n    ):\n    \"\"\"Function to generate labels (targets) for a given aspect (BPO, CCO, or MFO) for each protein id in ids, based on labels dataframe.\n    NB : For memory usage and models precision reasons, we only consider a subset of all GO terms labels\n    We consider to top K most frequent for each aspect (based on go_terms_per_aspects input dictionnary)\n\n    Args:\n        ids_path (str): path to protein ids \n        labels_paths (str): path to labels annotations dataframe for each protein in ids\n        weights_path (str): path to IA weights for GO terms metric computation\n        go_obo_path (str) : path to obo graph file\n        evidence_codes_path (str) : path to evidence codes to filter for EXP GO terms\n        targets_path (str) : path where to save the target labels\n        aspects_list (str): list of aspects to consider\n        go_terms_per_aspects (dict[int]): number of GO term classes to consider per aspect\n    \"\"\"\n    ids = np.load(ids_path)\n    labels = pd.read_csv(labels_path, sep = \"\\t\")\n    colnames = [\"term\", \"weight\"]\n    ia_weights = pd.read_csv(weights_path, sep = \"\\t\", names = colnames, header=None)\n    evidence_codes = pd.read_parquet(evidence_codes_path)\n    evidence_codes = evidence_codes[evidence_codes[\"EvidenceCode\"].notnull()][\"term\"].unique().tolist()\n    \n    for aspect in aspects_list:\n        print(\"=\"*25)\n        print(\"START LOADING FOR ASPECT {}\".format(aspect))\n        aspects_labels = labels[labels[\"aspect\"] == aspect]\n        aspects_labels = aspects_labels[aspects_labels[\"term\"].isin(evidence_codes)]\n        top_terms = aspects_labels.groupby(\"term\")[\"EntryID\"].count().sort_values(ascending=False).to_frame()\n        map_go_terms_aspects = extract_go_terms_and_branches(\n                                file_path=go_obo_path\n                                )\n        top_terms[\"aspect\"] = top_terms.index.map(map_go_terms_aspects)\n        labels_names = top_terms[:go_terms_per_aspects[aspect]]\n        labels_names = labels_names.index.values\n        weights_df = pd.DataFrame(data={\"term\" : labels_names})\n        weights_df = weights_df.merge(ia_weights, on = \"term\", how = \"left\")\n        print(\"NUMBER OF GO TERMS IN {} SUBSET : {}\".format(aspect,str(len(labels_names))))\n        train_labels_sub = labels[(labels.term.isin(labels_names)) & (labels.EntryID.isin(ids))]\n        id_labels = train_labels_sub.groupby('EntryID')['term'].apply(list).to_dict()\n        go_terms_map = {label: i for i, label in enumerate(labels_names)}\n        labels_matrix = generate_labels_matrix(\n            ids=ids,\n            labels_names=labels_names,\n            id_labels=id_labels,\n            go_terms_map=go_terms_map\n        )\n        labels_list = []\n        for l in range(labels_matrix.shape[0]):\n            labels_list.append(labels_matrix[l, :])\n\n        labels_df = pd.DataFrame(data={\"EntryID\":ids, \"labels_vect\":labels_list})\n        labels_df.to_pickle(\"/kaggle/working/train_targets_{}.pkl\".format(aspect))\n        weights_df.to_csv(\"/kaggle/working/weights_{}.csv\".format(aspect),index=None)\n        del aspects_labels, labels_df, weights_df, labels_list, labels_matrix, go_terms_map\n        gc.collect()\n        print(\"GENERATION FINISHED FOR ASPECT {}\".format(aspect))\n    \n    del ids, labels, ia_weights, evidence_codes\n    gc.collect()\n    print(\"GENERATION FINISHED ! :D\")\n    return ","metadata":{"execution":{"iopub.status.busy":"2023-08-20T22:27:36.991856Z","iopub.execute_input":"2023-08-20T22:27:36.992409Z","iopub.status.idle":"2023-08-20T22:27:37.030479Z","shell.execute_reply.started":"2023-08-20T22:27:36.992363Z","shell.execute_reply":"2023-08-20T22:27:37.029475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"generate_targets(\n    ids_path = CONFIG.TRAIN_IDS,\n    labels_path = CONFIG.TRAIN_LABELS,\n    weights_path = CONFIG.IA_WEIGHTS,\n    go_obo_path = CONFIG.GO_OBO_FILE,\n    evidence_codes_path = CONFIG.EVIDENCE_CODES,\n    targets_path = CONFIG.TARGETS_PATH,\n    aspects_list = CONFIG.ASPECTS,\n    go_terms_per_aspects = CONFIG.ASPECTS_LABELS\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T22:27:37.899133Z","iopub.execute_input":"2023-08-20T22:27:37.8995Z","iopub.status.idle":"2023-08-20T22:30:23.228524Z","shell.execute_reply.started":"2023-08-20T22:27:37.89947Z","shell.execute_reply":"2023-08-20T22:30:23.227539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Build the Pytorch *Dataset* instance","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom torch.utils.data import Dataset\n\nclass ProteinSequenceDataset(Dataset):\n    \n    embeds_map = {\n    \"T5\" : \"/kaggle/input/t5embeds\",\n    \"ProtBERT\" : \"/kaggle/input/protbert-embeddings-for-cafa5/\",\n    \"ESM2\" : \"/kaggle/input/4637427/\"\n    }\n    \n    def __init__(self, aspect, datatype, embeddings_source):\n        super(ProteinSequenceDataset).__init__()\n        self.datatype = datatype\n        \n        if embeddings_source == \"ProtBERT\":\n            embeds = np.load(ProteinSequenceDataset.embeds_map[embeddings_source]+datatype+\"_embeddings.npy\")\n            ids = np.load(ProteinSequenceDataset.embeds_map[embeddings_source]+datatype+\"_ids.npy\")\n        if embeddings_source == \"T5\":\n            embeds = np.load(ProteinSequenceDataset.embeds_map[embeddings_source]+datatype+\"_embeds.npy\")\n            ids = np.load(ProteinSequenceDataset.embeds_map[embeddings_source]+datatype+\"_ids.npy\")\n        if embeddings_source == \"ESM2\":\n            embeds = np.load(ProteinSequenceDataset.embeds_map[embeddings_source]+datatype+\"_embeds_esm2_t36_3B_UR50D.npy\")\n            ids = np.load(ProteinSequenceDataset.embeds_map[embeddings_source]+datatype+\"_ids_esm2_t36_3B_UR50D.npy\")\n        \n            \n        embeds_list = []\n        for l in range(embeds.shape[0]):\n            embeds_list.append(embeds[l,:])\n        self.df = pd.DataFrame(data={\"EntryID\": ids, \"embed\" : embeds_list})\n        del embeds_list\n        \n        if datatype==\"train\":\n            \n            df_labels = pd.read_pickle(\n                \"/kaggle/working/train_targets_{}.pkl\".format(aspect)\n            )\n            \n            self.df = self.df.merge(df_labels, how=\"right\", on=\"EntryID\")\n        \n    def __len__(self):\n        return len(self.df)\n    \n    def __getitem__(self, index):\n        embed = torch.tensor(self.df.iloc[index][\"embed\"] , dtype = torch.float32)\n        if self.datatype==\"train\":\n            targets = torch.tensor(self.df.iloc[index][\"labels_vect\"], dtype = torch.float32)\n            return embed, targets\n        if self.datatype==\"test\":\n            id = self.df.iloc[index][\"EntryID\"]\n            return embed, id","metadata":{"execution":{"iopub.status.busy":"2023-08-20T22:32:15.342176Z","iopub.execute_input":"2023-08-20T22:32:15.343218Z","iopub.status.idle":"2023-08-20T22:32:15.357721Z","shell.execute_reply.started":"2023-08-20T22:32:15.343156Z","shell.execute_reply":"2023-08-20T22:32:15.356448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Define the Pytorch *Model* instance","metadata":{}},{"cell_type":"code","source":"from torch import nn\n\nembeds_dim = {\n    \"T5\" : 1024,\n    \"ProtBERT\" : 1024,\n    \"EMS2\" : 2560\n}\n\nclass LinearModel(torch.nn.Module):\n\n    def __init__(self, input_dim, hidden_dim, num_classes):\n        super(LinearModel, self).__init__()\n        self.linear1 = torch.nn.Linear(input_dim, hidden_dim)\n        self.activation1 = torch.nn.ReLU()\n        self.linear2 = torch.nn.Linear(hidden_dim, num_classes)\n\n    def forward(self, x):\n        x = self.linear1(x)\n        x = self.activation1(x)\n        x = self.linear2(x)\n        return x\n    \n\nclass CNN1D(nn.Module):\n    def __init__(self, input_dim, num_classes):\n        super(CNN1D, self).__init__()\n        self.conv1 = nn.Conv1d(in_channels=1, out_channels=3, kernel_size=3, dilation=1, padding=1, stride=1)\n        self.pool1 = nn.MaxPool1d(kernel_size=2, stride=2)\n        self.conv2 = nn.Conv1d(in_channels=3, out_channels=8, kernel_size=3, dilation=1, padding=1, stride=1)\n        self.pool2 = nn.MaxPool1d(kernel_size=2, stride=2)\n        self.fc1 = nn.Linear(in_features=int(8 * input_dim/4), out_features=128)\n        self.fc2 = nn.Linear(in_features=128, out_features=num_classes)\n\n    def forward(self, x):\n        x = x.reshape(x.shape[0], 1, x.shape[1])\n        x = self.pool1(nn.functional.relu(self.conv1(x)))\n        x = self.pool2(nn.functional.relu(self.conv2(x)))\n        x = torch.flatten(x, 1)\n        x = nn.functional.relu(self.fc1(x))\n        x = self.fc2(x)\n        return x","metadata":{"execution":{"iopub.status.busy":"2023-08-20T22:32:16.688563Z","iopub.execute_input":"2023-08-20T22:32:16.689462Z","iopub.status.idle":"2023-08-20T22:32:16.703487Z","shell.execute_reply.started":"2023-08-20T22:32:16.689422Z","shell.execute_reply":"2023-08-20T22:32:16.702482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Make predictions from expert models","metadata":{}},{"cell_type":"code","source":"embeds_dim = {\n    \"T5\" : 1024,\n    \"ProtBERT\" : 1024,\n    \"ESM2\" : 2560\n}\n\n\ndef make_predictions(\n    aspect,\n    prob_threshold,\n    embeddings_source,\n    device\n    ):\n    \n    test_dataset = ProteinSequenceDataset(aspect=aspect, datatype=\"test\", embeddings_source = embeddings_source)\n    test_dataloader = torch.utils.data.DataLoader(test_dataset, batch_size=1, shuffle=False)\n    df_weights = pd.read_csv(\"/kaggle/working/weights_{}.csv\".format(aspect))\n    labels_names = list(df_weights[\"term\"].values)\n    num_labels = len(labels_names)\n    \n    model = LinearModel(\n        input_dim=embeds_dim[embeddings_source],\n        hidden_dim=CONFIG.ASPECTS_HIDDEN_SIZE[aspect],\n        num_classes=num_labels).to(device)\n\n    model_path = \"/kaggle/input/expert-models-cafa5/{}/expert_model.pt\".format(aspect)\n    model.load_state_dict(torch.load(model_path))\n    model.eval()\n    \n    print(\"=\"*25)\n    print(\"GENERATE PREDICTION FOR ASPECT {}\".format(aspect))\n\n    ids_ = np.empty(shape=(len(test_dataloader)*num_labels,), dtype=object)\n    go_terms_ = np.empty(shape=(len(test_dataloader)*num_labels,), dtype=object)\n    confs_ = np.empty(shape=(len(test_dataloader)*num_labels,), dtype=np.float32)\n\n    for i, (embed, id) in tqdm(enumerate(test_dataloader)):\n        embed = embed.to(device)\n        if device == \"cpu\":\n            confs_[i*num_labels:(i+1)*num_labels] = torch.sigmoid(model(embed)).squeeze().detach().numpy()\n        else:\n            confs_[i*num_labels:(i+1)*num_labels] = torch.sigmoid(model(embed)).squeeze().detach().cpu().numpy()\n        ids_[i*num_labels:(i+1)*num_labels] = id[0]\n        go_terms_[i*num_labels:(i+1)*num_labels] = labels_names\n    \n    len_before_delete = len(ids_)\n    rows_to_delete = confs_ < prob_threshold\n    confs_ = confs_[~rows_to_delete]\n    ids_ = ids_[~rows_to_delete]\n    go_terms_ = go_terms_[~rows_to_delete]\n    len_after_delete = len(ids_)\n    print(\"NUMBER OF ROWS DELETED (THRESHOLD) : {}\".format(len_before_delete - len_after_delete))\n    \n    del model, rows_to_delete\n    gc.collect()\n    submission_df = pd.DataFrame(data={\"Id\" : ids_, \"GO term\" : go_terms_, \"Confidence\" : confs_})\n    \n    del confs_, ids_, go_terms_\n    gc.collect()\n    submission_df.to_csv(\"/kaggle/working/predictions_{}.csv\".format(aspect))\n    print(\"PREDICTIONS DONE FOR ASPECT {}\".format(aspect))\n    return submission_df","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-20T22:32:18.40036Z","iopub.execute_input":"2023-08-20T22:32:18.400719Z","iopub.status.idle":"2023-08-20T22:32:18.416812Z","shell.execute_reply.started":"2023-08-20T22:32:18.40069Z","shell.execute_reply":"2023-08-20T22:32:18.414263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\nsub = pd.DataFrame()\nfor aspect in CONFIG.ASPECTS:\n    temp = make_predictions(\n            aspect = aspect,\n            prob_threshold = CONFIG.PROB_THRESHOLD,\n            embeddings_source = CONFIG.EMBEDDINGS_SOURCE,\n            device = CONFIG.DEVICE\n            )\n    sub = pd.concat([sub, temp])\n    del temp\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T22:32:18.993262Z","iopub.execute_input":"2023-08-20T22:32:18.994447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub","metadata":{"execution":{"iopub.status.busy":"2023-08-08T13:24:02.177821Z","iopub.execute_input":"2023-08-08T13:24:02.179065Z","iopub.status.idle":"2023-08-08T13:24:02.22009Z","shell.execute_reply.started":"2023-08-08T13:24:02.179009Z","shell.execute_reply":"2023-08-08T13:24:02.218674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.tsv', sep='\\t', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-08-08T13:25:47.535736Z","iopub.execute_input":"2023-08-08T13:25:47.536249Z","iopub.status.idle":"2023-08-08T13:36:36.193427Z","shell.execute_reply.started":"2023-08-08T13:25:47.536214Z","shell.execute_reply":"2023-08-08T13:36:36.192001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thanks for reading :D dont forget to upvote and give feedback !","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://media.giphy.com/media/10LKovKon8DENq/giphy.gif\">","metadata":{}}]}