{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Work in progress...","metadata":{}},{"cell_type":"markdown","source":"## Introduction\n\nHello [@vukpetar](https://www.kaggle.com/vukpetar) and I  worked on a little bit different approach. We tried to implement something new. The motivation for this is simple. From resources we have some basics GPU on laptops, Google Colab and Kaggle. With these resources we cannot compete against people who has much more resources. As this competition is very sensitive to image size we wanted to try something new, to learn new stuff, but also to check others opinions and how to improve this approach. Two weeks ago we found this [paper](https://arxiv.org/pdf/2301.13817.pdf) (**Patch Gradient Descent: Training Neural Networks on Very Large Images**). As it is relatively new, we had to make implementation for it and check its results on this dataset. One of the reasons for that is because it reported good results on Prostate cANcer graDe Assessment (PANDA) dataset. \nThey introduced a PatchGD as a novel method that allows to train the existing CNN architectures on large-scale images in an end-to-end manner. As its name says, here the gradient updates are done on small parts of the image at a time, ensuring that the majority of it is covered over the course of iterations and not on the whole image at once. \nAs we saw in previous disccussions, [notebooks](https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example), NextViT achieved some great results so our idea was to implement those two as some type of „proof of concept“ as we started last week working on this, so there is no time for some exhibitions.\nOn the following [link](https://github.com/vukpetar/rsna-screening-mammography-cancer-detection), you can find our github repo of work.\n","metadata":{}},{"cell_type":"markdown","source":"## Dataset\n\nIn this section, there is not much to say, as we used dataset that [@hengck23](https://www.kaggle.com/hengck23) proposed in his [notebook](https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example). As the scans are taken with different machines, every machine has its own encoding rules for scans. In this dataset we had two different types of encoding: JPEG Lossless, Nonhierarchical, First- Order Prediction (Processes 14) and JPEG 2000 Image Compression (Lossless Only). So just like he explained, with knowing this, we converted them to the png format, centered the breasts. On these images, we applied **resnet34d** model for windowing the breasts ([notebook](https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region)). And whoala, that is the dataset.","metadata":{}},{"cell_type":"markdown","source":"## Algorithm\n\nOn the image below it is depicted a scheme of PatchGD algorithm (Note: Image is taken from original paper).\n\n![Algorithm](https://i.imgur.com/vm7UCKC.png)\n\nThe main part of the algorithm is the filling of Z block, which actually represents a deep latent representation of full input image. Z is primarily an encoding of an input image X obtained using any given model parameterized with weights. The size of Z is constant (m x n x s). The process of Z-filling spans over multiple steps, where every step involves sampling k patches and their respective positions from X and passing them as a batch to the model for processing. When all patches are sampled, we can say that the filled form of Z is obtained. Also, as it is said that an end-to-end CNN models are built, it is added a small network. Task of that second network is to transform information contained in Z into a vector of probabilities for our task of classification.\nSome pseudocode is depicted on image below:\n\n![Pseudocode](https://i.imgur.com/bfStmIp.png)","metadata":{}},{"cell_type":"markdown","source":"## Our contribution","metadata":{}},{"cell_type":"markdown","source":"In our code, models θ1 and θ2 are called as Q models. In our case Q1 represents the **NextViT** model and Q2 is **TransformerEncoderLayer**. Main difference between our approach and other approaches is that we trained patches on patient-laterality level, which means that all patient-laterality images have one label.","metadata":{}},{"cell_type":"markdown","source":"## Code","metadata":{}},{"cell_type":"code","source":"!pip -q install timm\n!pip -q install einops","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-28T02:09:03.432687Z","iopub.execute_input":"2023-02-28T02:09:03.433098Z","iopub.status.idle":"2023-02-28T02:09:24.815304Z","shell.execute_reply.started":"2023-02-28T02:09:03.433014Z","shell.execute_reply":"2023-02-28T02:09:24.813965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir /kaggle/tmp/\n!mkdir /kaggle/tmp/~png\n!unzip -qq \"/kaggle/input/zip-dataset/tensorrt-nextvit-dataset.zip\" -d \"/kaggle/tmp/~png/\"\n!unzip -qq \"/kaggle/input/zip-dataset-non/tensorrt-nextvit-dataset-non.zip\" -d \"/kaggle/tmp/~png/\"\n!mv \"/kaggle/tmp/~png/kaggle/input/tensorrt-nextvit-dataset-non/~png/\"* \"/kaggle/tmp/~png/kaggle/input/tensorrt-nextvit-dataset/~png/\"","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:15:55.998312Z","iopub.execute_input":"2023-02-27T19:15:55.998642Z","iopub.status.idle":"2023-02-27T19:23:06.747903Z","shell.execute_reply.started":"2023-02-27T19:15:55.99861Z","shell.execute_reply":"2023-02-27T19:23:06.745983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%cd /kaggle/input/patch-gd-nextvit-guthub\nimport sys\nsys.path.append('./')\nsys.path.append('/kaggle/input/patch-gd-nextvit-guthub')","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:23:06.750508Z","iopub.execute_input":"2023-02-27T19:23:06.750963Z","iopub.status.idle":"2023-02-27T19:23:06.775237Z","shell.execute_reply.started":"2023-02-27T19:23:06.750913Z","shell.execute_reply":"2023-02-27T19:23:06.771122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nos.environ['CUDA_MODULE_LOADING']='LAZY'\n\nfrom types import SimpleNamespace\nimport pandas as pd\nfrom sklearn.model_selection import KFold\n\nfrom timm.optim import create_optimizer\nfrom timm.scheduler import create_scheduler\n\nimport numpy as np\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import DataLoader, SequentialSampler\n\nfrom rsna_utils import RsnaDataset\nfrom rsna_nets import Q1Net, Q2Net, load_q1_pretrained\nfrom rsna_engine import train_one_epoch, evaluate\n\nprint(torch.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:23:06.777026Z","iopub.execute_input":"2023-02-27T19:23:06.7777Z","iopub.status.idle":"2023-02-27T19:23:12.812975Z","shell.execute_reply.started":"2023-02-27T19:23:06.777655Z","shell.execute_reply":"2023-02-27T19:23:12.810369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"??RsnaDataset","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:41:18.894664Z","iopub.execute_input":"2023-02-27T19:41:18.895035Z","iopub.status.idle":"2023-02-27T19:41:18.965829Z","shell.execute_reply.started":"2023-02-27T19:41:18.895005Z","shell.execute_reply":"2023-02-27T19:41:18.964833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DF_PATH = \"/kaggle/input/rsna-mammography-train-csv-with-box/train_with_box.csv\"\nIMG_PATH = \"/kaggle/tmp/~png/kaggle/input/tensorrt-nextvit-dataset/~png/\"\ntrain_df = pd.read_csv(DF_PATH)\ntrain_df[\"id\"] = train_df.apply(lambda x: str(x.patient_id) + \"_\" + str(x.laterality), axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:23:12.816936Z","iopub.execute_input":"2023-02-27T19:23:12.817257Z","iopub.status.idle":"2023-02-27T19:23:14.173313Z","shell.execute_reply.started":"2023-02-27T19:23:12.817224Z","shell.execute_reply":"2023-02-27T19:23:14.17235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"device = \"cuda\"\nlr = 5e-5\nweight_decay = 0.05\nmomentum = 0.9\nbatch_size = 12\nEPOCHES = 1 #just to save notebook, but number need to be higher\nbatch_size = 8\n\nq1_model = Q1Net(stem_chs=[64, 32, 64], depths=[3, 4, 20, 3], path_dropout=0.2)\nq1_model.to(device)\n# load_q1_pretrained(\"/kaggle/input/nextvit-base-in1k6m-384/nextvit_base_in1k6m_384.pth\", q1_model)\nq1_model_pretrained = torch.load(\"/kaggle/input/nextvit-patchgd-f1-e2/fold-1-q1_model_2.pth\")\nq1_model.load_state_dict(q1_model_pretrained.state_dict(), strict=False)\n\n\nq2_model = Q2Net(dim_feedforward =2048)\nq2_model.to(device)\nq2_model_pretrained = torch.load(\"/kaggle/input/nextvit-patchgd-f1-e2/fold-1-q2_model_2.pth\")\nq2_model.load_state_dict(q2_model_pretrained.state_dict(), strict=False)\n\nargs = SimpleNamespace()\nargs.opt = \"adamw\"\nargs.lr = lr\nargs.weight_decay = weight_decay\nargs.batch_size = batch_size\nargs.decay_rate = 0.1\nargs.decay_epochs = 5\nargs.warmup_epochs = 2\nargs.cooldown_epochs = 3\nargs.epochs = EPOCHES\nargs.min_lr = 1e-5\nargs.warmup_lr = 1e-6\nargs.sched = \"cosine\"\nargs.momentum = momentum\n\nq1_optimizer = create_optimizer(args, q1_model)\nq1_lr_scheduler, _ = create_scheduler(args, q1_optimizer)\n\nq2_optimizer = create_optimizer(args, q2_model)\nq2_lr_scheduler, _ = create_scheduler(args, q2_optimizer)\n\ncriterion = nn.CrossEntropyLoss()\n\ninner_iterations = 4\npatch_size = 384\npatches_per_in_iter = 12\ngrad_acc_steps = 2","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:49:11.621383Z","iopub.execute_input":"2023-02-27T19:49:11.621752Z","iopub.status.idle":"2023-02-27T19:49:13.004385Z","shell.execute_reply.started":"2023-02-27T19:49:11.621721Z","shell.execute_reply":"2023-02-27T19:49:13.003395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FOLD_NUM = 1\nall_patient_ids = train_df.id.unique()\npatient_ids_w_cancer = train_df[train_df.cancer == 1].id.unique()\npatient_ids_wo_cancer = np.array([p for p in all_patient_ids if p not in patient_ids_w_cancer])\nkf = KFold(n_splits=5)\nkf.get_n_splits(patient_ids_w_cancer)\npatients_w_iter = kf.split(patient_ids_w_cancer)\nkf.get_n_splits(patient_ids_wo_cancer)\npatients_wo_iter = kf.split(patient_ids_wo_cancer)\nfor _ in range(FOLD_NUM):\n    first_fold_positive_train, first_fold_positive_valid = next(patients_w_iter)\n    first_fold_negative_train, first_fold_negative_valid = next(patients_wo_iter)\n    \ntrain_patient_ids = np.concatenate(\n    (\n        patient_ids_w_cancer[first_fold_positive_train],\n        patient_ids_wo_cancer[first_fold_negative_train]\n    )\n)\ntrain_labels = np.concatenate((\n    np.ones((len(first_fold_positive_train), ), dtype=np.int64),\n    np.zeros((len(first_fold_negative_train),), dtype=np.int64)\n))\nvalid_patient_ids = np.concatenate(\n    (\n        patient_ids_w_cancer[first_fold_positive_valid],\n        patient_ids_wo_cancer[first_fold_negative_valid]\n    )\n)\nvalid_labels = np.concatenate((\n    np.ones((len(first_fold_positive_valid), ), dtype=np.int64),\n    np.zeros((len(first_fold_negative_valid),), dtype=np.int64)\n))","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:49:35.265819Z","iopub.execute_input":"2023-02-27T19:49:35.266179Z","iopub.status.idle":"2023-02-27T19:49:35.616606Z","shell.execute_reply.started":"2023-02-27T19:49:35.266148Z","shell.execute_reply":"2023-02-27T19:49:35.615337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset = RsnaDataset(train_patient_ids, train_labels, \n                            is_train=True, positive_ratio=\"1in1\")\ntrain_loader = DataLoader(\n    train_dataset,\n    sampler=SequentialSampler(train_dataset),\n    batch_size=batch_size,\n    drop_last=False,\n    num_workers=2,\n    pin_memory=True,\n)\n\nvalid_dataset = RsnaDataset(valid_patient_ids, valid_labels, is_train=False)\nvalid_loader = DataLoader(\n    train_dataset,\n    sampler=SequentialSampler(train_dataset),\n    batch_size=batch_size,\n    drop_last=False,\n    num_workers=2,\n    pin_memory=True,\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:49:36.062303Z","iopub.execute_input":"2023-02-27T19:49:36.062693Z","iopub.status.idle":"2023-02-27T19:49:36.06996Z","shell.execute_reply.started":"2023-02-27T19:49:36.062661Z","shell.execute_reply":"2023-02-27T19:49:36.068871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_num_patches = 32\nimg_mean = 0\nimg_std = 1\nEPOCHES = 12","metadata":{"execution":{"iopub.status.busy":"2023-02-27T20:25:08.06375Z","iopub.execute_input":"2023-02-27T20:25:08.064153Z","iopub.status.idle":"2023-02-27T20:25:08.070384Z","shell.execute_reply.started":"2023-02-27T20:25:08.064117Z","shell.execute_reply":"2023-02-27T20:25:08.069417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for epoch in range(EPOCHES):\n    train_one_epoch(\n        train_df, IMG_PATH, patch_size, patches_per_in_iter, grad_acc_steps,\n        q1_model, q2_model, criterion, train_loader,\n        q1_optimizer, q2_optimizer, inner_iterations, \n        device, epoch, max_num_patches=max_num_patches,\n        img_mean=img_mean, img_std=img_std\n    )\n    q1_lr_scheduler.step(epoch)\n    q2_lr_scheduler.step(epoch)\n    \n    evaluate(\n        train_df, IMG_PATH, valid_loader,\n        q1_model, q2_model, patch_size, device,\n        max_num_patches=max_num_patches, img_mean=img_mean, img_std=img_std\n    )","metadata":{"execution":{"iopub.status.busy":"2023-02-27T20:25:08.388547Z","iopub.execute_input":"2023-02-27T20:25:08.388918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.save(q1_model, f\"/kaggle/working/fold-{FOLD_NUM}-q1_model_{epoch+1}.pth\")\ntorch.save(q2_model, f\"/kaggle/working/fold-{FOLD_NUM}-q2_model_{epoch+1}.pth\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for epoch in range(0, EPOCHES):\n    train_one_epoch(\n        train_df, IMG_PATH, patch_size, patches_per_in_iter, grad_acc_steps,\n        q1_model, q2_model, criterion, train_loader,\n        q1_optimizer, q2_optimizer, inner_iterations, \n        device, epoch\n    )\n    q1_lr_scheduler.step(epoch)\n    q2_lr_scheduler.step(epoch)\n    \n    evaluate(train_df, IMG_PATH, valid_loader, q1_model, q2_model, patch_size, device)\n    torch.save(q1_model, f\"/kaggle/working/fold-{FOLD_NUM}-q1_model_{epoch+1}.pth\")\n    torch.save(q2_model, f\"/kaggle/working/fold-{FOLD_NUM}-q2_model_{epoch+1}.pth\")","metadata":{"execution":{"iopub.status.busy":"2023-02-27T19:50:47.87607Z","iopub.execute_input":"2023-02-27T19:50:47.876522Z","iopub.status.idle":"2023-02-27T19:50:47.899461Z","shell.execute_reply.started":"2023-02-27T19:50:47.876482Z","shell.execute_reply":"2023-02-27T19:50:47.897027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}