{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Acknowledgment\nMain code and models in this notebook are from [Egor Trushin](https://www.kaggle.com/code/egortrushin/gr-icrgw-pytorch-lightning-baseline-unet-resnest) and [Shashwat Raman](https://www.kaggle.com/code/shashwatraman/simple-unet-baseline-infer-lb-0-580). Thanks to their amazing notebooks, I can try out ensemble here.\n\nThanks to [@imeintanis](https://www.kaggle.com/imeintanis)'s patient guidance, I've corrected a serious mistake.\n## Method Description\nI tried to apply **and operation** and **or operation** to combine the masks of different models, but the resulting was poor. In this notebook, I calculate the **weighted sum of probabilities** and then apply a threshold to get masks. \nNotice that there are K hyperparameters instead of K+1, where K is the number of models. This is because that I don't force the sum of weights equals to 1. In addition, I **scale their probabilities** based on the thresholds in their original notebooks.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Experiments\n\n|        | model_1 | model_2 |\n|--------|---------:|---------:|\n| **lb**     | 0.580   | 0.618   |\n\n\n| ratio_1  | ratio_2  | lb    |\n|----------:|----------:|-------:|\n| 0.3      | 0.7      | 0.626 |\n| 0.4      | 0.6      | 0.626 |\n| 0.45     | 0.55     | 0.626 |\n| 0.5      | 0.5      | 0.626 |\n| 0.55      | 0.45      | 0.587 |\n| 0.6      | 0.4      | 0.585 |\n\n\n**Lb remains 0.626 as long as the weight of the better model is greater than or equals to the weight of another model.\nThis pattern exists when I try to ensemble my model.**\n\n\n|        | model_1 | model_2 |\n|--------|---------:|---------:|\n| **lb**     | 0.552   | 0.580   |\n\n| ratio_1 | ratio_2 | lb    |\n|---------:|---------:|-------:|\n| 0.5     | 0.5     | 0.592 |\n| 0.6     | 0.5     | 0.589 |\n| 0.5     | 0.6     | 0.592 |\n\n<!-- |        | logical and | logical or |\n|--------|---------:|---------:|\n| **lb**     | 0.564   | 0.565   | -->","metadata":{}},{"cell_type":"markdown","source":"## Hyperparameters","metadata":{}},{"cell_type":"code","source":"ratio_1 = 0.55\nratio_2 = 0.45\nthr_1 = 0.02\nthr_2 = 0.5","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:44:58.319978Z","iopub.execute_input":"2023-06-17T09:44:58.320422Z","iopub.status.idle":"2023-06-17T09:44:58.325936Z","shell.execute_reply.started":"2023-06-17T09:44:58.320388Z","shell.execute_reply":"2023-06-17T09:44:58.324722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport os\nimport math\nimport random\nfrom PIL import Image\nfrom collections import defaultdict\n\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\nfrom tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt\n\nimport torch\nfrom torch import nn\nimport pytorch_lightning as pl\nfrom torchvision import transforms\nimport torchvision.transforms as T\nfrom torch.utils.data import Dataset, DataLoader\nfrom transformers import get_cosine_schedule_with_warmup\n\nimport sys\nsys.path.append(\"../input/pretrained-models-pytorch\")\nsys.path.append(\"../input/efficientnet-pytorch\")\nsys.path.append(\"/kaggle/input/smp-github/segmentation_models.pytorch-master\")\nsys.path.append(\"/kaggle/input/timm-pretrained-resnest/resnest/\")\nimport segmentation_models_pytorch as smp\n\nprint(f\"Segmentation Models version: {smp.__version__}\")\n\n!mkdir -p /root/.cache/torch/hub/checkpoints/\n!cp /kaggle/input/timm-pretrained-resnest/resnest/gluon_resnest26-50eb607c.pth /root/.cache/torch/hub/checkpoints/gluon_resnest26-50eb607c.pth","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:44:58.334227Z","iopub.execute_input":"2023-06-17T09:44:58.334622Z","iopub.status.idle":"2023-06-17T09:45:00.833939Z","shell.execute_reply.started":"2023-06-17T09:44:58.334591Z","shell.execute_reply":"2023-06-17T09:45:00.832657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## U_Net#1 [LB 0.580]","metadata":{}},{"cell_type":"code","source":"class Config:\n    batch_size = 32\n    seed = 42\n    thr = 0.02\n    \n    encoder = 'efficientnet-b0'\n    pretrained = False\n    weights = None\n    classes = ['contrail']\n    activation = None\n    in_chans = 3\n    \n    device = 'cuda' if torch.cuda.is_available() else 'cpu'\n    \n    image_size = 256\n    \n    model_ckpt = '/kaggle/input/unet-model/epoch-9.pth'\n    \nclass Paths:\n    data = '/kaggle/input/google-research-identify-contrails-reduce-global-warming'\n    data_root = '/kaggle/input/google-research-identify-contrails-reduce-global-warming/test/'","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:00.836315Z","iopub.execute_input":"2023-06-17T09:45:00.836955Z","iopub.status.idle":"2023-06-17T09:45:00.844983Z","shell.execute_reply.started":"2023-06-17T09:45:00.836917Z","shell.execute_reply":"2023-06-17T09:45:00.843724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed(seed=1234):\n    random.seed(seed)\n    np.random.seed(seed)\n    os.environ[\"PYTHONHASHSEED\"] = str(seed)\n    \n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = False\n    torch.backends.cudnn.benchmark = True\n    \nset_seed(Config.seed)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:00.846817Z","iopub.execute_input":"2023-06-17T09:45:00.84756Z","iopub.status.idle":"2023-06-17T09:45:01.056931Z","shell.execute_reply.started":"2023-06-17T09:45:00.847517Z","shell.execute_reply":"2023-06-17T09:45:01.055595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filenames = os.listdir(Paths.data_root)\ntest_df = pd.DataFrame(filenames, columns=['record_id'])\n\ntest_df['path'] = Paths.data_root + test_df['record_id'].astype(str)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.059802Z","iopub.execute_input":"2023-06-17T09:45:01.060152Z","iopub.status.idle":"2023-06-17T09:45:01.072827Z","shell.execute_reply.started":"2023-06-17T09:45:01.060125Z","shell.execute_reply":"2023-06-17T09:45:01.071756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class ContrailsDataset(torch.utils.data.Dataset):\n    def __init__(self, df, train=True):\n        \n        self.df = df\n        self.trn = train\n    \n    def read_record(self, directory):\n        record_data = {}\n        for x in [\n            \"band_11\", \n            \"band_14\", \n            \"band_15\"\n        ]:\n\n            record_data[x] = np.load(os.path.join(directory, x + \".npy\"))\n\n        return record_data\n\n    def normalize_range(self, data, bounds):\n        \"\"\"Maps data to the range [0, 1].\"\"\"\n        return (data - bounds[0]) / (bounds[1] - bounds[0])\n    \n    def get_false_color(self, record_data):\n        _T11_BOUNDS = (243, 303)\n        _CLOUD_TOP_TDIFF_BOUNDS = (-4, 5)\n        _TDIFF_BOUNDS = (-4, 2)\n        \n        N_TIMES_BEFORE = 4\n\n        r = self.normalize_range(record_data[\"band_15\"] - record_data[\"band_14\"], _TDIFF_BOUNDS)\n        g = self.normalize_range(record_data[\"band_14\"] - record_data[\"band_11\"], _CLOUD_TOP_TDIFF_BOUNDS)\n        b = self.normalize_range(record_data[\"band_14\"], _T11_BOUNDS)\n        false_color = np.clip(np.stack([r, g, b], axis=2), 0, 1)\n        img = false_color[..., N_TIMES_BEFORE]\n\n        return img\n    \n    def __getitem__(self, index):\n        row = self.df.iloc[index]\n        con_path = row.path\n        data = self.read_record(con_path)    \n        \n        img = self.get_false_color(data)\n        \n        img = torch.tensor(img)\n        img = img.permute(2, 0, 1)\n            \n        return img.float()\n    \n    def __len__(self):\n        return len(self.df)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.074173Z","iopub.execute_input":"2023-06-17T09:45:01.075049Z","iopub.status.idle":"2023-06-17T09:45:01.089242Z","shell.execute_reply.started":"2023-06-17T09:45:01.075018Z","shell.execute_reply":"2023-06-17T09:45:01.088357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_ds = ContrailsDataset(\n        test_df,\n        train = False\n    )\n \ntest_dl = DataLoader(test_ds, batch_size=Config.batch_size, num_workers = 2)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.090754Z","iopub.execute_input":"2023-06-17T09:45:01.091875Z","iopub.status.idle":"2023-06-17T09:45:01.108146Z","shell.execute_reply.started":"2023-06-17T09:45:01.09184Z","shell.execute_reply":"2023-06-17T09:45:01.106928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class UNet(nn.Module):\n    def __init__(self, cfg):\n        super(UNet, self).__init__()\n        \n        self.cfg = cfg\n        self.training = True\n        \n        self.model = smp.Unet(\n            encoder_name=cfg.encoder, \n            encoder_weights=cfg.weights, \n            decoder_use_batchnorm=True,\n            classes=len(cfg.classes), \n            activation=cfg.activation,\n        )\n        \n        self.loss_fn = smp.losses.DiceLoss(mode='binary')\n    \n    def forward(self, imgs):\n        \n        x = imgs\n        logits = self.model(x)\n        \n        return {\"logits\": logits.sigmoid()}","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.110591Z","iopub.execute_input":"2023-06-17T09:45:01.111339Z","iopub.status.idle":"2023-06-17T09:45:01.120517Z","shell.execute_reply.started":"2023-06-17T09:45:01.111255Z","shell.execute_reply":"2023-06-17T09:45:01.119438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = UNet(Config).to(Config.device)\nmodel.load_state_dict(torch.load(Config.model_ckpt, map_location=torch.device(Config.device)))","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.122733Z","iopub.execute_input":"2023-06-17T09:45:01.123064Z","iopub.status.idle":"2023-06-17T09:45:01.376832Z","shell.execute_reply.started":"2023-06-17T09:45:01.123036Z","shell.execute_reply":"2023-06-17T09:45:01.375562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.eval()\ntorch.set_grad_enabled(False)\n\nval_data = defaultdict(list)\npbar = tqdm(enumerate(test_dl), total=len(test_dl), desc='Test')\nfor step, X in pbar: \n    X = X.to(Config.device)\n\n    output = model(X)\n    for key, val in output.items():\n        val_data[key] += [output[key]]\n\nfor key, val in output.items():\n    value = val_data[key]\n    if len(value[0].shape) == 0:\n        val_data[key] = torch.stack(value)\n    else:\n        val_data[key] = torch.cat(value, dim=0).cpu().detach().numpy()","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.378087Z","iopub.execute_input":"2023-06-17T09:45:01.37849Z","iopub.status.idle":"2023-06-17T09:45:01.89474Z","shell.execute_reply.started":"2023-06-17T09:45:01.378456Z","shell.execute_reply":"2023-06-17T09:45:01.893505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def rle_encode(x, fg_val=1):\n    \"\"\"\n    Args:\n        x:  numpy array of shape (height, width), 1 - mask, 0 - background\n    Returns: run length encoding as list\n    \"\"\"\n\n    dots = np.where(\n        x.T.flatten() == fg_val)[0]  # .T sets Fortran order down-then-right\n    run_lengths = []\n    prev = -2\n    for b in dots:\n        if b > prev + 1:\n            run_lengths.extend((b + 1, 0))\n        run_lengths[-1] += 1\n        prev = b\n    return run_lengths\n\ndef list_to_string(x):\n    \"\"\"\n    Converts list to a string representation\n    Empty list returns '-'\n    \"\"\"\n    if x: # non-empty list\n        s = str(x).replace(\"[\", \"\").replace(\"]\", \"\").replace(\",\", \"\")\n    else:\n        s = '-'\n    return s","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.900327Z","iopub.execute_input":"2023-06-17T09:45:01.900714Z","iopub.status.idle":"2023-06-17T09:45:01.910071Z","shell.execute_reply.started":"2023-06-17T09:45:01.90068Z","shell.execute_reply":"2023-06-17T09:45:01.908941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Ids = []\nProbs_1 = []","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.911808Z","iopub.execute_input":"2023-06-17T09:45:01.912156Z","iopub.status.idle":"2023-06-17T09:45:01.925178Z","shell.execute_reply.started":"2023-06-17T09:45:01.91213Z","shell.execute_reply":"2023-06-17T09:45:01.92405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, pred in enumerate(val_data['logits']):\n    rec = test_df['record_id'][i]\n    Ids.append(rec)\n    Probs_1.append(pred[0])","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.926607Z","iopub.execute_input":"2023-06-17T09:45:01.927366Z","iopub.status.idle":"2023-06-17T09:45:01.938929Z","shell.execute_reply.started":"2023-06-17T09:45:01.927331Z","shell.execute_reply":"2023-06-17T09:45:01.937684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ndel filenames, test_df, test_ds, test_dl, model, val_data, pbar\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:01.940463Z","iopub.execute_input":"2023-06-17T09:45:01.941274Z","iopub.status.idle":"2023-06-17T09:45:02.371771Z","shell.execute_reply.started":"2023-06-17T09:45:01.941216Z","shell.execute_reply":"2023-06-17T09:45:02.37093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## U_Net#2 [LB 0.618]","metadata":{}},{"cell_type":"code","source":"batch_size = 16\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\ndata = '/kaggle/input/google-research-identify-contrails-reduce-global-warming'\ndata_root = '/kaggle/input/google-research-identify-contrails-reduce-global-warming/test/'","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:02.372957Z","iopub.execute_input":"2023-06-17T09:45:02.373905Z","iopub.status.idle":"2023-06-17T09:45:02.378916Z","shell.execute_reply.started":"2023-06-17T09:45:02.373871Z","shell.execute_reply":"2023-06-17T09:45:02.377824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filenames = os.listdir(data_root)\ntest_df = pd.DataFrame(filenames, columns=['record_id'])\ntest_df['path'] = data_root + test_df['record_id'].astype(str)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:02.380609Z","iopub.execute_input":"2023-06-17T09:45:02.38098Z","iopub.status.idle":"2023-06-17T09:45:02.398252Z","shell.execute_reply.started":"2023-06-17T09:45:02.38095Z","shell.execute_reply":"2023-06-17T09:45:02.396727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class ContrailsDataset(torch.utils.data.Dataset):\n    def __init__(self, df, train=True):\n        \n        self.df = df\n        self.trn = train\n        self.df_idx: pd.DataFrame = pd.DataFrame({'idx': os.listdir(f'/kaggle/input/google-research-identify-contrails-reduce-global-warming/test')})\n        self.normalize_image = T.Normalize((0.485, 0.456, 0.406), (0.229, 0.224, 0.225))\n    \n    def read_record(self, directory):\n        record_data = {}\n        for x in [\n            \"band_11\", \n            \"band_14\", \n            \"band_15\"\n        ]:\n\n            record_data[x] = np.load(os.path.join(directory, x + \".npy\"))\n\n        return record_data\n\n    def normalize_range(self, data, bounds):\n        \"\"\"Maps data to the range [0, 1].\"\"\"\n        return (data - bounds[0]) / (bounds[1] - bounds[0])\n    \n    def get_false_color(self, record_data):\n        _T11_BOUNDS = (243, 303)\n        _CLOUD_TOP_TDIFF_BOUNDS = (-4, 5)\n        _TDIFF_BOUNDS = (-4, 2)\n        \n        N_TIMES_BEFORE = 4\n\n        r = self.normalize_range(record_data[\"band_15\"] - record_data[\"band_14\"], _TDIFF_BOUNDS)\n        g = self.normalize_range(record_data[\"band_14\"] - record_data[\"band_11\"], _CLOUD_TOP_TDIFF_BOUNDS)\n        b = self.normalize_range(record_data[\"band_14\"], _T11_BOUNDS)\n        false_color = np.clip(np.stack([r, g, b], axis=2), 0, 1)\n        img = false_color[..., N_TIMES_BEFORE]\n\n        return img\n    \n    def __getitem__(self, index):\n        row = self.df.iloc[index]\n        con_path = row.path\n        data = self.read_record(con_path)    \n        \n        img = self.get_false_color(data)\n        \n        img = torch.tensor(np.reshape(img, (256, 256, 3))).to(torch.float32).permute(2, 0, 1)\n        \n        img = self.normalize_image(img)\n        \n        image_id = int(self.df_idx.iloc[index]['idx'])\n            \n        return img.float(), torch.tensor(image_id)\n    \n    def __len__(self):\n        return len(self.df)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:02.400571Z","iopub.execute_input":"2023-06-17T09:45:02.400974Z","iopub.status.idle":"2023-06-17T09:45:02.417728Z","shell.execute_reply.started":"2023-06-17T09:45:02.400938Z","shell.execute_reply":"2023-06-17T09:45:02.416556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_ds = ContrailsDataset(\n        test_df,\n        train = False\n    )\n \ntest_dl = DataLoader(test_ds, batch_size=batch_size, num_workers = 1)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:02.419398Z","iopub.execute_input":"2023-06-17T09:45:02.420692Z","iopub.status.idle":"2023-06-17T09:45:02.438218Z","shell.execute_reply.started":"2023-06-17T09:45:02.420654Z","shell.execute_reply":"2023-06-17T09:45:02.437057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class LightningModule(pl.LightningModule):\n\n    def __init__(self):\n        super().__init__()\n        self.model = smp.Unet(encoder_name=\"timm-resnest26d\",\n                              encoder_weights=None,\n                              in_channels=3,\n                              classes=1,\n                              activation=None,\n                              )\n\n    def forward(self, batch):\n        return self.model(batch)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:02.440062Z","iopub.execute_input":"2023-06-17T09:45:02.440445Z","iopub.status.idle":"2023-06-17T09:45:02.453983Z","shell.execute_reply.started":"2023-06-17T09:45:02.440414Z","shell.execute_reply":"2023-06-17T09:45:02.452774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if torch.cuda.is_available():\n    model = LightningModule().load_from_checkpoint(\"/kaggle/input/models/model.ckpt\")\nelse:\n    model = LightningModule().load_from_checkpoint(\"/kaggle/input/models/model.ckpt\", map_location=torch.device('cpu'))\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel.to(device)\nmodel.eval()\nmodel.zero_grad()","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:02.455343Z","iopub.execute_input":"2023-06-17T09:45:02.455728Z","iopub.status.idle":"2023-06-17T09:45:03.877627Z","shell.execute_reply.started":"2023-06-17T09:45:02.455684Z","shell.execute_reply":"2023-06-17T09:45:03.876697Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Probs_2 = []","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:03.879227Z","iopub.execute_input":"2023-06-17T09:45:03.879875Z","iopub.status.idle":"2023-06-17T09:45:03.883805Z","shell.execute_reply.started":"2023-06-17T09:45:03.879843Z","shell.execute_reply":"2023-06-17T09:45:03.882915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, data in enumerate(test_dl):\n    images, image_id = data\n    \n    # Predict mask for this instance\n    images = images.to(device)\n    predicated_mask = model.forward(images[:, :, :, :])\n    predicated_mask = torch.sigmoid(predicated_mask).cpu().detach().numpy()\n    \n    # Apply threshold\n#     predicated_mask_with_threshold = np.zeros((images.shape[0], 256, 256))\n#     predicated_mask_with_threshold[predicated_mask[:, 0, :, :] < 0.5] = 0\n#     predicated_mask_with_threshold[predicated_mask[:, 0, :, :] > 0.5] = 1\n    \n    for img_num in range(0, images.shape[0]):\n        current_mask = predicated_mask[img_num, :, :]\n        current_image_id = image_id[img_num].item()\n        \n        Probs_2.append(current_mask)","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:03.885313Z","iopub.execute_input":"2023-06-17T09:45:03.885952Z","iopub.status.idle":"2023-06-17T09:45:04.697563Z","shell.execute_reply.started":"2023-06-17T09:45:03.88592Z","shell.execute_reply":"2023-06-17T09:45:04.695907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del filenames, test_df, test_ds, model\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:04.6998Z","iopub.execute_input":"2023-06-17T09:45:04.700554Z","iopub.status.idle":"2023-06-17T09:45:05.140176Z","shell.execute_reply.started":"2023-06-17T09:45:04.700513Z","shell.execute_reply":"2023-06-17T09:45:05.139032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Ensemble & Submission","metadata":{}},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/google-research-identify-contrails-reduce-global-warming/sample_submission.csv', index_col='record_id')","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:05.141824Z","iopub.execute_input":"2023-06-17T09:45:05.142275Z","iopub.status.idle":"2023-06-17T09:45:05.151922Z","shell.execute_reply.started":"2023-06-17T09:45:05.142232Z","shell.execute_reply":"2023-06-17T09:45:05.150865Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def ensemble(p1, p2):\n    w1 = ratio_1 * (1 / thr_1)\n    w2 = ratio_2 * (1 / thr_2)\n    thr = 1  # fixed \n    p = p1*w1 + p2*w2\n    p = (p > thr).astype(np.int32)\n    return p","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:05.153267Z","iopub.execute_input":"2023-06-17T09:45:05.15361Z","iopub.status.idle":"2023-06-17T09:45:05.166879Z","shell.execute_reply.started":"2023-06-17T09:45:05.153582Z","shell.execute_reply":"2023-06-17T09:45:05.165702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for img_id, p1, p2 in zip(Ids, Probs_1, Probs_2):\n    submission.loc[int(img_id), 'encoded_pixels'] = list_to_string(rle_encode(ensemble(p1, p2)))","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:05.16862Z","iopub.execute_input":"2023-06-17T09:45:05.168953Z","iopub.status.idle":"2023-06-17T09:45:05.185473Z","shell.execute_reply.started":"2023-06-17T09:45:05.168926Z","shell.execute_reply":"2023-06-17T09:45:05.184219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:05.186603Z","iopub.execute_input":"2023-06-17T09:45:05.186918Z","iopub.status.idle":"2023-06-17T09:45:05.201204Z","shell.execute_reply.started":"2023-06-17T09:45:05.186892Z","shell.execute_reply":"2023-06-17T09:45:05.200022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2023-06-17T09:45:05.20299Z","iopub.execute_input":"2023-06-17T09:45:05.203466Z","iopub.status.idle":"2023-06-17T09:45:05.216398Z","shell.execute_reply.started":"2023-06-17T09:45:05.203426Z","shell.execute_reply":"2023-06-17T09:45:05.214943Z"},"trusted":true},"execution_count":null,"outputs":[]}]}