{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# How to create a HDF5 Dataset\nHey guys! As already mentioned in the discussions <br>\n( https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/409401 )<br>\n, **HDF5** ist a great data format to manage such huge datasets and achieve fast read access.\n\nThat's why I wrote some code to convert the dataset to HDF5. As I couldn't find a way to upload datasets of this size (300+GB) on kaggle, this notebook only processes the first 2500 instances as an example. I'm currently working on finding a nice way to share the full HDF5-file.\n<br><br>\nI also want to mention that this is V1, which means **it is only a prototype**. I'm not an HDF5 expert myself, and I will constantly improve this version. So if you have any advice, **feel free to help!!!**\nSome things I want to implement soon are for example parallel processing, fundamental preprocessing, and a better dataloader. \n<br><br>\nTo demonstrate the perfomance speedup, I used this Pytorch dataloader: <br>\nhttps://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example\n<br>So **check out this notebook** for more details about this part!!\n","metadata":{"execution":{"iopub.status.busy":"2023-05-13T20:50:40.341203Z","iopub.execute_input":"2023-05-13T20:50:40.341615Z","iopub.status.idle":"2023-05-13T20:50:44.968488Z","shell.execute_reply.started":"2023-05-13T20:50:40.341583Z","shell.execute_reply":"2023-05-13T20:50:44.967254Z"}}},{"cell_type":"markdown","source":"### Versions: \nV1: <br>- 2x faster with num_workers = 1 <br>\n    - Problems with num_workers = 4, almost no speedup <br>\n    (Probably the bottleneck is opening the hdf5-file) <br>\n    \nV2: <br>- Now dtype float16 is used since we don't lose much information, as mentioned in discussions <br>\n    - hdf5-file isn't opened in the get_item method but in the constructor. <br>\n    - Much faster with num_workers = 1 <br>\n    - 1.5x faster with num_workers = 4","metadata":{}},{"cell_type":"code","source":"# Imports\nimport numpy as np\nimport pandas as pd\nimport os\nimport json\nimport h5py\nfrom tqdm import tqdm\nimport torch\nfrom torch.utils.data import Dataset\nfrom torch.utils.data import DataLoader","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:07.753843Z","iopub.execute_input":"2023-05-14T10:18:07.754505Z","iopub.status.idle":"2023-05-14T10:18:07.762153Z","shell.execute_reply.started":"2023-05-14T10:18:07.754441Z","shell.execute_reply":"2023-05-14T10:18:07.761087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Environment setup and creating Dataset folder\n\nFirst, I need to configure the Kaggle environment to save the large output to a Kaggle dataset. Shoutout to:\n\nhttps://www.kaggle.com/code/xhlulu/how-to-create-very-large-datasets-from-a-notebook/notebook","metadata":{}},{"cell_type":"code","source":"from kaggle_secrets import UserSecretsClient\nsecrets = UserSecretsClient()\n\nos.environ['KAGGLE_USERNAME'] = secrets.get_secret(\"KAGGLE_USERNAME\")\nos.environ['KAGGLE_KEY'] = secrets.get_secret(\"KAGGLE_KEY\")","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:07.76427Z","iopub.execute_input":"2023-05-14T10:18:07.764791Z","iopub.status.idle":"2023-05-14T10:18:08.270733Z","shell.execute_reply.started":"2023-05-14T10:18:07.764757Z","shell.execute_reply":"2023-05-14T10:18:08.269297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# make sure dataset dir does not exist\n!rm -rf /kaggle/dataset/","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:08.272244Z","iopub.execute_input":"2023-05-14T10:18:08.272645Z","iopub.status.idle":"2023-05-14T10:18:10.07219Z","shell.execute_reply.started":"2023-05-14T10:18:08.272613Z","shell.execute_reply":"2023-05-14T10:18:10.070637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.makedirs('/kaggle/dataset/', exist_ok=True)\n\n\n# Change below\nmeta = dict(\n    id=f\"{os.getenv('KAGGLE_USERNAME')}/hdf5-dataset\",\n    title=\"Google Contrails HDF5 Dataset\",\n    isPrivate=True,\n    licenses=[dict(name=\"other\")]\n)\n\nwith open('/kaggle/dataset/dataset-metadata.json', 'w') as f:\n    json.dump(meta, f)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.076785Z","iopub.execute_input":"2023-05-14T10:18:10.077739Z","iopub.status.idle":"2023-05-14T10:18:10.090587Z","shell.execute_reply.started":"2023-05-14T10:18:10.077674Z","shell.execute_reply":"2023-05-14T10:18:10.089378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating HDF5 File\n\nnow I create the HDF5 file. Each instance gets its own hdf5 group, and all bands are concatenated into a single array (changing dtype to float16)\n<br><br>\n(As mentioned earlier, I'm using only a small portion of the dataset for demonstration since I couldn't find a way to save such big datasets on Kaggle.)\n","metadata":{}},{"cell_type":"code","source":"BASE_DIR = \"/kaggle/input/google-research-identify-contrails-reduce-global-warming\"\nHDF_DIR = \"/kaggle/dataset\"","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.092035Z","iopub.execute_input":"2023-05-14T10:18:10.093139Z","iopub.status.idle":"2023-05-14T10:18:10.111796Z","shell.execute_reply.started":"2023-05-14T10:18:10.093051Z","shell.execute_reply":"2023-05-14T10:18:10.110453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_ids = os.listdir(BASE_DIR +'/'+ \"train\")[:2500]\nval_ids = os.listdir(BASE_DIR +'/'+ \"validation\")[:100]\ntest_ids = os.listdir(BASE_DIR +'/'+ \"test\")","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.113275Z","iopub.execute_input":"2023-05-14T10:18:10.11442Z","iopub.status.idle":"2023-05-14T10:18:10.148691Z","shell.execute_reply.started":"2023-05-14T10:18:10.114368Z","shell.execute_reply":"2023-05-14T10:18:10.147398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"num train samples: \\t\\t{len(train_ids)}\")\nprint(f\"num validation samples: \\t{len(val_ids)}\")\nprint(f\"num test samples: \\t\\t{len(test_ids)}\")","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.150208Z","iopub.execute_input":"2023-05-14T10:18:10.150695Z","iopub.status.idle":"2023-05-14T10:18:10.157185Z","shell.execute_reply.started":"2023-05-14T10:18:10.150663Z","shell.execute_reply":"2023-05-14T10:18:10.155943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Note:** Every Instance saves arrays for band_08-16 and two arrays for the \"ground truth\". <br>\nValidation instances dont have human_individual_masks.npy and test instances obviously don't have ground truth arrays:","metadata":{}},{"cell_type":"code","source":"sorted(os.listdir(BASE_DIR + \"/train/\" + train_ids[0]))","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.159004Z","iopub.execute_input":"2023-05-14T10:18:10.159909Z","iopub.status.idle":"2023-05-14T10:18:10.182979Z","shell.execute_reply.started":"2023-05-14T10:18:10.159863Z","shell.execute_reply":"2023-05-14T10:18:10.182128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def saveToHDF5(id_, type_, group):\n    \"\"\"Save the train/validation/test (-> type_) instance in a unique subgroup of group\"\"\"\n    \n    if type_ not in [\"train\", \"validation\", \"test\"]:\n        raise ValueError(\"type has do be one of ['train', 'validation', 'test']\")\n    \n    # instance is the new subgroup where we put the data of id_\n    instance = group.create_group(id_)\n    path = BASE_DIR + f\"/{type_}/\" + id_\n    \n    # All bands are concatenated to one array\n    bands_data = []\n    for i in range(8, 17):\n        band = f\"band_{str(i).zfill(2)}\"\n        array = np.load(f\"{path}/{band}.npy\")\n        bands_data.append(array)\n    bands_data = np.stack(bands_data, axis=-1).astype(np.float16) \n    instance.create_dataset(\"bands_data\", data=bands_data)\n    \n    # Only train and validation have human_pixel_masks\n    if type_ in [\"train\", \"validation\"]:\n        agg_masks = np.load(path + \"/human_pixel_masks.npy\").astype(np.float16)\n        instance.create_dataset(\"human_pixel_masks\", data=agg_masks)\n    if type_ == \"train\":\n        ind_masks = np.load(path + \"/human_individual_masks.npy\").astype(np.float16)\n        instance.create_dataset(\"human_individual_masks\", data = ind_masks)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.184468Z","iopub.execute_input":"2023-05-14T10:18:10.185598Z","iopub.status.idle":"2023-05-14T10:18:10.196199Z","shell.execute_reply.started":"2023-05-14T10:18:10.185561Z","shell.execute_reply":"2023-05-14T10:18:10.195092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the HDF5-file\nwith h5py.File(HDF_DIR + \"/Google_Contrails.hdf5\", \"w\") as f:\n    \n    print(\"Processing Training Data...\")\n    train = f.create_group(\"train\")\n    for train_id in tqdm(train_ids):\n        saveToHDF5(train_id, \"train\", train)\n        \n    print(\"\\nProcessing Validation Data...\")\n    validation = f.create_group(\"validation\")\n    for val_id in tqdm(val_ids):\n        saveToHDF5(val_id, \"validation\", validation)\n    \n    print(\"\\nProcessing Test Data...\")\n    test = f.create_group(\"test\")\n    for test_id in tqdm(test_ids):\n        saveToHDF5(test_id, \"test\", test)\n        \n    print(\"\\nFinished all jobs!\")  ","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:18:10.200145Z","iopub.execute_input":"2023-05-14T10:18:10.200901Z","iopub.status.idle":"2023-05-14T10:31:31.513464Z","shell.execute_reply.started":"2023-05-14T10:18:10.200853Z","shell.execute_reply":"2023-05-14T10:31:31.511936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_dir_size(path='.'):\n    total = 0\n    with os.scandir(path) as it:\n        for entry in it:\n            if entry.is_file():\n                total += entry.stat().st_size\n            elif entry.is_dir():\n                total += get_dir_size(entry.path)\n    return total\n\ngb = get_dir_size(HDF_DIR) * 10**(-9)\nnum = len(train_ids) + len(val_ids) + len(test_ids)\nprint(f\"HDF5 file with {num} instances uses {gb:.5}GB of memory\")","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:31:31.520899Z","iopub.execute_input":"2023-05-14T10:31:31.521391Z","iopub.status.idle":"2023-05-14T10:31:31.533373Z","shell.execute_reply.started":"2023-05-14T10:31:31.521351Z","shell.execute_reply":"2023-05-14T10:31:31.531936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Now let's compare the speedup!<br>\nAs mentioned earlier, I'm comparing the speedup by iterating over a dataloader once. I used this dataloader:<br>\nhttps://www.kaggle.com/code/thomasrochefort/pytorch-dataloader-example \n<br>\nFor the HDF5-Dataloader, I had to make some small adjustments.","metadata":{}},{"cell_type":"markdown","source":"### 1) No HDF5","metadata":{}},{"cell_type":"code","source":"class ContrailDataset(Dataset):\n    def __init__(self, base_dir, data_type='train', transform=None):\n        assert data_type in ['train', 'validation', 'test'], \\\n            \"'data_type' should be one of 'train', 'validation', or 'test'\"\n\n        self.base_dir = base_dir\n        self.data_type = data_type\n        self.transform = transform\n        self.record = os.listdir(self.base_dir +'/'+ self.data_type)[:2500]\n\n    def __len__(self):\n        return len(self.record)\n\n    def __getitem__(self, idx):\n        record_id = self.record[idx]\n        record_dir = os.path.join(self.base_dir, self.data_type, record_id)\n\n        # Load the necessary .npy files\n        bands_data = []\n        for i in range(8, 17):\n            band_file = os.path.join(record_dir, f'band_{str(i).zfill(2)}.npy')\n            band_data = np.load(band_file)\n            bands_data.append(band_data)\n\n        # Stack band data along the channel axis\n        bands_data = np.stack(bands_data, axis=-1)\n\n        # If the data type is 'train' or 'validation', load the masks\n        if self.data_type in ['train', 'validation']:\n            pixel_masks_file = os.path.join(record_dir, 'human_pixel_masks.npy')\n            pixel_masks = np.load(pixel_masks_file)\n        else:\n            pixel_masks = None  # No masks for 'test' data\n\n            \n        return bands_data, pixel_masks\n\n\ndef get_dataloader(base_dir, data_type, batch_size, transform=None):\n    dataset = ContrailDataset(base_dir, data_type=data_type, transform=transform)\n    dataloader = DataLoader(dataset, batch_size=batch_size, num_workers = 4)\n    return dataloader","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:41:44.603719Z","iopub.execute_input":"2023-05-14T10:41:44.604357Z","iopub.status.idle":"2023-05-14T10:41:44.620368Z","shell.execute_reply.started":"2023-05-14T10:41:44.604285Z","shell.execute_reply":"2023-05-14T10:41:44.619154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataloader = get_dataloader(BASE_DIR, 'train', batch_size=16)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:41:44.786744Z","iopub.execute_input":"2023-05-14T10:41:44.787516Z","iopub.status.idle":"2023-05-14T10:41:44.980975Z","shell.execute_reply.started":"2023-05-14T10:41:44.787451Z","shell.execute_reply":"2023-05-14T10:41:44.979869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for bands, masks in tqdm(train_dataloader):\n    continue","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:41:45.265384Z","iopub.execute_input":"2023-05-14T10:41:45.266692Z","iopub.status.idle":"2023-05-14T10:53:12.985666Z","shell.execute_reply.started":"2023-05-14T10:41:45.266643Z","shell.execute_reply":"2023-05-14T10:53:12.98414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2) With HDF5","metadata":{}},{"cell_type":"code","source":"class ContrailDatasetHDF5(Dataset):\n    def __init__(self, hdf5_file, data_type='train', transform=None):\n        assert data_type in ['train', 'validation', 'test'], \\\n            \"'data_type' should be one of 'train', 'validation', or 'test'\"\n\n        self.hdf5_file = hdf5_file\n        self.data_type = data_type\n        self.transform = transform\n        # Open the HDF5 file and get the appropriate data group\n        f = h5py.File(hdf5_file, \"r\")\n        #with h5py.File(self.hdf5_file, \"r\") as f:\n        self.group = f[self.data_type]\n        \n        # Get a list of the instance IDs in the data group\n        self.records = list(self.group.keys())\n\n    def __len__(self):\n        return len(self.records)\n\n    def __getitem__(self, idx):\n        record_id = self.records[idx]\n        bands = self.group[f\"{record_id}/bands_data\"][()]\n        if self.data_type in ['train', 'validation']:\n            pixel_masks = self.group[f\"{record_id}/human_pixel_masks\"][()]\n        else: \n            pixel_masks = None\n        return bands, pixel_masks\n    \n\ndef get_dataloader_hdf5(path, data_type, batch_size, transform=None):\n    dataset = ContrailDatasetHDF5(path, data_type=data_type, transform=transform)\n    dataloader = DataLoader(dataset, batch_size=batch_size, num_workers=4)\n    return dataloader","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:57:34.371435Z","iopub.execute_input":"2023-05-14T10:57:34.372894Z","iopub.status.idle":"2023-05-14T10:57:34.387213Z","shell.execute_reply.started":"2023-05-14T10:57:34.372838Z","shell.execute_reply":"2023-05-14T10:57:34.385709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path = HDF_DIR + \"/Google_Contrails.hdf5\"\ntrain_dataloader_hdf5 = get_dataloader_hdf5(path, 'train', batch_size=16)","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:57:34.507477Z","iopub.execute_input":"2023-05-14T10:57:34.508338Z","iopub.status.idle":"2023-05-14T10:57:34.515628Z","shell.execute_reply.started":"2023-05-14T10:57:34.508274Z","shell.execute_reply":"2023-05-14T10:57:34.513762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for bands, masks in tqdm(train_dataloader_hdf5):\n    continue","metadata":{"execution":{"iopub.status.busy":"2023-05-14T10:57:35.608514Z","iopub.execute_input":"2023-05-14T10:57:35.608974Z","iopub.status.idle":"2023-05-14T11:04:01.625435Z","shell.execute_reply.started":"2023-05-14T10:57:35.608941Z","shell.execute_reply":"2023-05-14T11:04:01.623948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Results\nAs you can see, the speedup is by a factor of 1.5. I'm confident we can do much better, and there are many things to improve. But the HDF5 file isn't just faster in reading. We can do our preprocessing to save even more time, and the bands are now already concatenated\n<br><br>\nIt would be great if somebody could tell me if there is a way to upload the full HDF5-file (300+GB) to kaggle. Otherwise, I will share a link to my Google Drive. <br><br><br>\nPS: If you want to save the dataset folder of this notebook, use: <br>\n!kaggle datasets create -p \"/kaggle/dataset\" --dir-mode zip","metadata":{}}]}