{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":10200,"databundleVersionId":868375,"sourceType":"competition"}],"dockerImageVersionId":30746,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Doodle Classifier Dataset Preparation\n\n![Doodle Classifier](https://storage.googleapis.com/gweb-cloudblog-publish/images/quick-draw-1u639.max-900x900.PNG)\n\n## Overview\n\nThis notebook prepares a dataset for a doodle classifier. It involves loading and filtering recognized doodle drawings from multiple CSV files, sampling a fixed number of images per class, converting the drawing data from JSON strings to Python objects, and saving the processed images in an organized format.\n\n## Steps\n\n### 1. Loading and Filtering Data\n\n- Load each CSV file containing doodle drawings.\n- Filter to retain only the recognized doodles.\n- Sample a fixed number of images from each class.\n\n### 2. Data Conversion\n\n- Convert the drawing data from JSON format to Python objects.\n\n### 3. Image Generation\n\n- Plot the doodles using Matplotlib.\n- Convert the plots to grayscale images.\n- Save the images using OpenCV.\n\n### 4. Parallel Processing\n\n- Utilize `joblib.Parallel` to speed up the data loading and image processing tasks.\n\n### 5. Organized Storage\n\n- Store the processed images in a structured directory format, with each class having its own folder.\n- Save metadata about the images in a CSV file.\n\n### 6. Compression\n\n- Compress the directory containing the processed images into a ZIP file for easy distribution.\n\n## Functions\n\n### `get_samples(file_path, master_df, no_of_images=5000)`\n\n- Loads data from a CSV file.\n- Filters recognized doodles.\n- Samples a fixed number of images.\n- Concatenates the results into a master dataframe.\n\n### `apply_parallel(df, func, n_jobs=-1)`\n\n- Applies a function in parallel across a dataframe using `joblib.Parallel`.\n\n### `get_imgs(doodle_data)`\n\n- Converts doodle data into grayscale images by plotting the strokes with Matplotlib and converting the plot to an image.\n\n### `save_img(doodle_data, image_path)`\n\n- Saves doodle data as an image to the specified path.\n\n## Usage\n\n1. **Data Loading**:\n    ```python\n    master_df = pd.DataFrame()\n    file_paths = [os.path.join(path, i) for i in os.listdir(path)]\n    for file_path in tqdm(file_paths):\n        master_df = get_samples(file_path, master_df, no_of_images=1000)\n    ```\n\n2. **Data Conversion**:\n    ```python\n    master_df['drawing_pr'] = master_df['drawing'].progress_apply(json.loads)\n    ```\n\n3. **Image Generation**:\n    ```python\n    master_df['drawing_img'] = apply_parallel(master_df['drawing_pr'], get_imgs)\n    ```\n\n4. **Image Saving**:\n    ```python\n    for index, row in master_df.iterrows():\n        save_img(row['drawing_pr'], row['image_path'])\n    ```\n\n5. **Compression**:\n    ```python\n    shutil.make_archive('doodle_images', 'zip', 'data')\n    ```","metadata":{}},{"cell_type":"code","source":"import os\nimport cv2\nimport ast\nimport json\nimport shutil\nimport numpy as np\nimport pandas as pd\nfrom tqdm.notebook import tqdm\n\nfrom joblib import Parallel, delayed\n\nimport matplotlib.pyplot as plt\n\npath = '/kaggle/input/quickdraw-doodle-recognition/train_simplified/'","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-08-04T09:24:43.853159Z","iopub.execute_input":"2024-08-04T09:24:43.853632Z","iopub.status.idle":"2024-08-04T09:24:43.86081Z","shell.execute_reply.started":"2024-08-04T09:24:43.853594Z","shell.execute_reply":"2024-08-04T09:24:43.859359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 1. Making Single CSV with Fix Number of Samples\n\n#### 1.1) One by One","metadata":{}},{"cell_type":"code","source":"master_df = pd.DataFrame()\n\ndef get_samples(file_path, master_df, no_of_images=5000):\n    df_sim = pd.read_csv(file_path)\n    df_sim = df_sim[df_sim['recognized'] == True]\n    df_sim = df_sim.sample(frac=1).reset_index(drop=True)\n    df_sim = df_sim.head(no_of_images)\n\n    master_df = pd.concat([master_df, df_sim], ignore_index=True)\n\n    return master_df\n\nfile_paths = [os.path.join(path, i) for i in os.listdir(path)]\n\n\n# for file_path in tqdm(file_paths):\n#     master_df = get_samples(file_path, master_df, no_of_images=1000)\n\n# print(master_df.head())","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:08:08.917116Z","iopub.execute_input":"2024-08-03T03:08:08.917582Z","iopub.status.idle":"2024-08-03T03:08:09.038006Z","shell.execute_reply.started":"2024-08-03T03:08:08.917544Z","shell.execute_reply":"2024-08-03T03:08:09.036773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.2) Multiprocessing","metadata":{}},{"cell_type":"code","source":"dataframes = []\nno_of_images = 3000\npath = '/kaggle/input/quickdraw-doodle-recognition/train_simplified/'\n\n\ndef get_samples(file_path, no_of_images=5000):\n    df_sim = pd.read_csv(file_path)\n    df_sim = df_sim[df_sim['recognized'] == True]\n    df_sim = df_sim.sample(frac=1).reset_index(drop=True)\n    df_sim = df_sim.head(no_of_images)\n    return df_sim\n\nfile_paths = [os.path.join(path, i) for i in os.listdir(path)]\n\nresults = Parallel(n_jobs=-1)(delayed(get_samples)(file_path, no_of_images = no_of_images) for file_path in tqdm(file_paths))\n\nmaster_df = pd.concat(results, ignore_index=True)\n\nprint(f'Total samples in master dataframe: {len(master_df)}')\n\nmaster_df.to_csv('master_doodle_dataframe.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:08:12.905302Z","iopub.execute_input":"2024-08-03T03:08:12.905727Z","iopub.status.idle":"2024-08-03T03:11:51.032389Z","shell.execute_reply.started":"2024-08-03T03:08:12.905693Z","shell.execute_reply":"2024-08-03T03:11:51.030589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.3) Loading the Strokes","metadata":{}},{"cell_type":"code","source":"tqdm.pandas()\n\nmaster_df['drawing'] = master_df['drawing'].progress_apply(json.loads)\n\ndel master_df['timestamp']\n\nmaster_df.to_csv('master_doodle_dataframe.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:11:51.035643Z","iopub.execute_input":"2024-08-03T03:11:51.036175Z","iopub.status.idle":"2024-08-03T03:13:10.797936Z","shell.execute_reply.started":"2024-08-03T03:11:51.036111Z","shell.execute_reply":"2024-08-03T03:13:10.795686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.4) Sample Image","metadata":{}},{"cell_type":"code","source":"doodle_data = master_df['drawing'][100]\n\nfor stroke in doodle_data:\n    x, y = stroke\n    \n    plt.plot(x,y)\n\nplt.gca().invert_yaxis()\nplt.axis('equal')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:13:10.801348Z","iopub.execute_input":"2024-08-03T03:13:10.80202Z","iopub.status.idle":"2024-08-03T03:13:11.123875Z","shell.execute_reply.started":"2024-08-03T03:13:10.801971Z","shell.execute_reply":"2024-08-03T03:13:11.122401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.5) Creating Images from Strokes","metadata":{}},{"cell_type":"code","source":"def get_imgs(doodle_data):\n    \n    fig, ax = plt.subplots(figsize=(2.55, 2.55))\n\n    for stroke in doodle_data:\n        x, y = stroke  \n        plt.plot(x,y)\n\n    plt.gca().invert_yaxis()\n    ax.axis('off')\n\n    fig.canvas.draw()\n    image = np.frombuffer(fig.canvas.tostring_rgb(), dtype=np.uint8)\n    image = image.reshape(fig.canvas.get_width_height()[::-1] + (3,))\n    plt.close(fig)\n\n    image_gray= cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n\n    return image_gray\n\nplt.imshow(get_imgs(master_df['drawing'][100]),cmap = 'gray')\n\n# tqdm.pandas()\n# df_sim['drawing_img'] = master_df['drawing_pr'].progress_apply(get_imgs)","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:13:11.126593Z","iopub.execute_input":"2024-08-03T03:13:11.126993Z","iopub.status.idle":"2024-08-03T03:13:11.494614Z","shell.execute_reply.started":"2024-08-03T03:13:11.12696Z","shell.execute_reply":"2024-08-03T03:13:11.493356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.6) Loading Images in Parallel","metadata":{}},{"cell_type":"code","source":"def get_imgs(doodle_data):\n    fig, ax = plt.subplots(figsize=(2.55, 2.55))\n\n    for stroke in doodle_data:\n        x, y = stroke  \n        plt.plot(x, y)\n\n    plt.gca().invert_yaxis()\n    ax.axis('off')\n\n    fig.canvas.draw()\n    image = np.frombuffer(fig.canvas.tostring_rgb(), dtype=np.uint8)\n    image = image.reshape(fig.canvas.get_width_height()[::-1] + (3,))\n    plt.close(fig)\n\n    image_gray = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n\n    return image_gray\n\n# tqdm.pandas()\n\n# def apply_parallel(df, func, n_jobs=-1):\n#     results = Parallel(n_jobs=n_jobs)(delayed(func)(doodle) for doodle in tqdm(df))\n#     return results\n\n# # master_df['drawing_img'] = apply_parallel(master_df['drawing'], get_imgs)\n\n# # master_df.to_csv('master_doodle_dataframe.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:13:11.496Z","iopub.execute_input":"2024-08-03T03:13:11.496407Z","iopub.status.idle":"2024-08-03T03:13:11.505193Z","shell.execute_reply.started":"2024-08-03T03:13:11.496363Z","shell.execute_reply":"2024-08-03T03:13:11.504023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.7) Creating Folder for each Class","metadata":{}},{"cell_type":"code","source":"base_dir = 'data'\nos.makedirs(base_dir, exist_ok=True)\n\nfor label in master_df['word'].unique():\n    os.makedirs(os.path.join(base_dir, label), exist_ok=True)","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:15:08.764158Z","iopub.execute_input":"2024-08-03T03:15:08.766178Z","iopub.status.idle":"2024-08-03T03:15:08.906765Z","shell.execute_reply.started":"2024-08-03T03:15:08.766102Z","shell.execute_reply":"2024-08-03T03:15:08.90491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.8) Saving Images in Folders","metadata":{}},{"cell_type":"code","source":"def save_img(doodle_data, image_path):\n    fig, ax = plt.subplots(figsize=(2.55, 2.55))\n\n    for stroke in doodle_data:\n        x, y = stroke  \n        plt.plot(x, y)\n\n    plt.gca().invert_yaxis()\n    ax.axis('off')\n\n    fig.canvas.draw()\n    image = np.frombuffer(fig.canvas.tostring_rgb(), dtype=np.uint8)\n    image = image.reshape(fig.canvas.get_width_height()[::-1] + (3,))\n    plt.close(fig)\n\n    image_gray = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n    cv2.imwrite(image_path, image_gray)\n\n\ndef apply_parallel(df, func, n_jobs=-1):\n    Parallel(n_jobs=n_jobs)(delayed(func)(doodle, path) for doodle, path in tqdm(zip(df['drawing'], df['image_path']), total=len(df)))\n\n\nmaster_df['image_path'] = master_df.apply(lambda row: os.path.join(base_dir, row['word'], f\"{row['key_id']}.png\"), axis=1)\n\napply_parallel(master_df, save_img)\n\nmaster_df.to_csv('master_doodle_dataframe.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-08-03T03:15:30.100306Z","iopub.execute_input":"2024-08-03T03:15:30.100773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 1.9) Zipping Images","metadata":{}},{"cell_type":"code","source":"shutil.make_archive('doodle', 'zip', base_dir)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\n\nThis notebook provides a comprehensive pipeline for preparing a doodle dataset, making it ready for machine learning tasks. By leveraging parallel processing and efficient data handling techniques, it ensures that the dataset is clean, balanced, and easy to use.","metadata":{}}]}