{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\nData preprocessing is a crucial step in machine learning projects, including competitions. This notebook provides a preprocessing template specifically designed for the American Sign Language (ASL) Fingerspelling Recognition competition on Kaggle. The code in this notebook aims to facilitate the preprocessing of the ASL Fingerspelling dataset, enabling participants to focus on developing effective machine learning models.","metadata":{}},{"cell_type":"markdown","source":"# Notebook Structure\nThis notebook is organized as follows:\n\n1. Data and Environment Setup:\n    - Importing necessary libraries\n    - Defining file paths and constants\n    - Checking the execution environment (Kaggle or local)\n    \n2. Data Loading:\n    - Loading the raw data from CSV and Parquet files\n    - Creating the output directory for preprocessed data\n    \n3. Preprocessing Functions:\n    - Defining custom preprocessing functions\n    - Placeholder functions for feature engineering, data transformation, and usefulness checks\n    \n4. Preprocessing Pipeline:\n    - Preprocessing the data using the defined functions\n    - Writing preprocessed data to TFRecords\n    \n5. Execution Time Measurement:\n    - Calculating the execution time of the preprocessing pipeline\n    \n6. Conclusion:\n    - Recap of the purpose and benefits of using this preprocessing template\n    - Encouragement to customize the functions based on specific requirements","metadata":{}},{"cell_type":"markdown","source":"# Customization Instructions\nTo use this preprocessing template for your project, follow these steps:\n\n1. Customize the preprocessing functions:\n    - Modify the `preprocess_frames()`, `preprocess_phrase()`, and `is_useful()` functions to suit your preprocessing requirements.\n    - Implement feature engineering, data transformation, and usefulness checks based on your project's needs.\n    \n2. Update file paths and constants:\n    - Modify the `PATH_KAGGLE`, `PATH_LOCAL`, `INPUT_TRAIN_CSV`, `INPUT_LANDMARKS`, and `OUTPUT_DATASET` variables according to your file locations.\n    \n3. Execute the preprocessing pipeline:\n    - Run the `preprocess()` function to perform the data preprocessing.\n    - The preprocessed data will be written to the specified TFRecords output directory.\n    \n4. Evaluate the execution time:\n    - Observe the execution time of the preprocessing pipeline to assess its efficiency.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\nData preprocessing is a critical step in machine learning projects, and this notebook provides a standardized template to streamline the process. By customizing the preprocessing functions and following the provided instructions, you can efficiently preprocess your data and save valuable time. Enjoy using this preprocessing template for your projects! 😁","metadata":{}},{"cell_type":"markdown","source":"# 1. Data and Environment Setup\nThe code below is code that does not need to be modified. Edit only when necessary.","metadata":{}},{"cell_type":"code","source":"import os\nimport time\nimport pandas as pd\nimport tensorflow as tf\nfrom tqdm import tqdm\n\n\nPATH_KAGGLE = '/kaggle/input/asl-fingerspelling/'\nPATH_LOCAL = 'D:/dataset/kaggle/asl-fingerspelling/'\n\nINPUT_TRAIN_CSV = 'train.csv'\nINPUT_LANDMARKS = 'train_landmarks/'\nOUTPUT_DATASET = 'tfds'\n\n\ndef is_kaggle_env():\n    return os.environ.get('KAGGLE_KERNEL_RUN_TYPE') is not None\n\n\ndef get_path():\n    return PATH_KAGGLE if is_kaggle_env() else PATH_LOCAL\n\n\ndef create_output_path():\n    if not os.path.isdir(OUTPUT_DATASET):\n        os.mkdir(OUTPUT_DATASET)\n\n\ndef load_train_csv():\n    path = os.path.join(get_path(), INPUT_TRAIN_CSV)\n    return pd.read_csv(path)\n\n\ndef load_parquet(file_id, columns=None):\n    path = os.path.join(get_path(), f'{INPUT_LANDMARKS}{file_id}.parquet')\n\n    if columns is None:\n        return pd.read_parquet(path)\n\n    return pd.read_parquet(path, columns=columns)\n\n\ndef get_record_writer(file_id):\n    path = f'{OUTPUT_DATASET}/{file_id}.tfrecord'\n    return tf.io.TFRecordWriter(path)\n\n\ndef preprocess_file(file_id, file_df):\n    parquet = load_parquet(file_id, get_columns())\n    record_writer = get_record_writer(file_id)\n\n    for seq_id, phrase in zip(file_df['sequence_id'], file_df['phrase']):\n        frames = parquet.iloc[parquet.index == seq_id]\n        frames = preprocess_frames(frames)\n        phrase = preprocess_phrase(phrase)\n\n        if is_useful(frames, phrase):\n            features = to_features(frames, phrase)\n            record_writer.write(features)\n\n    record_writer.close()\n\n\ndef preprocess():\n    create_output_path()\n\n    train_df = load_train_csv()\n\n    for file_id in tqdm(train_df['file_id'].unique()):\n        file_df = train_df.loc[train_df['file_id'] == file_id]\n        preprocess_file(file_id, file_df)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Preprocessing Functions (cutom functions)\nCustomize these functions according to your requirements","metadata":{}},{"cell_type":"code","source":"def get_columns():\n    # Placeholder for get columns\n    POSE = [0, 1, 2, 3, 22]\n\n    columns = [f'x_left_hand_{i}' for i in range(21)]\n    columns += [f'x_right_hand_{i}' for i in range(21)]\n    columns += [f'x_pose_{i}' for i in POSE]\n\n    columns += [f'y_left_hand_{i}' for i in range(21)]\n    columns += [f'y_right_hand_{i}' for i in range(21)]\n    columns += [f'y_pose_{i}' for i in POSE]\n\n    columns += [f'z_left_hand_{i}' for i in range(21)]\n    columns += [f'z_right_hand_{i}' for i in range(21)]\n    columns += [f'z_pose_{i}' for i in POSE]\n\n    return columns\n\n\ndef to_features(frames, phrase):\n    features = {}\n\n    for col in frames.columns:\n        value = frames[col]\n        float_list = tf.train.FloatList(value=value)\n        features[col] = tf.train.Feature(float_list=float_list)\n\n    bytes_list = tf.train.BytesList(value=[phrase])\n    features['phrase'] = tf.train.Feature(bytes_list=bytes_list)\n    features = tf.train.Example(features=tf.train.Features(feature=features))\n    return features.SerializeToString()\n\n\ndef preprocess_frames(frames):\n    # Placeholder for preprocessing frames\n    return frames\n\n\ndef preprocess_phrase(phrase):\n    # Placeholder for preprocessing phrase\n    return phrase.encode('utf-8')\n\n\ndef is_useful(frames, phrase):\n    # Placeholder for usefulness check\n    return True","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. main function","metadata":{}},{"cell_type":"code","source":"start_time = time.time()\npreprocess()\nend_time = time.time()\nexecution_time = end_time - start_time\nprint(f\"Execution time of preprocess: {execution_time} seconds\")","metadata":{},"execution_count":null,"outputs":[]}]}