{
  "id": 413451,
  "title": "Better data representation",
  "url": "/competitions/asl-fingerspelling/discussion/413451",
  "author_name": "Vadim Irtlach",
  "post_date": "2023-05-28T17:57:31.573000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Data is presented as Pandas DataFrames with an average size of 1.4GB disk memory, which contain 1k (exception is the last with only 400) samples. So, training models with such data format is very uncomfortable, because you are required to read 1.4GB files. I decided to convert each sample into numpy files (TF users can use TF Records) with <code>(num_frames, num_landmarks, num_channels)</code> shape and load them as it is done in CV competitions, i.e. read samples one by one. However, I suppose there is a more convenient way to do this because I am sure I will struggle with reading data during Inference because it is also presented as DataFrame. So, if you know a better way for doing this, please share.</p>\n<p>P.S. I was looking for previous competition's notebooks and as I can see, participants also did convertation into TF records or numpy. </p>\n<p>Thanks! </p>",
  "messages": [
    {
      "id": 2278467,
      "postDate": "2023-05-28T17:57:31.573Z",
      "content": "<p>Data is presented as Pandas DataFrames with an average size of 1.4GB disk memory, which contain 1k (exception is the last with only 400) samples. So, training models with such data format is very uncomfortable, because you are required to read 1.4GB files. I decided to convert each sample into numpy files (TF users can use TF Records) with <code>(num_frames, num_landmarks, num_channels)</code> shape and load them as it is done in CV competitions, i.e. read samples one by one. However, I suppose there is a more convenient way to do this because I am sure I will struggle with reading data during Inference because it is also presented as DataFrame. So, if you know a better way for doing this, please share.</p>\n<p>P.S. I was looking for previous competition's notebooks and as I can see, participants also did convertation into TF records or numpy. </p>\n<p>Thanks! </p>",
      "rawMarkdown": "Data is presented as Pandas DataFrames with an average size of 1.4GB disk memory, which contain 1k (exception is the last with only 400) samples. So, training models with such data format is very uncomfortable, because you are required to read 1.4GB files. I decided to convert each sample into numpy files (TF users can use TF Records) with `(num_frames, num_landmarks, num_channels)` shape and load them as it is done in CV competitions, i.e. read samples one by one. However, I suppose there is a more convenient way to do this because I am sure I will struggle with reading data during Inference because it is also presented as DataFrame. So, if you know a better way for doing this, please share.\n\nP.S. I was looking for previous competition's notebooks and as I can see, participants also did convertation into TF records or numpy. \n\nThanks! ",
      "votes": 4
    },
    {
      "id": 2279198,
      "postDate": "2023-05-29T07:14:01.313Z",
      "content": "<p>I can share my own experience from the previous competition. I wished to have fast data loading and the opportunity to experiment with features, so I needed a flexible dataset.  </p>\n<p>As a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native <code>load_relevant_data_subset</code> consumes 35min on average, while <strong>TFRecords takes only ~35sec</strong> (sic!).</p>\n<p>Moreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.</p>\n<p>What I achieved with this:</p>\n<ul>\n<li>Having a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch</li>\n<li>Having a single .tfrecord per file I could create any CV split I wished within seconds. </li>\n</ul>\n<p>What about limitations:</p>\n<ul>\n<li>Applying random augmentation was a little bit tricky due to the cached mechansim of the tensorflow datasets, but still was doable.</li>\n</ul>\n<p>I can share a notebook with dataset creation and usages If you wish. Hope, that would help</p>",
      "rawMarkdown": "I can share my own experience from the previous competition. I wished to have fast data loading and the opportunity to experiment with features, so I needed a flexible dataset.  \n\nAs a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native `load_relevant_data_subset` consumes 35min on average, while **TFRecords takes only ~35sec** (sic!).\n\nMoreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.\n\nWhat I achieved with this:\n* Having a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch\n* Having a single .tfrecord per file I could create any CV split I wished within seconds. \n\nWhat about limitations:\n* Applying random augmentation was a little bit tricky due to the cached mechansim of the tensorflow datasets, but still was doable.\n\nI can share a notebook with dataset creation and usages If you wish. Hope, that would help",
      "votes": 1,
      "replies": [
        {
          "id": 2279279,
          "postDate": "2023-05-29T08:29:58.207Z",
          "content": "<pre><code>As a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native load_relevant_data_subset consumes 35min on average, while TFRecords takes only ~35sec (sic!).\n</code></pre>\n<p>I've observed no such speed-up with numpy reading (<code>np.load</code>). </p>\n<pre><code>Moreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.\n\nWhat I achieved with this:\n\nHaving a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch\nHaving a single .tfrecord per file I could create any CV split I wished within seconds.\n</code></pre>\n<p>Yep, exactly. This is why I did it in the same way. </p>\n<pre><code>I can share a notebook with dataset creation and usages If you wish. Hope, that would help\n</code></pre>\n<p>I wouldn't mind! Thank you!</p>",
          "rawMarkdown": "```\nAs a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native load_relevant_data_subset consumes 35min on average, while TFRecords takes only ~35sec (sic!).\n```\n\nI've observed no such speed-up with numpy reading (`np.load`). \n\n```\nMoreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.\n\nWhat I achieved with this:\n\nHaving a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch\nHaving a single .tfrecord per file I could create any CV split I wished within seconds.\n```\n\nYep, exactly. This is why I did it in the same way. \n\n```\nI can share a notebook with dataset creation and usages If you wish. Hope, that would help\n```\n\nI wouldn't mind! Thank you!",
          "replies": [
            {
              "id": 2279292,
              "postDate": "2023-05-29T08:44:54.737Z",
              "content": "<p>For reading 10k samples one by one with <code>np.load</code> takes ~289 seconds. So, for reading all the training samples I need around 28 minutes. Maybe, TF records are really fast? </p>",
              "rawMarkdown": "For reading 10k samples one by one with `np.load` takes ~289 seconds. So, for reading all the training samples I need around 28 minutes. Maybe, TF records are really fast? "
            },
            {
              "id": 2279594,
              "postDate": "2023-05-29T13:19:20.577Z",
              "content": "<blockquote>\n  <p>I wouldn't mind! Thank you!</p>\n</blockquote>\n<p>Here I share a link to the notebook I used in the previous competition. I guess a little modification is required to adapt it to the current data structure - <a href=\"https://www.kaggle.com/code/meowmeowmeowmeowmeow/asl-sratified-full-dataset-in-tfrecords-format/notebook\" target=\"_blank\">[ASL] Sratified Full dataset in TFRecords format</a></p>",
              "rawMarkdown": "> I wouldn't mind! Thank you!\n\nHere I share a link to the notebook I used in the previous competition. I guess a little modification is required to adapt it to the current data structure - [[ASL] Sratified Full dataset in TFRecords format](https://www.kaggle.com/code/meowmeowmeowmeowmeow/asl-sratified-full-dataset-in-tfrecords-format/notebook)",
              "votes": 1
            },
            {
              "id": 2279603,
              "postDate": "2023-05-29T13:23:31.940Z",
              "content": "<p>I was very surprised as well. I guess, reading data with TFRecords is very efficient through the dataset. One of the reasons may be a prefetch feature.<br>\nOnce I switched to the separate files per each sample, I was surprised, that the data loading speed was increased. I was worried, that opening/closing too many files would slow down the pipeline, but it did the opposite. </p>\n<p>What about np.load, I believe tf.data is much more efficient rather than np.load. (Disclaimer, no studies have been performed on that, just my personal inuition :) )</p>",
              "rawMarkdown": "I was very surprised as well. I guess, reading data with TFRecords is very efficient through the dataset. One of the reasons may be a prefetch feature.\nOnce I switched to the separate files per each sample, I was surprised, that the data loading speed was increased. I was worried, that opening/closing too many files would slow down the pipeline, but it did the opposite. \n\nWhat about np.load, I believe tf.data is much more efficient rather than np.load. (Disclaimer, no studies have been performed on that, just my personal inuition :) )",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2278700,
      "postDate": "2023-05-28T22:41:34.960Z",
      "content": "<p>At first you should drop unused landmarks. You typically need only around 100 landmarks, at least for previous competition. Then convert them to numpy for pytorch or tfrecords for tensorflow.</p>",
      "rawMarkdown": "At first you should drop unused landmarks. You typically need only around 100 landmarks, at least for previous competition. Then convert them to numpy for pytorch or tfrecords for tensorflow.",
      "votes": 2,
      "replies": [
        {
          "id": 2279272,
          "postDate": "2023-05-29T08:25:37.513Z",
          "content": "<p>Yes, I know, <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>, but for now, I would like to experiment with different combinations of them, so I need to save them all. At least I have enough disk space for saving another 200GB 😄. </p>",
          "rawMarkdown": "Yes, I know, @bamps53, but for now, I would like to experiment with different combinations of them, so I need to save them all. At least I have enough disk space for saving another 200GB 😄. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2279198,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2023-05-29T07:14:01.313000",
      "content": "<p>I can share my own experience from the previous competition. I wished to have fast data loading and the opportunity to experiment with features, so I needed a flexible dataset.  </p>\n<p>As a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native <code>load_relevant_data_subset</code> consumes 35min on average, while <strong>TFRecords takes only ~35sec</strong> (sic!).</p>\n<p>Moreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.</p>\n<p>What I achieved with this:</p>\n<ul>\n<li>Having a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch</li>\n<li>Having a single .tfrecord per file I could create any CV split I wished within seconds. </li>\n</ul>\n<p>What about limitations:</p>\n<ul>\n<li>Applying random augmentation was a little bit tricky due to the cached mechansim of the tensorflow datasets, but still was doable.</li>\n</ul>\n<p>I can share a notebook with dataset creation and usages If you wish. Hope, that would help</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2279279,
          "author_name": "Vadim Irtlach",
          "author_url": "",
          "post_date": "2023-05-29T08:29:58.207000",
          "content": "<pre><code>As a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native load_relevant_data_subset consumes 35min on average, while TFRecords takes only ~35sec (sic!).\n</code></pre>\n<p>I've observed no such speed-up with numpy reading (<code>np.load</code>). </p>\n<pre><code>Moreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.\n\nWhat I achieved with this:\n\nHaving a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch\nHaving a single .tfrecord per file I could create any CV split I wished within seconds.\n</code></pre>\n<p>Yep, exactly. This is why I did it in the same way. </p>\n<pre><code>I can share a notebook with dataset creation and usages If you wish. Hope, that would help\n</code></pre>\n<p>I wouldn't mind! Thank you!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2279292,
              "author_name": "Vadim Irtlach",
              "author_url": "",
              "post_date": "2023-05-29T08:44:54.737000",
              "content": "<p>For reading 10k samples one by one with <code>np.load</code> takes ~289 seconds. So, for reading all the training samples I need around 28 minutes. Maybe, TF records are really fast? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2279594,
              "author_name": "Mykola",
              "author_url": "",
              "post_date": "2023-05-29T13:19:20.577000",
              "content": "<blockquote>\n  <p>I wouldn't mind! Thank you!</p>\n</blockquote>\n<p>Here I share a link to the notebook I used in the previous competition. I guess a little modification is required to adapt it to the current data structure - <a href=\"https://www.kaggle.com/code/meowmeowmeowmeowmeow/asl-sratified-full-dataset-in-tfrecords-format/notebook\" target=\"_blank\">[ASL] Sratified Full dataset in TFRecords format</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2279603,
              "author_name": "Mykola",
              "author_url": "",
              "post_date": "2023-05-29T13:23:31.940000",
              "content": "<p>I was very surprised as well. I guess, reading data with TFRecords is very efficient through the dataset. One of the reasons may be a prefetch feature.<br>\nOnce I switched to the separate files per each sample, I was surprised, that the data loading speed was increased. I was worried, that opening/closing too many files would slow down the pipeline, but it did the opposite. </p>\n<p>What about np.load, I believe tf.data is much more efficient rather than np.load. (Disclaimer, no studies have been performed on that, just my personal inuition :) )</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2278700,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2023-05-28T22:41:34.960000",
      "content": "<p>At first you should drop unused landmarks. You typically need only around 100 landmarks, at least for previous competition. Then convert them to numpy for pytorch or tfrecords for tensorflow.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2279272,
          "author_name": "Vadim Irtlach",
          "author_url": "",
          "post_date": "2023-05-29T08:25:37.513000",
          "content": "<p>Yes, I know, <a href=\"https://www.kaggle.com/bamps53\" target=\"_blank\">@bamps53</a>, but for now, I would like to experiment with different combinations of them, so I need to save them all. At least I have enough disk space for saving another 200GB 😄. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2278467": "Data is presented as Pandas DataFrames with an average size of 1.4GB disk memory, which contain 1k (exception is the last with only 400) samples. So, training models with such data format is very uncomfortable, because you are required to read 1.4GB files. I decided to convert each sample into numpy files (TF users can use TF Records) with `(num_frames, num_landmarks, num_channels)` shape and load them as it is done in CV competitions, i.e. read samples one by one. However, I suppose there is a more convenient way to do this because I am sure I will struggle with reading data during Inference because it is also presented as DataFrame. So, if you know a better way for doing this, please share.\n\nP.S. I was looking for previous competition's notebooks and as I can see, participants also did convertation into TF records or numpy. \n\nThanks! ",
    "2279198": "I can share my own experience from the previous competition. I wished to have fast data loading and the opportunity to experiment with features, so I needed a flexible dataset.  \n\nAs a result, I converted the dataset to the TFRecord format. It reduced the size of the dataset from ~40GB to ~20GB. This is due to the high redundancy of the parquet files. Moreover, it allows me to reduce iteration time over the dataset. Native `load_relevant_data_subset` consumes 35min on average, while **TFRecords takes only ~35sec** (sic!).\n\nMoreover, I converted the dataset in such a way, that each sample (gesture) is stored under a single .tfrecord file. Thus, a number of files corresponded to a number of samples.\n\nWhat I achieved with this:\n* Having a raw dataset gives me an opportunity to easily deal with feature engineering. I could drop any features I want and produce any features I like. TensorFlow dataset caching mechanism caches everything after 1 epoch, so all the other epochs didn't require processing from scratch\n* Having a single .tfrecord per file I could create any CV split I wished within seconds. \n\nWhat about limitations:\n* Applying random augmentation was a little bit tricky due to the cached mechansim of the tensorflow datasets, but still was doable.\n\nI can share a notebook with dataset creation and usages If you wish. Hope, that would help",
    "2278700": "At first you should drop unused landmarks. You typically need only around 100 landmarks, at least for previous competition. Then convert them to numpy for pytorch or tfrecords for tensorflow."
  }
}