{
  "id": 414549,
  "title": "Load numpy arrays 1.7x faster",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/414549",
  "author_name": "JEANMPIA",
  "post_date": "2023-06-02T06:53:55.433000",
  "votes": 42,
  "comment_count": 14,
  "views": 0,
  "content": "<p><strong>This load function comes from <a href=\"https://github.com/divideconcept/fastnumpyio\" target=\"_blank\">this</a> git repo:</strong></p>\n<pre><code> :\n     ():\n        file=(file,)\n        header = file.read()\n        descr = (header[:], ).replace(,).replace(,)\n        shape = ((num)  num  (header[:], ).replace(, ).replace(, ).replace(, ).split())\n        datasize = np.lib..descr_to_dtype(descr).itemsize\n         dimension  shape:\n            datasize *= dimension\n         np.ndarray(shape, dtype=descr, buffer=file.read(datasize))\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F95ec86cd2452068bde95ff4b09235f75%2FScreenshot_96.jpg?generation=1685688703795772&amp;alt=media\" alt=\"\"></p>\n<p><em>my epoch time went from 4mins to 2:30mins with this simple change</em></p>",
  "messages": [
    {
      "id": 2284589,
      "postDate": "2023-06-02T06:53:55.433Z",
      "content": "<p><strong>This load function comes from <a href=\"https://github.com/divideconcept/fastnumpyio\" target=\"_blank\">this</a> git repo:</strong></p>\n<pre><code> :\n     ():\n        file=(file,)\n        header = file.read()\n        descr = (header[:], ).replace(,).replace(,)\n        shape = ((num)  num  (header[:], ).replace(, ).replace(, ).replace(, ).split())\n        datasize = np.lib..descr_to_dtype(descr).itemsize\n         dimension  shape:\n            datasize *= dimension\n         np.ndarray(shape, dtype=descr, buffer=file.read(datasize))\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F95ec86cd2452068bde95ff4b09235f75%2FScreenshot_96.jpg?generation=1685688703795772&amp;alt=media\" alt=\"\"></p>\n<p><em>my epoch time went from 4mins to 2:30mins with this simple change</em></p>",
      "rawMarkdown": "**This load function comes from [this](https://github.com/divideconcept/fastnumpyio) git repo:**\n\n```python\nclass fastnumpyio:\n    def load(file):\n        file=open(file,\"rb\")\n        header = file.read(128)\n        descr = str(header[19:25], 'utf-8').replace(\"'\",\"\").replace(\" \",\"\")\n        shape = tuple(int(num) for num in str(header[60:120], 'utf-8').replace(', }', '').replace('(', '').replace(')', '').split(','))\n        datasize = np.lib.format.descr_to_dtype(descr).itemsize\n        for dimension in shape:\n            datasize *= dimension\n        return np.ndarray(shape, dtype=descr, buffer=file.read(datasize))\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F95ec86cd2452068bde95ff4b09235f75%2FScreenshot_96.jpg?generation=1685688703795772&alt=media)\n\n*my epoch time went from 4mins to 2:30mins with this simple change*",
      "votes": 41
    },
    {
      "id": 2284776,
      "postDate": "2023-06-02T09:13:31.890Z",
      "content": "<p>Wow, this appears to be highly useful! Thanks for the post <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a>!</p>",
      "rawMarkdown": "Wow, this appears to be highly useful! Thanks for the post @janmpia!",
      "votes": 3
    },
    {
      "id": 2284592,
      "postDate": "2023-06-02T06:58:58.500Z",
      "content": "<p><strong>If you are willing to sacrifice the adaptivness of your code, you can also do the following for a 2x:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F4d2739538ebb540155d2c405fe0798bb%2FScreenshot_97.jpg?generation=1685689109654626&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**If you are willing to sacrifice the adaptivness of your code, you can also do the following for a 2x:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F4d2739538ebb540155d2c405fe0798bb%2FScreenshot_97.jpg?generation=1685689109654626&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 2285139,
          "postDate": "2023-06-02T13:28:19.710Z",
          "content": "<p>What do you mean with \"adaptiveness\"?</p>",
          "rawMarkdown": "What do you mean with \"adaptiveness\"?\n",
          "votes": 1,
          "replies": [
            {
              "id": 2285289,
              "postDate": "2023-06-02T15:31:03.910Z",
              "content": "<p>hey <a href=\"https://www.kaggle.com/janhuebik\" target=\"_blank\">@janhuebik</a>,<br>\nfor this to work, you have a write the dtype and shape of thr array you are loading, meaning that if you want to upscale the images or change to dtype to float16 for example, you will have to change that too, hence the adaptiveness.<br>\nOn top of that, you will now be working with menmap and not numpy array, which will need be taken care of.<br>\nHope that makes my point clearer.</p>",
              "rawMarkdown": "hey @janhuebik,\nfor this to work, you have a write the dtype and shape of thr array you are loading, meaning that if you want to upscale the images or change to dtype to float16 for example, you will have to change that too, hence the adaptiveness.\nOn top of that, you will now be working with menmap and not numpy array, which will need be taken care of.\nHope that makes my point clearer.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2291790,
      "postDate": "2023-06-07T20:17:09.027Z",
      "content": "<p>this is helpful i think</p>",
      "rawMarkdown": "this is helpful i think\n",
      "votes": 1
    },
    {
      "id": 2286653,
      "postDate": "2023-06-03T16:06:39.527Z",
      "content": "<p>very useful content <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> </p>",
      "rawMarkdown": "very useful content @janmpia ",
      "votes": 1
    },
    {
      "id": 2294664,
      "postDate": "2023-06-10T08:00:38.217Z",
      "content": "<p>I have made kernel to compare all numpy methods (reading fp16 data) vs JPEG. <a href=\"https://www.kaggle.com/code/melgor/benchmark-numpy-vs-jpeg\" target=\"_blank\">https://www.kaggle.com/code/melgor/benchmark-numpy-vs-jpeg</a></p>\n<p>Options with scores:</p>\n<ul>\n<li>raw numpy: 1:48</li>\n<li>fastnumpy: 0:50</li>\n<li>numpy-memmap: 0:49</li>\n<li>jpeg: 1:13</li>\n</ul>\n<p>Timing vary based on run, but fastnumpy/memmap is enough for this competition.</p>",
      "rawMarkdown": "I have made kernel to compare all numpy methods (reading fp16 data) vs JPEG. https://www.kaggle.com/code/melgor/benchmark-numpy-vs-jpeg\n\nOptions with scores:\n- raw numpy: 1:48\n- fastnumpy: 0:50\n- numpy-memmap: 0:49\n- jpeg: 1:13\n\nTiming vary based on run, but fastnumpy/memmap is enough for this competition.",
      "votes": 2,
      "replies": [
        {
          "id": 2295114,
          "postDate": "2023-06-10T15:40:31.357Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/melgor\" target=\"_blank\">@melgor</a>, <br>\nI'm kind of confused by your conclusion after seeing your notebook:</p>\n<blockquote>\n  <p>Timing vary based on run, <strong>but numpy is enough for this competition.</strong></p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F792949947e52943c1478d181384d23cf%2FScreenshot_104.jpg?generation=1686410876904695&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F2266719ca8ebc1c6ead933b1eddbe7ad%2FScreenshot_105.jpg?generation=1686410886826231&amp;alt=media\" alt=\"\"></p>\n<p>from those screen shots of your notebook, fastnumpy looks much faster.</p>",
          "rawMarkdown": "Hey @melgor, \nI'm kind of confused by your conclusion after seeing your notebook:\n>Timing vary based on run, **but numpy is enough for this competition.**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F792949947e52943c1478d181384d23cf%2FScreenshot_104.jpg?generation=1686410876904695&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F2266719ca8ebc1c6ead933b1eddbe7ad%2FScreenshot_105.jpg?generation=1686410886826231&alt=media)\n\nfrom those screen shots of your notebook, fastnumpy looks much faster.",
          "votes": 2,
          "replies": [
            {
              "id": 2295356,
              "postDate": "2023-06-10T20:35:21.087Z",
              "content": "<p>I should be more specific as I have in my mind Numpy vs JPEG. </p>\n<p>But just comparing just Numpy methods, fastnumpy or memmap should be used.</p>",
              "rawMarkdown": "I should be more specific as I have in my mind Numpy vs JPEG. \n\nBut just comparing just Numpy methods, fastnumpy or memmap should be used.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2317537,
      "postDate": "2023-06-25T19:04:24.243Z",
      "content": "<p>Git repo link is broken.</p>",
      "rawMarkdown": "Git repo link is broken.",
      "replies": [
        {
          "id": 2343473,
          "postDate": "2023-07-13T17:27:02.100Z",
          "content": "<p>oh thank you for telling me. I updated the topic</p>",
          "rawMarkdown": "oh thank you for telling me. I updated the topic"
        }
      ]
    },
    {
      "id": 2286960,
      "postDate": "2023-06-04T02:47:37.550Z",
      "content": "<p>(From translator) Does the memory expense increase?</p>",
      "rawMarkdown": "(From translator) Does the memory expense increase?",
      "isDeleted": true,
      "replies": [
        {
          "id": 2287249,
          "postDate": "2023-06-04T08:56:17.227Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/mozattt\" target=\"_blank\">@mozattt</a>,<br>\nI didn't test that to be honest but my code only sped up without any other side effect.<br>\nNote that you get a speed up only with numpy arrays that don't have a complex structure, but for this competition there are none !</p>",
          "rawMarkdown": "Hey @mozattt,\nI didn't test that to be honest but my code only sped up without any other side effect.\nNote that you get a speed up only with numpy arrays that don't have a complex structure, but for this competition there are none !",
          "replies": [
            {
              "id": 2287283,
              "postDate": "2023-06-04T09:41:27.930Z",
              "content": "<p>Thanks. :D</p>",
              "rawMarkdown": "Thanks. :D",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2284776,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2023-06-02T09:13:31.890000",
      "content": "<p>Wow, this appears to be highly useful! Thanks for the post <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a>!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2284592,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-06-02T06:58:58.500000",
      "content": "<p><strong>If you are willing to sacrifice the adaptivness of your code, you can also do the following for a 2x:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F4d2739538ebb540155d2c405fe0798bb%2FScreenshot_97.jpg?generation=1685689109654626&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2285139,
          "author_name": "Jan H",
          "author_url": "",
          "post_date": "2023-06-02T13:28:19.710000",
          "content": "<p>What do you mean with \"adaptiveness\"?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2285289,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-06-02T15:31:03.910000",
              "content": "<p>hey <a href=\"https://www.kaggle.com/janhuebik\" target=\"_blank\">@janhuebik</a>,<br>\nfor this to work, you have a write the dtype and shape of thr array you are loading, meaning that if you want to upscale the images or change to dtype to float16 for example, you will have to change that too, hence the adaptiveness.<br>\nOn top of that, you will now be working with menmap and not numpy array, which will need be taken care of.<br>\nHope that makes my point clearer.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2291790,
      "author_name": "Priyanshu shukla",
      "author_url": "",
      "post_date": "2023-06-07T20:17:09.027000",
      "content": "<p>this is helpful i think</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2286653,
      "author_name": "Tamanna Akter Swarna",
      "author_url": "",
      "post_date": "2023-06-03T16:06:39.527000",
      "content": "<p>very useful content <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2294664,
      "author_name": "Bartek",
      "author_url": "",
      "post_date": "2023-06-10T08:00:38.217000",
      "content": "<p>I have made kernel to compare all numpy methods (reading fp16 data) vs JPEG. <a href=\"https://www.kaggle.com/code/melgor/benchmark-numpy-vs-jpeg\" target=\"_blank\">https://www.kaggle.com/code/melgor/benchmark-numpy-vs-jpeg</a></p>\n<p>Options with scores:</p>\n<ul>\n<li>raw numpy: 1:48</li>\n<li>fastnumpy: 0:50</li>\n<li>numpy-memmap: 0:49</li>\n<li>jpeg: 1:13</li>\n</ul>\n<p>Timing vary based on run, but fastnumpy/memmap is enough for this competition.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2295114,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-06-10T15:40:31.357000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/melgor\" target=\"_blank\">@melgor</a>, <br>\nI'm kind of confused by your conclusion after seeing your notebook:</p>\n<blockquote>\n  <p>Timing vary based on run, <strong>but numpy is enough for this competition.</strong></p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F792949947e52943c1478d181384d23cf%2FScreenshot_104.jpg?generation=1686410876904695&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F2266719ca8ebc1c6ead933b1eddbe7ad%2FScreenshot_105.jpg?generation=1686410886826231&amp;alt=media\" alt=\"\"></p>\n<p>from those screen shots of your notebook, fastnumpy looks much faster.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2295356,
              "author_name": "Bartek",
              "author_url": "",
              "post_date": "2023-06-10T20:35:21.087000",
              "content": "<p>I should be more specific as I have in my mind Numpy vs JPEG. </p>\n<p>But just comparing just Numpy methods, fastnumpy or memmap should be used.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2317537,
      "author_name": "Pankaj Joshi",
      "author_url": "",
      "post_date": "2023-06-25T19:04:24.243000",
      "content": "<p>Git repo link is broken.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2343473,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-07-13T17:27:02.100000",
          "content": "<p>oh thank you for telling me. I updated the topic</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2286960,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-04T02:47:37.550000",
      "content": "<p>(From translator) Does the memory expense increase?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2287249,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-06-04T08:56:17.227000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/mozattt\" target=\"_blank\">@mozattt</a>,<br>\nI didn't test that to be honest but my code only sped up without any other side effect.<br>\nNote that you get a speed up only with numpy arrays that don't have a complex structure, but for this competition there are none !</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2287283,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-04T09:41:27.930000",
              "content": "<p>Thanks. :D</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2284589": "**This load function comes from [this](https://github.com/divideconcept/fastnumpyio) git repo:**\n\n```python\nclass fastnumpyio:\n    def load(file):\n        file=open(file,\"rb\")\n        header = file.read(128)\n        descr = str(header[19:25], 'utf-8').replace(\"'\",\"\").replace(\" \",\"\")\n        shape = tuple(int(num) for num in str(header[60:120], 'utf-8').replace(', }', '').replace('(', '').replace(')', '').split(','))\n        datasize = np.lib.format.descr_to_dtype(descr).itemsize\n        for dimension in shape:\n            datasize *= dimension\n        return np.ndarray(shape, dtype=descr, buffer=file.read(datasize))\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F95ec86cd2452068bde95ff4b09235f75%2FScreenshot_96.jpg?generation=1685688703795772&alt=media)\n\n*my epoch time went from 4mins to 2:30mins with this simple change*",
    "2284776": "Wow, this appears to be highly useful! Thanks for the post @janmpia!",
    "2284592": "**If you are willing to sacrifice the adaptivness of your code, you can also do the following for a 2x:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12870466%2F4d2739538ebb540155d2c405fe0798bb%2FScreenshot_97.jpg?generation=1685689109654626&alt=media)",
    "2291790": "this is helpful i think\n",
    "2286653": "very useful content @janmpia ",
    "2294664": "I have made kernel to compare all numpy methods (reading fp16 data) vs JPEG. https://www.kaggle.com/code/melgor/benchmark-numpy-vs-jpeg\n\nOptions with scores:\n- raw numpy: 1:48\n- fastnumpy: 0:50\n- numpy-memmap: 0:49\n- jpeg: 1:13\n\nTiming vary based on run, but fastnumpy/memmap is enough for this competition.",
    "2317537": "Git repo link is broken.",
    "2286960": "(From translator) Does the memory expense increase?"
  }
}