{
  "id": 495128,
  "title": "Useful code and comparison result of save speeds between Polars and Pandas",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495128",
  "author_name": "chumajin",
  "post_date": "2024-04-19T18:51:40.474000",
  "votes": 20,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I will share some code that seems to be useful for this competition.<br>\nThis competition involves large data sets, and just opening or saving them can be time-consuming. Therefore, using polars can help reduce time and make the process more comfortable.</p>\n<h2>Targets extraction</h2>\n<pre><code> polars  pl\nsample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=)\n#    \ntargets = sample.[:]\n</code></pre>\n<h2>Save submission.csv</h2>\n<pre><code>## change pandas  polars  you use pandas\ntest_polars = pl.from\ntest_polars.write\n</code></pre>\n<h2>Result of time comparison between polars and pandas</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fdcdfef4534835d2f46668af6c290c2f5%2FClipboard03.jpg?generation=1713552587674532&amp;alt=media\"></p>\n<h2><strong>About 18 times faster using by polars.</strong></h2>",
  "messages": [
    {
      "id": 2761282,
      "postDate": "2024-04-19T18:51:40.473Z",
      "content": "<p>I will share some code that seems to be useful for this competition.<br>\nThis competition involves large data sets, and just opening or saving them can be time-consuming. Therefore, using polars can help reduce time and make the process more comfortable.</p>\n<h2>Targets extraction</h2>\n<pre><code> polars  pl\nsample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=)\n#    \ntargets = sample.[:]\n</code></pre>\n<h2>Save submission.csv</h2>\n<pre><code>## change pandas  polars  you use pandas\ntest_polars = pl.from\ntest_polars.write\n</code></pre>\n<h2>Result of time comparison between polars and pandas</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fdcdfef4534835d2f46668af6c290c2f5%2FClipboard03.jpg?generation=1713552587674532&amp;alt=media\"></p>\n<h2><strong>About 18 times faster using by polars.</strong></h2>",
      "rawMarkdown": "I will share some code that seems to be useful for this competition.\nThis competition involves large data sets, and just opening or saving them can be time-consuming. Therefore, using polars can help reduce time and make the process more comfortable.\n\n##  Targets extraction\n\n~~~\nimport polars as pl\nsample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=1)\n# read only 1 rows\ntargets = sample.columns[1:]\n~~~\n\n## Save submission.csv\n\n~~~\n## change pandas to polars if you use pandas\ntest_polars = pl.from_pandas(test[[\"sample_id\"]+targets])\ntest_polars.write_csv(\"submission.csv\")\n~~~\n\n## Result of time comparison between polars and pandas\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fdcdfef4534835d2f46668af6c290c2f5%2FClipboard03.jpg?generation=1713552587674532&alt=media)\n\n## **About 18 times faster using by polars.**",
      "votes": 20
    },
    {
      "id": 2762755,
      "postDate": "2024-04-20T04:36:21.537Z",
      "content": "<p>You can also convert .csv to .parquet with low RAM<br>\nJust use lazy execution of polars, - </p>\n<pre><code>import polars  pl\n\nstorage_dir = \n file  os.listdir(storage_dir):\n    df = pl.scan\n    df.sink)\n</code></pre>",
      "rawMarkdown": "You can also convert .csv to .parquet with low RAM\nJust use lazy execution of polars, - \n```\nimport polars as pl\n\nstorage_dir = Path('/path/to/dataset/')\nfor file in os.listdir(storage_dir):\n    df = pl.scan_csv(storage_dir / file)\n    df.sink_parquet(storage_dir / file.replace('.csv', '.parquet'))\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2768638,
          "postDate": "2024-04-23T00:53:41.853Z",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> Thank you. Indeed, converting to Parquet when saving small files seems useful. Thank you for including lazy processing as well. It will be helpful.</p>",
          "rawMarkdown": "@martynoveduard Thank you. Indeed, converting to Parquet when saving small files seems useful. Thank you for including lazy processing as well. It will be helpful.",
          "replies": [
            {
              "id": 2768917,
              "postDate": "2024-04-23T04:53:36.313Z",
              "content": "<p>In the end I went with this code:</p>\n<pre><code> tqdm\n os\n\nos.makedirs(storage_dir / )\nframe_size = _000 # total length  _091_520\n chunk_index  tqdm.tqdm(range(, (num_train_samples // frame_size) + )):\n     = frame_size * chunk_index\n    df_sub = df.(, frame_size)\n    df_sub.sink_parquet(storage_dir /  / f)\n</code></pre>\n<p>After, during dataset creation, you can also lazily load like this:</p>\n<pre><code>  [pl.scan_parquet(data_dir /  / f) for i in indices]\n  pl.concat(dfs)       \n</code></pre>\n<p>Works fast!</p>",
              "rawMarkdown": "In the end I went with this code:\n```\nimport tqdm\nimport os\n\nos.makedirs(storage_dir / 'train_chunks')\nframe_size = 200_000 # total length is 10_091_520\nfor chunk_index in tqdm.tqdm(range(0, (num_train_samples // frame_size) + 1)):\n    offset = frame_size * chunk_index\n    df_sub = df.slice(offset, frame_size)\n    df_sub.sink_parquet(storage_dir / 'train_chunks' / f'train_{chunk_index}.parquet')\n```\n\nAfter, during dataset creation, you can also lazily load like this:\n```\ndfs = [pl.scan_parquet(data_dir / \"train_chunks\" / f\"train_{i}.parquet\") for i in indices]\ndf = pl.concat(dfs)       \n```\nWorks fast!",
              "votes": 10
            },
            {
              "id": 2769189,
              "postDate": "2024-04-23T07:52:32.113Z",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> Feels good! Thank you!!</p>",
              "rawMarkdown": "@martynoveduard Feels good! Thank you!!"
            }
          ]
        }
      ]
    },
    {
      "id": 2761330,
      "postDate": "2024-04-19T19:20:16.887Z",
      "content": "<p>Try DuckDB also, I think it is also a good option for such large datasets <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "rawMarkdown": "Try DuckDB also, I think it is also a good option for such large datasets @chumajin ",
      "votes": 1,
      "replies": [
        {
          "id": 2768633,
          "postDate": "2024-04-23T00:52:14.440Z",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Thank you for the addition. Yes, I have seen DuckDB performing well in other competitions as well. It will be helpful in this competition too.</p>",
          "rawMarkdown": "@ravi20076 Thank you for the addition. Yes, I have seen DuckDB performing well in other competitions as well. It will be helpful in this competition too."
        }
      ]
    },
    {
      "id": 2761352,
      "postDate": "2024-04-19T19:32:55.503Z",
      "content": "<p>Hi chumajin,<br>\nOne of the things that contributes to the speed of the Polars library is that it is in operation <code>Eager and Lazy Execution options</code> and It is derived from the <code>Rust programming language</code>.<br>\n<code>Polars   that follow the Lazy Execution system</code> Unlike<br>\n<code>Pandas that follow the Eager Execution system</code><br>\n<a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "rawMarkdown": "Hi chumajin,\nOne of the things that contributes to the speed of the Polars library is that it is in operation `Eager and Lazy Execution options` and It is derived from the `Rust programming language`.\n`Polars   that follow the Lazy Execution system` Unlike\n`Pandas that follow the Eager Execution system`\n@chumajin ",
      "votes": 2,
      "replies": [
        {
          "id": 2761706,
          "postDate": "2024-04-20T03:14:51.583Z",
          "content": "<p><a href=\"https://www.kaggle.com/adnanalaref\" target=\"_blank\">@adnanalaref</a> Hi.  You are right. Thank you for comment !</p>",
          "rawMarkdown": "@adnanalaref Hi.  You are right. Thank you for comment !"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2762755,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-20T04:36:21.537000",
      "content": "<p>You can also convert .csv to .parquet with low RAM<br>\nJust use lazy execution of polars, - </p>\n<pre><code>import polars  pl\n\nstorage_dir = \n file  os.listdir(storage_dir):\n    df = pl.scan\n    df.sink)\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2768638,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-04-23T00:53:41.853000",
          "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> Thank you. Indeed, converting to Parquet when saving small files seems useful. Thank you for including lazy processing as well. It will be helpful.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2768917,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2024-04-23T04:53:36.313000",
              "content": "<p>In the end I went with this code:</p>\n<pre><code> tqdm\n os\n\nos.makedirs(storage_dir / )\nframe_size = _000 # total length  _091_520\n chunk_index  tqdm.tqdm(range(, (num_train_samples // frame_size) + )):\n     = frame_size * chunk_index\n    df_sub = df.(, frame_size)\n    df_sub.sink_parquet(storage_dir /  / f)\n</code></pre>\n<p>After, during dataset creation, you can also lazily load like this:</p>\n<pre><code>  [pl.scan_parquet(data_dir /  / f) for i in indices]\n  pl.concat(dfs)       \n</code></pre>\n<p>Works fast!</p>",
              "votes": 10,
              "replies": []
            },
            {
              "id": 2769189,
              "author_name": "chumajin",
              "author_url": "",
              "post_date": "2024-04-23T07:52:32.113000",
              "content": "<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> Feels good! Thank you!!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2761330,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-04-19T19:20:16.887000",
      "content": "<p>Try DuckDB also, I think it is also a good option for such large datasets <a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2768633,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-04-23T00:52:14.440000",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Thank you for the addition. Yes, I have seen DuckDB performing well in other competitions as well. It will be helpful in this competition too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2761352,
      "author_name": "Adnan Alaref",
      "author_url": "",
      "post_date": "2024-04-19T19:32:55.503000",
      "content": "<p>Hi chumajin,<br>\nOne of the things that contributes to the speed of the Polars library is that it is in operation <code>Eager and Lazy Execution options</code> and It is derived from the <code>Rust programming language</code>.<br>\n<code>Polars   that follow the Lazy Execution system</code> Unlike<br>\n<code>Pandas that follow the Eager Execution system</code><br>\n<a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2761706,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-04-20T03:14:51.583000",
          "content": "<p><a href=\"https://www.kaggle.com/adnanalaref\" target=\"_blank\">@adnanalaref</a> Hi.  You are right. Thank you for comment !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2761282": "I will share some code that seems to be useful for this competition.\nThis competition involves large data sets, and just opening or saving them can be time-consuming. Therefore, using polars can help reduce time and make the process more comfortable.\n\n##  Targets extraction\n\n~~~\nimport polars as pl\nsample = pl.read_csv(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\",n_rows=1)\n# read only 1 rows\ntargets = sample.columns[1:]\n~~~\n\n## Save submission.csv\n\n~~~\n## change pandas to polars if you use pandas\ntest_polars = pl.from_pandas(test[[\"sample_id\"]+targets])\ntest_polars.write_csv(\"submission.csv\")\n~~~\n\n## Result of time comparison between polars and pandas\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2Fdcdfef4534835d2f46668af6c290c2f5%2FClipboard03.jpg?generation=1713552587674532&alt=media)\n\n## **About 18 times faster using by polars.**",
    "2762755": "You can also convert .csv to .parquet with low RAM\nJust use lazy execution of polars, - \n```\nimport polars as pl\n\nstorage_dir = Path('/path/to/dataset/')\nfor file in os.listdir(storage_dir):\n    df = pl.scan_csv(storage_dir / file)\n    df.sink_parquet(storage_dir / file.replace('.csv', '.parquet'))\n```",
    "2761330": "Try DuckDB also, I think it is also a good option for such large datasets @chumajin ",
    "2761352": "Hi chumajin,\nOne of the things that contributes to the speed of the Polars library is that it is in operation `Eager and Lazy Execution options` and It is derived from the `Rust programming language`.\n`Polars   that follow the Lazy Execution system` Unlike\n`Pandas that follow the Eager Execution system`\n@chumajin "
  }
}