{
  "id": 499447,
  "title": "How to detect whether context is competition submission or notebook run",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/499447",
  "author_name": "Chau YH",
  "post_date": "2024-05-01T18:18:15.094000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In other Kaggle competitions, we can detect whether the context is competition submission or notebook run by checking the size of the test set, thereby saving time and GPU hours from the 30hr weekly limit. However, it seems that the publicly available test.csv and sample_submission.csv is the same as the hidden ones.</p>\n<p>This makes GPU-based approaches (such as neural networks) submissions take up the 30hr weekly limit very quickly.</p>\n<p>Does anyone know how to detect the context in notebook code using other methods? Thanks!</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a> <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
  "messages": [
    {
      "id": 2787486,
      "postDate": "2024-05-01T18:40:50.363Z",
      "content": "<p>There is an environment variable that will tell you if it is a submission run</p>\n<pre><code> os\nos.getenv()\n</code></pre>\n<p>you can check this and make it do different behavior based on if it is a submission run or not</p>",
      "rawMarkdown": "There is an environment variable that will tell you if it is a submission run\n\n```python\nimport os\nos.getenv('KAGGLE_IS_COMPETITION_RERUN')\n```\n\nyou can check this and make it do different behavior based on if it is a submission run or not",
      "votes": 3,
      "replies": [
        {
          "id": 2787950,
          "postDate": "2024-05-02T03:02:33.040Z",
          "content": "<p>Thanks for the information! However, this does not seem to work.</p>\n<pre><code> polars  pl\n pandas  pd\n os\n shutil\n\n os.getenv():\n    assert , \n:\n    shutil.copy(, )\n</code></pre>\n<p>My notebook and submission is able to run without errors in both competition submission and notebook run.</p>",
          "rawMarkdown": "Thanks for the information! However, this does not seem to work.\n\n```\nimport polars as pl\nimport pandas as pd\nimport os\nimport shutil\n\nif os.getenv(\"KAGGLE_IS_COMPETITION_RERUN\"):\n    assert False, \"Temp error!\"\nelse:\n    shutil.copy(\"/kaggle/input/leap-atmospheric-physics-ai-climsim/sample_submission.csv\", \"submission.csv\")\n```\n\nMy notebook and submission is able to run without errors in both competition submission and notebook run.",
          "votes": 1,
          "replies": [
            {
              "id": 2788008,
              "postDate": "2024-05-02T04:09:58.753Z",
              "content": "<p>Since this isnt a code competition dont you just run the notebook once on our publicly available test data and then submit the csv? It isnt run a second time in the background on hidden data. We have all the test data</p>",
              "rawMarkdown": "Since this isnt a code competition dont you just run the notebook once on our publicly available test data and then submit the csv? It isnt run a second time in the background on hidden data. We have all the test data",
              "votes": 1
            },
            {
              "id": 2788026,
              "postDate": "2024-05-02T04:22:15.133Z",
              "content": "<p>Ah I see😂. I have only participated in code competitions before. Didn't see the upload .csv option in this competition. Thanks so much for the information!</p>\n<p>Does this mean we can predict using our own computer and its enough to upload the submission.csv file without relying on Kaggle notebooks?</p>",
              "rawMarkdown": "Ah I see😂. I have only participated in code competitions before. Didn't see the upload .csv option in this competition. Thanks so much for the information!\n\nDoes this mean we can predict using our own computer and its enough to upload the submission.csv file without relying on Kaggle notebooks?"
            },
            {
              "id": 2788050,
              "postDate": "2024-05-02T04:46:43.077Z",
              "content": "<p>Yep exactly. You can run inference anywhere and then just upload the file. I havent seen any non-code competitions in a while. Seemed to become the standard. Not sure why they didnt on this one</p>",
              "rawMarkdown": "Yep exactly. You can run inference anywhere and then just upload the file. I havent seen any non-code competitions in a while. Seemed to become the standard. Not sure why they didnt on this one",
              "votes": 1
            },
            {
              "id": 2789974,
              "postDate": "2024-05-03T00:05:40.633Z",
              "content": "<p>You don't see a lot of non code because it's problematic with speech/LLM/CV due to the possibility for hand labeling beating models, impossible with time series future prediction (obviously) and not very wise with small test set that the host want to generalize to (instead of overfitting). For the rest of the competitions, it's not so rare…BELKA is also non-code, as was also Ribonanza. Especially for competitions with large data set, if it's code then it puts limits on ensembles size and model complexity, then the focus shifts from 'best possible solution' (with reasonable compute) to 'best solution that I can run on a notebook in 9 hours'+it's puts huge strain on Kaggle resources for conpute-heavy competitions (9*5*7=315(!) GPU weekly 'test quota' which is more than ten times the regular quota…)</p>",
              "rawMarkdown": "You don't see a lot of non code because it's problematic with speech/LLM/CV due to the possibility for hand labeling beating models, impossible with time series future prediction (obviously) and not very wise with small test set that the host want to generalize to (instead of overfitting). For the rest of the competitions, it's not so rare...BELKA is also non-code, as was also Ribonanza. Especially for competitions with large data set, if it's code then it puts limits on ensembles size and model complexity, then the focus shifts from 'best possible solution' (with reasonable compute) to 'best solution that I can run on a notebook in 9 hours'+it's puts huge strain on Kaggle resources for conpute-heavy competitions (9\\*5\\*7=315(!) GPU weekly 'test quota' which is more than ten times the regular quota...)",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2787455,
      "postDate": "2024-05-01T18:18:15.093Z",
      "content": "<p>In other Kaggle competitions, we can detect whether the context is competition submission or notebook run by checking the size of the test set, thereby saving time and GPU hours from the 30hr weekly limit. However, it seems that the publicly available test.csv and sample_submission.csv is the same as the hidden ones.</p>\n<p>This makes GPU-based approaches (such as neural networks) submissions take up the 30hr weekly limit very quickly.</p>\n<p>Does anyone know how to detect the context in notebook code using other methods? Thanks!</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/mylesoneill\" target=\"_blank\">@mylesoneill</a> <a href=\"https://www.kaggle.com/jerrylin96\" target=\"_blank\">@jerrylin96</a> </p>",
      "rawMarkdown": "In other Kaggle competitions, we can detect whether the context is competition submission or notebook run by checking the size of the test set, thereby saving time and GPU hours from the 30hr weekly limit. However, it seems that the publicly available test.csv and sample_submission.csv is the same as the hidden ones.\n\nThis makes GPU-based approaches (such as neural networks) submissions take up the 30hr weekly limit very quickly.\n\nDoes anyone know how to detect the context in notebook code using other methods? Thanks!\n\n@sohier @inversion @mylesoneill @jerrylin96 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2787486,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2024-05-01T18:40:50.363000",
      "content": "<p>There is an environment variable that will tell you if it is a submission run</p>\n<pre><code> os\nos.getenv()\n</code></pre>\n<p>you can check this and make it do different behavior based on if it is a submission run or not</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2787950,
          "author_name": "Chau YH",
          "author_url": "",
          "post_date": "2024-05-02T03:02:33.040000",
          "content": "<p>Thanks for the information! However, this does not seem to work.</p>\n<pre><code> polars  pl\n pandas  pd\n os\n shutil\n\n os.getenv():\n    assert , \n:\n    shutil.copy(, )\n</code></pre>\n<p>My notebook and submission is able to run without errors in both competition submission and notebook run.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2788008,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2024-05-02T04:09:58.753000",
              "content": "<p>Since this isnt a code competition dont you just run the notebook once on our publicly available test data and then submit the csv? It isnt run a second time in the background on hidden data. We have all the test data</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2788026,
              "author_name": "Chau YH",
              "author_url": "",
              "post_date": "2024-05-02T04:22:15.133000",
              "content": "<p>Ah I see😂. I have only participated in code competitions before. Didn't see the upload .csv option in this competition. Thanks so much for the information!</p>\n<p>Does this mean we can predict using our own computer and its enough to upload the submission.csv file without relying on Kaggle notebooks?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2788050,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2024-05-02T04:46:43.077000",
              "content": "<p>Yep exactly. You can run inference anywhere and then just upload the file. I havent seen any non-code competitions in a while. Seemed to become the standard. Not sure why they didnt on this one</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2789974,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-05-03T00:05:40.633000",
              "content": "<p>You don't see a lot of non code because it's problematic with speech/LLM/CV due to the possibility for hand labeling beating models, impossible with time series future prediction (obviously) and not very wise with small test set that the host want to generalize to (instead of overfitting). For the rest of the competitions, it's not so rare…BELKA is also non-code, as was also Ribonanza. Especially for competitions with large data set, if it's code then it puts limits on ensembles size and model complexity, then the focus shifts from 'best possible solution' (with reasonable compute) to 'best solution that I can run on a notebook in 9 hours'+it's puts huge strain on Kaggle resources for conpute-heavy competitions (9*5*7=315(!) GPU weekly 'test quota' which is more than ten times the regular quota…)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2787486": "There is an environment variable that will tell you if it is a submission run\n\n```python\nimport os\nos.getenv('KAGGLE_IS_COMPETITION_RERUN')\n```\n\nyou can check this and make it do different behavior based on if it is a submission run or not",
    "2787455": "In other Kaggle competitions, we can detect whether the context is competition submission or notebook run by checking the size of the test set, thereby saving time and GPU hours from the 30hr weekly limit. However, it seems that the publicly available test.csv and sample_submission.csv is the same as the hidden ones.\n\nThis makes GPU-based approaches (such as neural networks) submissions take up the 30hr weekly limit very quickly.\n\nDoes anyone know how to detect the context in notebook code using other methods? Thanks!\n\n@sohier @inversion @mylesoneill @jerrylin96 "
  }
}