{
  "id": 529229,
  "title": "Are the calibration files the same for all the planets?",
  "url": "/competitions/ariel-data-challenge-2024/discussion/529229",
  "author_name": "Reza R. Choubeh",
  "post_date": "2024-08-19T15:18:12.266000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a><br>\nI checked the calibration files for <strong>AIRS_CH0</strong> sensor and they are all the same for all the planets. That is the <strong>dark.parquet</strong>, <strong>flat.parquet</strong>, <strong>read.parquet</strong>, <strong>dead.parquet</strong> files of one planet is the same as those of others. The same is true with the <strong>FGS1</strong> sensor calibration data.<br>\nI am wondering if I have made a mistake or is this intended?</p>\n<p>For <strong>AIRS_CH0</strong> I used:</p>\n<pre><code> os\n pathlib  Path\n\n pandas  pd\n numpy  np\n\ntrain_dir = Path()\n\nplanet_ids = os.listdir(train_dir)\nfiles = [, , , ]\n\n file  files:\n     planet_id  planet_ids:\n         planet_id==planet_ids[]:\n            compare_with_the_rest_file = pd.read_parquet(train_dir/)\n        :\n            other_file = pd.read_parquet(train_dir/)\n             np.allclose(compare_with_the_rest_file.values, other_file.values)\n</code></pre>\n<p>Checking some individual files manually resulted in the same conclusion.</p>",
  "messages": [
    {
      "id": 2964176,
      "postDate": "2024-08-19T15:52:02.057Z",
      "content": "<p>They are for the train set, but not for the test set.</p>",
      "rawMarkdown": "They are for the train set, but not for the test set.",
      "votes": 3
    },
    {
      "id": 2964157,
      "postDate": "2024-08-19T15:18:12.267Z",
      "content": "<p><a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a><br>\nI checked the calibration files for <strong>AIRS_CH0</strong> sensor and they are all the same for all the planets. That is the <strong>dark.parquet</strong>, <strong>flat.parquet</strong>, <strong>read.parquet</strong>, <strong>dead.parquet</strong> files of one planet is the same as those of others. The same is true with the <strong>FGS1</strong> sensor calibration data.<br>\nI am wondering if I have made a mistake or is this intended?</p>\n<p>For <strong>AIRS_CH0</strong> I used:</p>\n<pre><code> os\n pathlib  Path\n\n pandas  pd\n numpy  np\n\ntrain_dir = Path()\n\nplanet_ids = os.listdir(train_dir)\nfiles = [, , , ]\n\n file  files:\n     planet_id  planet_ids:\n         planet_id==planet_ids[]:\n            compare_with_the_rest_file = pd.read_parquet(train_dir/)\n        :\n            other_file = pd.read_parquet(train_dir/)\n             np.allclose(compare_with_the_rest_file.values, other_file.values)\n</code></pre>\n<p>Checking some individual files manually resulted in the same conclusion.</p>",
      "rawMarkdown": "@gordonyip @lorenzomugnai\nI checked the calibration files for **AIRS_CH0** sensor and they are all the same for all the planets. That is the **dark.parquet**, **flat.parquet**, **read.parquet**, **dead.parquet** files of one planet is the same as those of others. The same is true with the **FGS1** sensor calibration data.\nI am wondering if I have made a mistake or is this intended?\n\nFor **AIRS_CH0** I used:\n\n```python\nimport os\nfrom pathlib import Path\n\nimport pandas as pd\nimport numpy as np\n\ntrain_dir = Path('/kaggle/input/ariel-data-challenge-2024/train')\n\nplanet_ids = os.listdir(train_dir)\nfiles = ['dead', 'flat', 'read', 'dark']\n\nfor file in files:\n    for planet_id in planet_ids:\n        if planet_id==planet_ids[0]:\n            compare_with_the_rest_file = pd.read_parquet(train_dir/f'{planet_id}/AIRS-CH0_calibration/{file}.parquet')\n        else:\n            other_file = pd.read_parquet(train_dir/f'{planet_id}/AIRS-CH0_calibration/{file}.parquet')\n            assert np.allclose(compare_with_the_rest_file.values, other_file.values)\n```\n\nChecking some individual files manually resulted in the same conclusion.",
      "votes": 1
    },
    {
      "id": 2964160,
      "postDate": "2024-08-19T15:26:22.007Z",
      "content": "<p>Look at the calibration files for the test planet and you'll be surprised…</p>",
      "rawMarkdown": "Look at the calibration files for the test planet and you'll be surprised...",
      "votes": 2
    },
    {
      "id": 2965298,
      "postDate": "2024-08-20T18:17:52.107Z",
      "content": "<p>This is working as intended. I've updated the data description to note the duplications.</p>",
      "rawMarkdown": "This is working as intended. I've updated the data description to note the duplications.",
      "replies": [
        {
          "id": 2965852,
          "postDate": "2024-08-21T11:38:55.823Z",
          "content": "<p>So, yet another competition with a very different distribution for the test set and no way to perform proper local validation… Do public and private test sets at least have the same distribution? Can we get a confirmation for that? My nose starts to smell another Novozymes/Belka, lol. I hope to be wrong.</p>",
          "rawMarkdown": "So, yet another competition with a very different distribution for the test set and no way to perform proper local validation... Do public and private test sets at least have the same distribution? Can we get a confirmation for that? My nose starts to smell another Novozymes/Belka, lol. I hope to be wrong.",
          "votes": 2,
          "replies": [
            {
              "id": 2965984,
              "postDate": "2024-08-21T13:02:43.410Z",
              "content": "<p>The previous competitions in this series were without a shakeup at all, although the test was done the same way. This seems to be a good sign.</p>",
              "rawMarkdown": "The previous competitions in this series were without a shakeup at all, although the test was done the same way. This seems to be a good sign.",
              "votes": 7
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2964176,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-08-19T15:52:02.057000",
      "content": "<p>They are for the train set, but not for the test set.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2964160,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2024-08-19T15:26:22.007000",
      "content": "<p>Look at the calibration files for the test planet and you'll be surprised…</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2965298,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2024-08-20T18:17:52.107000",
      "content": "<p>This is working as intended. I've updated the data description to note the duplications.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2965852,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-08-21T11:38:55.823000",
          "content": "<p>So, yet another competition with a very different distribution for the test set and no way to perform proper local validation… Do public and private test sets at least have the same distribution? Can we get a confirmation for that? My nose starts to smell another Novozymes/Belka, lol. I hope to be wrong.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2965984,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-08-21T13:02:43.410000",
              "content": "<p>The previous competitions in this series were without a shakeup at all, although the test was done the same way. This seems to be a good sign.</p>",
              "votes": 7,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2964176": "They are for the train set, but not for the test set.",
    "2964157": "@gordonyip @lorenzomugnai\nI checked the calibration files for **AIRS_CH0** sensor and they are all the same for all the planets. That is the **dark.parquet**, **flat.parquet**, **read.parquet**, **dead.parquet** files of one planet is the same as those of others. The same is true with the **FGS1** sensor calibration data.\nI am wondering if I have made a mistake or is this intended?\n\nFor **AIRS_CH0** I used:\n\n```python\nimport os\nfrom pathlib import Path\n\nimport pandas as pd\nimport numpy as np\n\ntrain_dir = Path('/kaggle/input/ariel-data-challenge-2024/train')\n\nplanet_ids = os.listdir(train_dir)\nfiles = ['dead', 'flat', 'read', 'dark']\n\nfor file in files:\n    for planet_id in planet_ids:\n        if planet_id==planet_ids[0]:\n            compare_with_the_rest_file = pd.read_parquet(train_dir/f'{planet_id}/AIRS-CH0_calibration/{file}.parquet')\n        else:\n            other_file = pd.read_parquet(train_dir/f'{planet_id}/AIRS-CH0_calibration/{file}.parquet')\n            assert np.allclose(compare_with_the_rest_file.values, other_file.values)\n```\n\nChecking some individual files manually resulted in the same conclusion.",
    "2964160": "Look at the calibration files for the test planet and you'll be surprised...",
    "2965298": "This is working as intended. I've updated the data description to note the duplications."
  }
}