{
  "id": 357645,
  "title": "HDF files with indexed dataframes for the training set",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/357645",
  "author_name": "JohnM",
  "post_date": "2022-10-05T06:35:00.403000",
  "votes": 16,
  "comment_count": 5,
  "views": 0,
  "content": "<p>You can find processed HDF files for the training set at <a href=\"https://www.kaggle.com/datasets/jpmiller/simplified-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/jpmiller/simplified-dataset</a>. I processed each file so that it contains two datasets, one for  Hanford and one for Livingston.</p>\n<p>The index on each frame is the frequency array from the original file, and the columns are the timestamps. As you probably know, he Hanford and Livingston readings within a sample can have different timestamp counts.</p>\n<p>The files are easily readable by pandas. Here's an example:</p>\n<pre><code>filename = \"../input/simplified-dataset/001121a05.h5\"\ndf_h = pd.read_hdf(filename, key='h')\ndf_l = pd.read_hdf(filename, key='l')\n</code></pre>\n<p>The dataset description contains code to generate these files. </p>",
  "messages": [
    {
      "id": 1972411,
      "postDate": "2022-10-05T06:35:00.403Z",
      "content": "<p>You can find processed HDF files for the training set at <a href=\"https://www.kaggle.com/datasets/jpmiller/simplified-dataset\" target=\"_blank\">https://www.kaggle.com/datasets/jpmiller/simplified-dataset</a>. I processed each file so that it contains two datasets, one for  Hanford and one for Livingston.</p>\n<p>The index on each frame is the frequency array from the original file, and the columns are the timestamps. As you probably know, he Hanford and Livingston readings within a sample can have different timestamp counts.</p>\n<p>The files are easily readable by pandas. Here's an example:</p>\n<pre><code>filename = \"../input/simplified-dataset/001121a05.h5\"\ndf_h = pd.read_hdf(filename, key='h')\ndf_l = pd.read_hdf(filename, key='l')\n</code></pre>\n<p>The dataset description contains code to generate these files. </p>",
      "rawMarkdown": "You can find processed HDF files for the training set at https://www.kaggle.com/datasets/jpmiller/simplified-dataset. I processed each file so that it contains two datasets, one for  Hanford and one for Livingston.\n\nThe index on each frame is the frequency array from the original file, and the columns are the timestamps. As you probably know, he Hanford and Livingston readings within a sample can have different timestamp counts.\n\nThe files are easily readable by pandas. Here's an example:\n\n```\nfilename = \"../input/simplified-dataset/001121a05.h5\"\ndf_h = pd.read_hdf(filename, key='h')\ndf_l = pd.read_hdf(filename, key='l')\n```\n\nThe dataset description contains code to generate these files. ",
      "votes": 15
    },
    {
      "id": 1972632,
      "postDate": "2022-10-05T08:57:07.053Z",
      "content": "<p>Thank you ! It's will be easier. <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> do you plan to release the test set as well please in this format? Thank you.</p>",
      "rawMarkdown": "Thank you ! It's will be easier. @jpmiller do you plan to release the test set as well please in this format? Thank you.",
      "replies": [
        {
          "id": 1973262,
          "postDate": "2022-10-05T14:53:29.993Z",
          "content": "<p>I'd like to but haven't figured out a way to stay within the dataset size limits on Kaggle. </p>",
          "rawMarkdown": "I'd like to but haven't figured out a way to stay within the dataset size limits on Kaggle. "
        },
        {
          "id": 1973857,
          "postDate": "2022-10-05T22:34:21.897Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1972538,
      "postDate": "2022-10-05T07:44:18.437Z",
      "content": "<p>This is likely to be of great help, many thanks and regards <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> </p>",
      "rawMarkdown": "This is likely to be of great help, many thanks and regards @jpmiller "
    },
    {
      "id": 2053196,
      "postDate": "2022-12-03T00:38:18.330Z",
      "content": "<p>Thank you! </p>",
      "rawMarkdown": "Thank you! "
    }
  ],
  "comments": [
    {
      "id": 1972632,
      "author_name": "Defend Intelligence",
      "author_url": "",
      "post_date": "2022-10-05T08:57:07.053000",
      "content": "<p>Thank you ! It's will be easier. <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> do you plan to release the test set as well please in this format? Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1973262,
          "author_name": "JohnM",
          "author_url": "",
          "post_date": "2022-10-05T14:53:29.993000",
          "content": "<p>I'd like to but haven't figured out a way to stay within the dataset size limits on Kaggle. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1973857,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-05T22:34:21.897000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1972538,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-10-05T07:44:18.437000",
      "content": "<p>This is likely to be of great help, many thanks and regards <a href=\"https://www.kaggle.com/jpmiller\" target=\"_blank\">@jpmiller</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2053196,
      "author_name": "Temari",
      "author_url": "",
      "post_date": "2022-12-03T00:38:18.330000",
      "content": "<p>Thank you! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1972411": "You can find processed HDF files for the training set at https://www.kaggle.com/datasets/jpmiller/simplified-dataset. I processed each file so that it contains two datasets, one for  Hanford and one for Livingston.\n\nThe index on each frame is the frequency array from the original file, and the columns are the timestamps. As you probably know, he Hanford and Livingston readings within a sample can have different timestamp counts.\n\nThe files are easily readable by pandas. Here's an example:\n\n```\nfilename = \"../input/simplified-dataset/001121a05.h5\"\ndf_h = pd.read_hdf(filename, key='h')\ndf_l = pd.read_hdf(filename, key='l')\n```\n\nThe dataset description contains code to generate these files. ",
    "1972632": "Thank you ! It's will be easier. @jpmiller do you plan to release the test set as well please in this format? Thank you.",
    "1972538": "This is likely to be of great help, many thanks and regards @jpmiller ",
    "2053196": "Thank you! "
  }
}