{
  "id": 375961,
  "title": "12th place solution - from public 281st to private 12th!",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375961",
  "author_name": "Jungwoo Park",
  "post_date": "2023-01-04T07:22:55.710000",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, kagglers! I was really surprised that I was in 12th place. Even I thought there should be some shake-ups, but I didn't expect that I would be the only man who climbed the high wall.</p>\n<p>My solution is simple:</p>\n<ul>\n<li>Generate 30k pure signals with sqrtSX=0.</li>\n<li>Combine the pure signals with random noise backgrounds and make 2m combinations. The pure signals are flipped and stretched randomly.</li>\n<li>Train a ConvNext-Small model with 5 epochs, AdamW, lr=3e-4 cosine annealing, and batch size=128.</li>\n<li>Random vertical &amp; horizontal shuffling, random flips, random virtual horizontal lines and beams are used for data augmentation.</li>\n<li>Use 4-way flips as test-time augmentation.</li>\n</ul>\n<p>I also tried an ensemble with many models, but the final score is bad. A single convnext-small model achieves public lb 0.76 and private lb 0.78. In addition, a single convnext-tiny model also achieves public lb 0.757 and private lb 0.78. Very weird 🤔🤔🤔🤔</p>\n<p>You can find my code on <a href=\"https://github.com/affjljoo3581/G2Net-Detecting-Continuous-Gravitational-Waves\" target=\"_blank\">my github repo</a>.</p>",
  "messages": [
    {
      "id": 2085471,
      "postDate": "2023-01-04T07:22:55.710Z",
      "content": "<p>Hi, kagglers! I was really surprised that I was in 12th place. Even I thought there should be some shake-ups, but I didn't expect that I would be the only man who climbed the high wall.</p>\n<p>My solution is simple:</p>\n<ul>\n<li>Generate 30k pure signals with sqrtSX=0.</li>\n<li>Combine the pure signals with random noise backgrounds and make 2m combinations. The pure signals are flipped and stretched randomly.</li>\n<li>Train a ConvNext-Small model with 5 epochs, AdamW, lr=3e-4 cosine annealing, and batch size=128.</li>\n<li>Random vertical &amp; horizontal shuffling, random flips, random virtual horizontal lines and beams are used for data augmentation.</li>\n<li>Use 4-way flips as test-time augmentation.</li>\n</ul>\n<p>I also tried an ensemble with many models, but the final score is bad. A single convnext-small model achieves public lb 0.76 and private lb 0.78. In addition, a single convnext-tiny model also achieves public lb 0.757 and private lb 0.78. Very weird 🤔🤔🤔🤔</p>\n<p>You can find my code on <a href=\"https://github.com/affjljoo3581/G2Net-Detecting-Continuous-Gravitational-Waves\" target=\"_blank\">my github repo</a>.</p>",
      "rawMarkdown": "Hi, kagglers! I was really surprised that I was in 12th place. Even I thought there should be some shake-ups, but I didn't expect that I would be the only man who climbed the high wall.\n\nMy solution is simple:\n- Generate 30k pure signals with sqrtSX=0.\n- Combine the pure signals with random noise backgrounds and make 2m combinations. The pure signals are flipped and stretched randomly.\n- Train a ConvNext-Small model with 5 epochs, AdamW, lr=3e-4 cosine annealing, and batch size=128.\n- Random vertical & horizontal shuffling, random flips, random virtual horizontal lines and beams are used for data augmentation.\n- Use 4-way flips as test-time augmentation.\n\nI also tried an ensemble with many models, but the final score is bad. A single convnext-small model achieves public lb 0.76 and private lb 0.78. In addition, a single convnext-tiny model also achieves public lb 0.757 and private lb 0.78. Very weird 🤔🤔🤔🤔\n\nYou can find my code on [my github repo](https://github.com/affjljoo3581/G2Net-Detecting-Continuous-Gravitational-Waves).",
      "votes": 24
    },
    {
      "id": 2085486,
      "postDate": "2023-01-04T07:35:42.407Z",
      "content": "<p>And I found that the overfitting is occurred when the model is trying hard to memorize weak signal patterns.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Fae8958d991c706080761df6cda5b0d91%2F2023-01-04%20163042.png?generation=1672817532203703&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Ff83a8b0a3190128331e3dac923bed262%2F2023-01-04%20163058.png?generation=1672817544082915&amp;alt=media\" alt=\"\"></p>\n<p>When the training accuracy of the 0.01-0.02 group suddenly rise, all validation metrics become worse. It was really hard to control the overfitting, even with 2m data. As <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> said <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375927\" target=\"_blank\">in his post</a>, it looks like generating more weak signal samples is important.</p>",
      "rawMarkdown": "And I found that the overfitting is occurred when the model is trying hard to memorize weak signal patterns.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Fae8958d991c706080761df6cda5b0d91%2F2023-01-04%20163042.png?generation=1672817532203703&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Ff83a8b0a3190128331e3dac923bed262%2F2023-01-04%20163058.png?generation=1672817544082915&alt=media)\n\nWhen the training accuracy of the 0.01-0.02 group suddenly rise, all validation metrics become worse. It was really hard to control the overfitting, even with 2m data. As @kozistr said [in his post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375927), it looks like generating more weak signal samples is important.",
      "votes": 2
    },
    {
      "id": 2085477,
      "postDate": "2023-01-04T07:26:34.397Z",
      "content": "<p>As I said in a different topic - you must be some kind of Master Yoda of self-control and believing in yourself and your CV! Great job!</p>",
      "rawMarkdown": "As I said in a different topic - you must be some kind of Master Yoda of self-control and believing in yourself and your CV! Great job!",
      "votes": 2
    },
    {
      "id": 2086335,
      "postDate": "2023-01-04T17:57:39.963Z",
      "content": "<p>Thanks for sharing your solution! Very impressive leaderboard jump. This truly shows the importance of cross validation. Congrats <a href=\"https://www.kaggle.com/affjljoo3581\" target=\"_blank\">@affjljoo3581</a>!</p>",
      "rawMarkdown": "Thanks for sharing your solution! Very impressive leaderboard jump. This truly shows the importance of cross validation. Congrats @affjljoo3581!"
    },
    {
      "id": 2085480,
      "postDate": "2023-01-04T07:31:13.373Z",
      "content": "<p>Congrats Jungwoo Park！Very excellent solution! </p>",
      "rawMarkdown": "Congrats Jungwoo Park！Very excellent solution! "
    },
    {
      "id": 2090121,
      "postDate": "2023-01-06T23:56:00.140Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2085486,
      "author_name": "Jungwoo Park",
      "author_url": "",
      "post_date": "2023-01-04T07:35:42.407000",
      "content": "<p>And I found that the overfitting is occurred when the model is trying hard to memorize weak signal patterns.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Fae8958d991c706080761df6cda5b0d91%2F2023-01-04%20163042.png?generation=1672817532203703&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Ff83a8b0a3190128331e3dac923bed262%2F2023-01-04%20163058.png?generation=1672817544082915&amp;alt=media\" alt=\"\"></p>\n<p>When the training accuracy of the 0.01-0.02 group suddenly rise, all validation metrics become worse. It was really hard to control the overfitting, even with 2m data. As <a href=\"https://www.kaggle.com/kozistr\" target=\"_blank\">@kozistr</a> said <a href=\"https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375927\" target=\"_blank\">in his post</a>, it looks like generating more weak signal samples is important.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2085477,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2023-01-04T07:26:34.397000",
      "content": "<p>As I said in a different topic - you must be some kind of Master Yoda of self-control and believing in yourself and your CV! Great job!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2086335,
      "author_name": "Ravi Shah",
      "author_url": "",
      "post_date": "2023-01-04T17:57:39.963000",
      "content": "<p>Thanks for sharing your solution! Very impressive leaderboard jump. This truly shows the importance of cross validation. Congrats <a href=\"https://www.kaggle.com/affjljoo3581\" target=\"_blank\">@affjljoo3581</a>!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2085480,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-01-04T07:31:13.373000",
      "content": "<p>Congrats Jungwoo Park！Very excellent solution! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2090121,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-06T23:56:00.140000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2085471": "Hi, kagglers! I was really surprised that I was in 12th place. Even I thought there should be some shake-ups, but I didn't expect that I would be the only man who climbed the high wall.\n\nMy solution is simple:\n- Generate 30k pure signals with sqrtSX=0.\n- Combine the pure signals with random noise backgrounds and make 2m combinations. The pure signals are flipped and stretched randomly.\n- Train a ConvNext-Small model with 5 epochs, AdamW, lr=3e-4 cosine annealing, and batch size=128.\n- Random vertical & horizontal shuffling, random flips, random virtual horizontal lines and beams are used for data augmentation.\n- Use 4-way flips as test-time augmentation.\n\nI also tried an ensemble with many models, but the final score is bad. A single convnext-small model achieves public lb 0.76 and private lb 0.78. In addition, a single convnext-tiny model also achieves public lb 0.757 and private lb 0.78. Very weird 🤔🤔🤔🤔\n\nYou can find my code on [my github repo](https://github.com/affjljoo3581/G2Net-Detecting-Continuous-Gravitational-Waves).",
    "2085486": "And I found that the overfitting is occurred when the model is trying hard to memorize weak signal patterns.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Fae8958d991c706080761df6cda5b0d91%2F2023-01-04%20163042.png?generation=1672817532203703&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2160097%2Ff83a8b0a3190128331e3dac923bed262%2F2023-01-04%20163058.png?generation=1672817544082915&alt=media)\n\nWhen the training accuracy of the 0.01-0.02 group suddenly rise, all validation metrics become worse. It was really hard to control the overfitting, even with 2m data. As @kozistr said [in his post](https://www.kaggle.com/competitions/g2net-detecting-continuous-gravitational-waves/discussion/375927), it looks like generating more weak signal samples is important.",
    "2085477": "As I said in a different topic - you must be some kind of Master Yoda of self-control and believing in yourself and your CV! Great job!",
    "2086335": "Thanks for sharing your solution! Very impressive leaderboard jump. This truly shows the importance of cross validation. Congrats @affjljoo3581!",
    "2085480": "Congrats Jungwoo Park！Very excellent solution! ",
    "2090121": ""
  }
}