{
  "competition": "rsna-knee-abnormality-detection",
  "topic_id": "742050",
  "comments": [
    {
      "id": 3525911,
      "authorName": "Dread Development",
      "votes": 4,
      "postDate": "2026-09-19T13:41:09.927000",
      "content": "<p>So I’m going to agree with <a href=\"/mattiaangeli\" target=\"_blank\">@mattiaangeli</a> on this, covered it very well.\nPersonally I have been working hard on developing my next iteration, which happens to be much slower that a few tweaks &amp; a submit.</p>\n<p>My current score &amp; best time (0.944 &amp; 13 minutes) has been a 2 week long training process. \nIn the last two weeks I’ve trained 40 different models, attacking this from every possible angle with a goal of catching Scott.\nThe weights I had made public as well as the datasets, I get on average 50 notifications a day (0.924 Raptor)\nSo that bulk section of teams sitting at 0.941 are sitting there due to the mass forking and minor tweaking not actual tuning nor training nor anything that’s not already being done.</p>\n<p>This makes this competition especially slow &amp; difficult for people like me, I personally can only really use the discussions with others to build on new ideas due to a large percentage of public code/notebooks just being forks of my earlier work.</p>\n<p>I did audit the competition a week or so ago &amp; in total, 10,000 downloads over my datasets with roughly 3,000 forks. Just about everyone gives credit so I am appreciative of that! \nJust takes teams like my team longer to produce genuine quality work.\n(I have been non stop working this for weeks &amp; just in the past 2 or so days broke my own best of 0.941 &amp; scoring time of 14 minutes)</p>"
    },
    {
      "id": 3525968,
      "authorName": "Scott Willis",
      "votes": 1,
      "postDate": "2026-09-19T18:11:24.570000",
      "content": "<p>Getting that down to 13 minutes is really impressive.  When I tried using CoAtNet (assuming that's what you're still using), my best submission time was around 25 minutes.</p>"
    },
    {
      "id": 3526030,
      "authorName": "Dread Development",
      "votes": 0,
      "postDate": "2026-09-20T01:17:56.257000",
      "content": "<p>Thank you sir! &amp; Yes, still CoAtNet, paired with a ConvNeXt now.\nThe thing that really surprised me was how little of our runtime is actually the model, with as little as 2 minutes of diff between a single arm &amp; a dual arm run.</p>"
    },
    {
      "id": 3526152,
      "authorName": "starkhushi",
      "votes": -2,
      "postDate": "2026-09-20T14:31:02.677000",
      "content": "<p>Absolutely — that’s the distinction I was trying to highlight. The consistency of the 0.941 score shows how much of the public leaderboard comes from the same recipe and minor tweaks.</p>\n<p>Respect for making your datasets and weights public despite the massive number of forks. And congrats on 0.944 in 13 minutes — that kind of improvement clearly comes from much deeper experimentation.</p>"
    },
    {
      "id": 3525875,
      "authorName": "Mattia Angeli",
      "votes": 4,
      "postDate": "2026-09-19T09:37:29.737000",
      "content": "<p>The bigger point is that basically <strong>no one besides <a href=\"https://www.kaggle.com/dreaddevelopment\" target=\"_blank\">@DreadDevelopment</a> and me has released genuinely new models recently</strong>.</p>\n<p>Building a new model, prototyping it, training it and running proper ablations takes hours or days. Changing <code>0.55/0.15</code> to <code>0.60/0.10</code>, re-ranking, or tweaking per-label blend weights takes minutes.</p>\n<p>Repeat those minute-scale changes across dozens or hundreds of submissions and you are simply overfitting the public LB. A <code>.940 → .941</code> move from another blend tweak is not strong evidence of a better model.</p>\n<p>That is also why I would be very cautious about reading too much into whether 0.55/0.15 or 0.60/0.10 is “better.” Without new models or an independent validation signal, repeatedly searching those blend weights is basically another hyperparameter search on the public split.</p>\n<p>Some of this is probably encouraged by the incentive to reach the top of the public leaderboard and make popular notebooks, but that is different from demonstrating better generalization on the hidden test.</p>"
    },
    {
      "id": 3525876,
      "authorName": "",
      "votes": -2,
      "postDate": "2026-09-19T09:48:05.897000",
      "content": ""
    }
  ],
  "messages": [],
  "raw_show": {
    "topic": {
      "id": 742050,
      "title": "Why ~200 public forks sit at exactly 0.941 ",
      "authorName": "starkhushi",
      "commentCount": 6,
      "votes": 2,
      "postDate": "2026-09-19T07:49:26.076000"
    },
    "comments": [
      {
        "id": 3525911,
        "authorName": "Dread Development",
        "votes": 4,
        "postDate": "2026-09-19T13:41:09.927000",
        "content": "<p>So I’m going to agree with <a href=\"/mattiaangeli\" target=\"_blank\">@mattiaangeli</a> on this, covered it very well.\nPersonally I have been working hard on developing my next iteration, which happens to be much slower that a few tweaks &amp; a submit.</p>\n<p>My current score &amp; best time (0.944 &amp; 13 minutes) has been a 2 week long training process. \nIn the last two weeks I’ve trained 40 different models, attacking this from every possible angle with a goal of catching Scott.\nThe weights I had made public as well as the datasets, I get on average 50 notifications a day (0.924 Raptor)\nSo that bulk section of teams sitting at 0.941 are sitting there due to the mass forking and minor tweaking not actual tuning nor training nor anything that’s not already being done.</p>\n<p>This makes this competition especially slow &amp; difficult for people like me, I personally can only really use the discussions with others to build on new ideas due to a large percentage of public code/notebooks just being forks of my earlier work.</p>\n<p>I did audit the competition a week or so ago &amp; in total, 10,000 downloads over my datasets with roughly 3,000 forks. Just about everyone gives credit so I am appreciative of that! \nJust takes teams like my team longer to produce genuine quality work.\n(I have been non stop working this for weeks &amp; just in the past 2 or so days broke my own best of 0.941 &amp; scoring time of 14 minutes)</p>"
      },
      {
        "id": 3525968,
        "authorName": "Scott Willis",
        "votes": 1,
        "postDate": "2026-09-19T18:11:24.570000",
        "content": "<p>Getting that down to 13 minutes is really impressive.  When I tried using CoAtNet (assuming that's what you're still using), my best submission time was around 25 minutes.</p>"
      },
      {
        "id": 3526030,
        "authorName": "Dread Development",
        "votes": 0,
        "postDate": "2026-09-20T01:17:56.257000",
        "content": "<p>Thank you sir! &amp; Yes, still CoAtNet, paired with a ConvNeXt now.\nThe thing that really surprised me was how little of our runtime is actually the model, with as little as 2 minutes of diff between a single arm &amp; a dual arm run.</p>"
      },
      {
        "id": 3526152,
        "authorName": "starkhushi",
        "votes": -2,
        "postDate": "2026-09-20T14:31:02.677000",
        "content": "<p>Absolutely — that’s the distinction I was trying to highlight. The consistency of the 0.941 score shows how much of the public leaderboard comes from the same recipe and minor tweaks.</p>\n<p>Respect for making your datasets and weights public despite the massive number of forks. And congrats on 0.944 in 13 minutes — that kind of improvement clearly comes from much deeper experimentation.</p>"
      },
      {
        "id": 3525875,
        "authorName": "Mattia Angeli",
        "votes": 4,
        "postDate": "2026-09-19T09:37:29.737000",
        "content": "<p>The bigger point is that basically <strong>no one besides <a href=\"https://www.kaggle.com/dreaddevelopment\" target=\"_blank\">@DreadDevelopment</a> and me has released genuinely new models recently</strong>.</p>\n<p>Building a new model, prototyping it, training it and running proper ablations takes hours or days. Changing <code>0.55/0.15</code> to <code>0.60/0.10</code>, re-ranking, or tweaking per-label blend weights takes minutes.</p>\n<p>Repeat those minute-scale changes across dozens or hundreds of submissions and you are simply overfitting the public LB. A <code>.940 → .941</code> move from another blend tweak is not strong evidence of a better model.</p>\n<p>That is also why I would be very cautious about reading too much into whether 0.55/0.15 or 0.60/0.10 is “better.” Without new models or an independent validation signal, repeatedly searching those blend weights is basically another hyperparameter search on the public split.</p>\n<p>Some of this is probably encouraged by the incentive to reach the top of the public leaderboard and make popular notebooks, but that is different from demonstrating better generalization on the hidden test.</p>"
      },
      {
        "id": 3525876,
        "authorName": "",
        "votes": -2,
        "postDate": "2026-09-19T09:48:05.897000",
        "content": ""
      }
    ]
  },
  "topic": {
    "id": 742050,
    "title": "Why ~200 public forks sit at exactly 0.941 ",
    "authorName": "starkhushi",
    "commentCount": 6,
    "votes": 2,
    "postDate": "2026-09-19T07:49:26.076000"
  },
  "index": {
    "id": "742050",
    "title": "Why ~200 public forks sit at exactly 0.941 ",
    "authorName": "",
    "commentCount": "6",
    "votes": "2",
    "postDate": "2026-09-19 07:49:26.076000"
  }
}