{
  "topic": {
    "id": 735304,
    "title": "Best single-model score",
    "authorName": "roy214",
    "commentCount": 85,
    "votes": 59,
    "postDate": "2026-08-14T23:38:09.839000"
  },
  "comments": [
    {
      "id": 3525138,
      "authorName": "Archit Konde",
      "votes": 5,
      "postDate": "2026-09-16T23:32:05.800000",
      "content": "<p>Got a single fold CoAtNet at 224px to 0.950 right now.</p>"
    },
    {
      "id": 3525155,
      "authorName": "huytngo",
      "votes": 1,
      "postDate": "2026-09-17T01:20:59.037000",
      "content": "<p>Nice! do you mind sharing a bit of your thought process behind getting that single fold? anyways congrats!</p>"
    },
    {
      "id": 3525193,
      "authorName": "Archit Konde",
      "votes": 2,
      "postDate": "2026-09-17T05:25:24.790000",
      "content": "<p>Thank you! The main thing was one check. On the 58 provided labels, my OOF predictions were already closer to the radiologists than the extracted labels the model was trained on. Bigger models weren't helping either, so I figured the labels were the bottleneck and started working on that. The model is pretty standard: 2.5D CoAtNet at 224 (a few neighbouring slices at a time).</p>"
    },
    {
      "id": 3525255,
      "authorName": "Salem Ali",
      "votes": 1,
      "postDate": "2026-09-17T09:04:15.410000",
      "content": "<p>That’s interesting! I got a mean AUC of 0.9208 on the 58 provided labels.Synovitis was my lowest-scoring class at 0.7969, followed by PF OA at 0.8533.\nwhat did you get on the 58 labels? And what were your Synovitis and PF OA scores? Which classes were the main bottlenecks for you?</p>"
    },
    {
      "id": 3525491,
      "authorName": "Tom Aindow",
      "votes": 2,
      "postDate": "2026-09-17T21:12:24.357000",
      "content": "<p>Congratulations of the great model performance! 0.950 single fold is amazing.</p>\n<p>Seems like the labels are really the key here, then. I've tried lots though and haven't had much success. I thought that perhaps using my models strongest predictions to highlight noisy labels might help, whether by masking them in training or by replacing them outright, but neither really have any noticeable effect. </p>\n<p>But then again, I suppose feeding a model back it's own predictions isn't exactly telling it anything new, it just makes it easier to reach the same point. </p>\n<p>I'm scratching my head trying to think about how to improve these labels further. Maybe multiple models with different strengths can teach each other when to listen to the teacher and when to ignore it? </p>\n<p>If you had any gentle hints to spare it would be much appreciated !</p>"
    },
    {
      "id": 3525700,
      "authorName": "Archit Konde",
      "votes": 1,
      "postDate": "2026-09-18T17:25:03.347000",
      "content": "<p>Single model (out of fold) is 0.930 on the 58 for me. Synovitis 0.797 and Lateral OA 0.816 are my weakest, PF OA 0.874. Our Synovitis numbers are nearly identical, which is interesting. I wouldn't read too much into the per-class values though, some classes only have 9-12 positives across the 58 studies.</p>"
    },
    {
      "id": 3525701,
      "authorName": "Archit Konde",
      "votes": 1,
      "postDate": "2026-09-18T17:26:25.743000",
      "content": "<p>Thank you! Of course, happy to help.</p>\n<p>Masking the noisy labels did nothing for me either. Replacing them with the model's predictions came out worse than mixing the two, so I ended up keeping the extracted label and mixing the prediction into it rather than choosing between them. Why that helps, I'm honestly not sure. My guess is it's less about new information and more about how confident the label is, but that's a guess.</p>\n<p>Rounding the labels to 0/1 also hurt; keeping them as fractions was better. And I found it useful to look at calibration on the 58 rather than only AUC. Mine happened to line up already (labels around 0.3 were positive roughly 30% of the time), which is what made me stop poking at the extraction.</p>\n<p>Have you found anything that moves the diffuse classes? Synovitis and Lateral OA are where I'm stuck, and nothing I've tried on the model side seems to touch them.</p>"
    },
    {
      "id": 3525942,
      "authorName": "Tom Aindow",
      "votes": 0,
      "postDate": "2026-09-19T16:29:49.190000",
      "content": "<p>Interesting I've also tried blending prediction and labels together but didn't have much luck, something to come back to perhaps.</p>\n<p>Synovitis I have tried to attack with little luck. It's clearly under reported I think, based on the golden data, so I suspect there are lots of false negatives in my labels. It's there, present, but they don't say anything about it. I tried to filter out these false positives by basically masking all negative synovitis labels except the ones where it was explicitly declared absent (not many), or was negative alongside lots of other negatives (I.e. the whole scan basically came back normal). This didn't do anything to my result though. I suspect masking false negatives isn't enough: the model needs more positives.</p>"
    },
    {
      "id": 3525225,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-09-17T07:30:17.437000",
      "content": ""
    },
    {
      "id": 3524183,
      "authorName": "Less",
      "votes": 4,
      "postDate": "2026-09-14T07:41:14.233000",
      "content": "<p>Single fold, CoAtNet, 384*384, LB 0.937, but fusion is ineffective.</p>"
    },
    {
      "id": 3520201,
      "authorName": "Jack",
      "votes": 3,
      "postDate": "2026-09-03T00:21:47.753000",
      "content": "<p>0.930 in 46 minutes, hoping a label improvement experiment pushes this further</p>"
    },
    {
      "id": 3519217,
      "authorName": "Yann Majewski",
      "votes": 6,
      "postDate": "2026-09-01T10:06:34.323000",
      "content": "<p>0.936lb, single fold, small resnet model, 224x224</p>"
    },
    {
      "id": 3519327,
      "authorName": "Salem Ali",
      "votes": 1,
      "postDate": "2026-09-01T14:48:30.450000",
      "content": "<p>Really impressive results with such a small resnet! I’m curious about your approach to handling label noise. Would you be willing to share some details?</p>"
    },
    {
      "id": 3520002,
      "authorName": "Yann Majewski",
      "votes": 2,
      "postDate": "2026-09-02T17:04:42.500000",
      "content": "<p>I did a little bit of label work, combined my labels with some public dataset and it increased lb score by 0.015</p>\n<p>Then some pre training on the encoder and training longer with more regularization helped to reach 0.936!</p>\n<p>Its a little hard to properly evaluate what really helped, when I made too many changes in between runs (e.g. changing the encoder architecture), I stopped seeing a correlation between validation and lb…</p>"
    },
    {
      "id": 3520075,
      "authorName": "Salem Ali",
      "votes": 1,
      "postDate": "2026-09-02T19:25:43.800000",
      "content": "<p>I’ve found the same thing in my experiments data quality and training strategy can matter more than just making the model more complex. 0.936 with one fold ,a small ResNet is impressive,\ncongrats!</p>"
    },
    {
      "id": 3520511,
      "authorName": "tennogh",
      "votes": 2,
      "postDate": "2026-09-03T15:13:03.147000",
      "content": "<p>So far 0.942 raw, 0.943 with TTA at 288x288. OOF pseudo-labels have been pretty well correlated with LB. </p>"
    },
    {
      "id": 3520560,
      "authorName": "Komil Parmar",
      "votes": 0,
      "postDate": "2026-09-03T16:54:13.597000",
      "content": "<p>That's really high! At just 288x288 resolution. If resolution was just 288x288 to get such high scores, I wonder what are the core drivers of such a high score then. If you would be willing to share any hints… Great job regardless!</p>"
    },
    {
      "id": 3520564,
      "authorName": "tennogh",
      "votes": 2,
      "postDate": "2026-09-03T17:00:32.597000",
      "content": "<p>Just a lot of incremental improvements. I have my own labels but they are not measurably better than the public ones. I'm not sure resolution is such a big driver.</p>"
    },
    {
      "id": 3520567,
      "authorName": "Tucker Arrants",
      "votes": 3,
      "postDate": "2026-09-03T17:09:59.857000",
      "content": "<p>Agreed with resolution not being that important. I'm sure its backbone dependent and I've only tried ResNets and EffNets, but 224 / 288 is the sweet spot so far. </p>"
    },
    {
      "id": 3520569,
      "authorName": "Komil Parmar",
      "votes": 1,
      "postDate": "2026-09-03T17:16:33.747000",
      "content": "<p>Wow, that's really interesting and worring at the same time. I gotta be finding the missing peice staying low-res. Thankyou <a href=\"https://www.kaggle.com/tennogh\" target=\"_blank\">@tennogh</a> <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> </p>"
    },
    {
      "id": 3520743,
      "authorName": "Shuolin Liu",
      "votes": 0,
      "postDate": "2026-09-04T03:32:02.427000",
      "content": "<p>Thanks for sharing! I'm planning to start training with public labels. Do you think it's better to start with hard (0/1) or soft labels?</p>"
    },
    {
      "id": 3520745,
      "authorName": "Tucker Arrants",
      "votes": 0,
      "postDate": "2026-09-04T03:35:34.647000",
      "content": "<p>Start with soft but worth trying both so you can measure the value of the soft labels. The delta between the two is itself insightful.</p>"
    },
    {
      "id": 3523155,
      "authorName": "Cody_Null",
      "votes": 0,
      "postDate": "2026-09-10T17:54:44.050000",
      "content": "<p>Have you found any magic slices/windows for 224 res? I am trying it now with 0.446 mm/px on 64 slices 62 windows but my model scores struggle. Cant make out if my data is the problem or I have just busted something in my pipeline hahaha </p>"
    },
    {
      "id": 3523270,
      "authorName": "Parag",
      "votes": 1,
      "postDate": "2026-09-11T00:06:07.497000",
      "content": "<p>At 100mm Field of View you may be losing some useful information. My sweet spot is 140mm to 160mm FOV</p>"
    },
    {
      "id": 3523293,
      "authorName": "Cody_Null",
      "votes": 0,
      "postDate": "2026-09-11T03:03:11.693000",
      "content": "<p>Ill have a think about that one! Thanks!</p>"
    },
    {
      "id": 3515147,
      "authorName": "Scott Willis",
      "votes": 6,
      "postDate": "2026-08-20T22:27:20.557000",
      "content": "<p>I'm at 0.938 with a single model, single fold.</p>\n<p>I'm also pretty new to all of this, so I'm going to assume that I'm missing a lot that I could be doing to help out my model.</p>"
    },
    {
      "id": 3515634,
      "authorName": "Tom Aindow",
      "votes": 0,
      "postDate": "2026-08-22T07:24:25.747000",
      "content": "<p>Really impressive result, congratulations! Would you be willing to share any details about how you are approaching the problem? It seems you are also top of the efficiency leaderboard, so perhaps you are doing something quite different to most !</p>"
    },
    {
      "id": 3515986,
      "authorName": "Scott Willis",
      "votes": 1,
      "postDate": "2026-08-22T16:55:33.927000",
      "content": "<p>I just tried to start with as small of a model as possible.  My whole goal was to start aiming for the efficiency LB and then go up from there if needed.  I also spent a lot of time trying to figure out how to work around the low quality labels I have.  I did break down and submit a 5 fold ensemble though, so my current 0.947 is no longer a single fold.</p>\n<p>I'm not too confident my score is going to hold on the private LB, though.  Should be interesting to see.</p>"
    },
    {
      "id": 3519159,
      "authorName": "Berat Kirbiyik",
      "votes": 1,
      "postDate": "2026-09-01T07:11:38.647000",
      "content": "<p>Congrats on holding #1 on the efficiency board — the \"start as small as possible\" framing is the opposite of what most of us drifted into, and clearly it worked.</p>\n<p>The part I'd love to hear more about is the labels. You said you spent a lot of time working around the low-quality labels. Were you working from one of the public LLM-extracted sets, or did you build your own extraction? And when you say \"working around\" — is that a labelling-side fix (better extraction, soft targets, per-class confidence) or a training-side one (loss weighting, sample reweighting, masking uncertain cells)?</p>\n<p>Asking because we're stuck at 0.924 and every architectural lever we've tried has come back inside noise: larger backbone, cross-slot attention, top-k pooling, EMA, mixup, longer schedules, slice-count sweeps. Seven ablations, all within 0.008 of each other. That pattern usually means the ceiling is upstream of the model, and your comment is the first thing I've read that points squarely at it.</p>\n<p>Happy to share our timing profile in return — we got a CoAtNet-384 run down from 65 minutes to 17 by pipelining preprocessing against GPU, if that's useful to anyone.</p>"
    },
    {
      "id": 3521384,
      "authorName": "Devguru Tiwari",
      "votes": 0,
      "postDate": "2026-09-05T17:11:36.443000",
      "content": "<p>please help me what to do to reduce time</p>"
    },
    {
      "id": 3521391,
      "authorName": "Scott Willis",
      "votes": 7,
      "postDate": "2026-09-05T17:40:26.257000",
      "content": "<p>I'm not using public labels.  I used Gemma 4 locally to extract labels from the reports.  I ran through it multiple times until it was \"good enough\" and they ended up being around .89.   I haven't tried public labels at all, like most people have said, I haven't really found a huge difference in the label quality right now.  My current LB best (0.949) is a single fold, single model and scores around .930 on gold.  I'm basically out of ideas right without sacrificing effeciency.</p>\n<p>Scoring usually takes around 5 minutes, if that's helpful. </p>"
    },
    {
      "id": 3521398,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-09-05T18:11:16.247000",
      "content": "<p>Single fold, single model, 5 minute sub 0.949 is unreal haha I cant even process the images that fast right now</p>"
    },
    {
      "id": 3522662,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-09-09T13:28:24.803000",
      "content": ""
    },
    {
      "id": 3522684,
      "authorName": "Scott Willis",
      "votes": 4,
      "postDate": "2026-09-09T14:54:03.640000",
      "content": "<p>I've really only ever spent time on efficentnet and resnet.  I spent a little time on others but at least this point I want to see how much I can get out of them.</p>"
    },
    {
      "id": 3522723,
      "authorName": "huytngo",
      "votes": 0,
      "postDate": "2026-09-09T16:56:31.147000",
      "content": "<blockquote>\n  <p>I'm at 0.938 with a single model, single fold.</p>\n  <p>I'm also pretty new to all of this, so I'm going to assume that I'm missing a lot that I could be doing to help out my model.</p>\n</blockquote>\n<p>Bro you are a beast man, appreciate the advice on this thread.</p>"
    },
    {
      "id": 3524586,
      "authorName": "MJ121212",
      "votes": 0,
      "postDate": "2026-09-15T10:37:19.463000",
      "content": "<p>what base model did you use for this and where did you get/assemble your labels from ?</p>"
    },
    {
      "id": 3524704,
      "authorName": "Scott Willis",
      "votes": 0,
      "postDate": "2026-09-15T16:53:47.940000",
      "content": "<p>I'm using a smaller resnet.  Honestly at this point I'd be shocked if I'm not overfitting and my score holds up on the private LB.</p>"
    },
    {
      "id": 3524919,
      "authorName": "Optimo",
      "votes": 3,
      "postDate": "2026-09-16T08:41:30.117000",
      "content": "<p><a href=\"/scottgwillis\" target=\"_blank\">@scottgwillis</a> What makes you worry about overfitting the public LB ? You do have a lot of subs, but a single fold single model running in 5 min and reaching such high scores seems like a solid solution instead of a worrying one. Are you relying a lot on LB to make changes to your approach ? Do you have a good CV/LB correlation ? Did you tweak any parameter on LB alone ?</p>"
    },
    {
      "id": 3525008,
      "authorName": "Lavin Wins",
      "votes": -7,
      "postDate": "2026-09-16T15:17:39.617000",
      "content": "<p><a href=\"/scottgwillis\" target=\"_blank\">@scottgwillis</a>  , Bro you are at the top , can you please give me some advice related how to get started , like I used public baseline at recieved 0.94 so what next step should I take or what was your strategy?</p>"
    },
    {
      "id": 3515050,
      "authorName": "k256.dev",
      "votes": 4,
      "postDate": "2026-08-20T16:49:30.603000",
      "content": "<p>single(5-fold) model 0.926 @ 336px</p>\n<p>LLM output labels only</p>"
    },
    {
      "id": 3515696,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-08-22T09:22:52.693000",
      "content": ""
    },
    {
      "id": 3513019,
      "authorName": "PC Jimmmy",
      "votes": 8,
      "postDate": "2026-08-15T01:32:37.950000",
      "content": "<p>In my option - 9 days into a competition is a very silly time to be wasting efforts refining ensemble models.  </p>"
    },
    {
      "id": 3513461,
      "authorName": "Tucker Arrants",
      "votes": 6,
      "postDate": "2026-08-16T15:23:04.520000",
      "content": "<p>Single model 0.934 @ 224px </p>\n<p>Update: 0.943 with some tweaks, still @ 224px</p>"
    },
    {
      "id": 3513553,
      "authorName": "Chikuwabu",
      "votes": 0,
      "postDate": "2026-08-17T00:32:44.853000",
      "content": "<p>Is this a single fold, or N-fold ensemble?</p>"
    },
    {
      "id": 3513573,
      "authorName": "Tucker Arrants",
      "votes": 3,
      "postDate": "2026-08-17T01:52:13.533000",
      "content": "<p>5 fold ensemble</p>"
    },
    {
      "id": 3513878,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-08-17T21:00:36.593000",
      "content": ""
    },
    {
      "id": 3513879,
      "authorName": "Tom Aindow",
      "votes": 1,
      "postDate": "2026-08-17T21:00:54.030000",
      "content": "<p>Really impressive result at 224x, are you willing to give any insight into your approach, either labeling or modelling? My current thinking is that labeling can be a significant source of error, since many labels in the gold set do not agree with the reports when the diagnosis rules are applied strictly (e.g. small effusions being labeled as 1). </p>"
    },
    {
      "id": 3513883,
      "authorName": "Tucker Arrants",
      "votes": 0,
      "postDate": "2026-08-17T21:33:40.020000",
      "content": "<p>My label extraction is very crude - haven’t spent much time on it. But yes, I agree, I think it’s important here.</p>\n<p>The model architecture is standard for these types of competitions and relatively small, currently taking around 25m for preprocessing and inference via submission scoring, including TTA. </p>"
    },
    {
      "id": 3518445,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-08-30T23:06:46.997000",
      "content": "<p>Are you guys only using the 58 to validate labels and then using the LB as your CV effectively? Or are you trusting the created labels to compute CV? Curious what some are getting for their CV and label scores compared to gold 58. </p>"
    },
    {
      "id": 3518449,
      "authorName": "Tucker Arrants",
      "votes": 3,
      "postDate": "2026-08-31T00:07:03.007000",
      "content": "<p>I only validate against the report generated labels, never the provided labels. The sample size of the gold labels is just too small to resolve anything.</p>\n<p>I've only used the LB to check correlation between it and my report generated labels CV, which is, so far, strongly correlated. If the correlation breaks, might just treat it like BirdCLEF and use LB for CV.</p>\n<p>My CV is around 0.87 - 0.9 depending on the generated labels used, all leading to the same ~0.94 leaderboard score. </p>"
    },
    {
      "id": 3518458,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-08-31T01:02:56.787000",
      "content": "<p>Great! That’s about what my plan was too! Have you tried comparing the score of your generated labels to the 58 gold? Same issues exist with small sample size and the domain issue of their being natural disagreement between experts. But curious if you’ve managed to score much differently than I have. I’m at about .9 AUC. And another method at .92 that I worry is probably just a bit too good to be true.</p>"
    },
    {
      "id": 3518459,
      "authorName": "Tucker Arrants",
      "votes": 2,
      "postDate": "2026-08-31T01:10:38.977000",
      "content": "<p>Yes I did a while back, the label extractor against gold is around yours, a bit lower I believe. Maybe 0.89. I've read the 58 reports of the gold labels, and my extraction is not missing anything - the reports genuinely do not contain the information required to generate the corresponding label, which is that domain issue you just mentioned. The radiologist writting the report either missed or simply did not care to report that particular label. </p>\n<p>Spent the last week or so getting \"better\" labels and none of it translated to the LB. Single prompt LLM extraction has not been beaten, outside of noise, by any more sophisticated extraction procedures that I have tried. </p>"
    },
    {
      "id": 3518461,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-08-31T01:15:10.047000",
      "content": "<p>I get the same result. Any time I try more than that my resulting curves definitely look like overfitting more than actual better labels. I assume the .89 you have also gets better results than the public .89? For me I did fair with public labels but using a different model got .9 and the cv is much much better. </p>"
    },
    {
      "id": 3518462,
      "authorName": "Tucker Arrants",
      "votes": 1,
      "postDate": "2026-08-31T01:25:26.263000",
      "content": "<p>Exactly, I try not to look at the gold reports too much to avoid overfitting prompts to them. I actually haven't tried the public labels at all, as I am not totally sure how they were generated. My best current LB model is from my first pass at extraction, with thousands of verified extraction failures. I correct them, no change in LB. Image model \"student\" seems to consistently outperform the label extractor \"teacher\". Where it struggles are on the labels that are just poorly reported in general.</p>"
    },
    {
      "id": 3518464,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-08-31T01:28:32.797000",
      "content": "<p>Maybe I am not as far off as I thought then. Thanks so much for sharing, bringing back the kaggle spirit! </p>"
    },
    {
      "id": 3518466,
      "authorName": "Tucker Arrants",
      "votes": 1,
      "postDate": "2026-08-31T01:40:14.840000",
      "content": "<p>Of course. I could be wrong, but I lean more \"extracted labels are probably fine\" than \"extracted labels are the main lever to pull here\". I'm not the best with this whole LLM thing, but based on all the work I've done with labels so far, it has not made a difference. The fundamental challenge is that report-&gt;label vs. image-&gt;label gap, which no amount of extraction tweaking can fix.</p>"
    },
    {
      "id": 3518471,
      "authorName": "Cody_Null",
      "votes": 1,
      "postDate": "2026-08-31T01:58:30.613000",
      "content": "<p>I am personally much better with LLMs than image models and I think LLMs would be the main lever if there wasnt such a grey area there in the middle. The actual difference between a .92 and a .89 cv on those 58 gold samples is really explained by just small disagreement between only a few samples which would be natural in the field. So I think here it could be a difference but THE difference would probably be more about how you actually handle labels in modeling. At least thats my take on it, going to be fun to see how it all goes in the end!</p>"
    },
    {
      "id": 3518522,
      "authorName": "tennogh",
      "votes": 1,
      "postDate": "2026-08-31T07:05:15.310000",
      "content": "<p>It does seem like in many cases the reports just don't contain the information and clever \"fixes\" end up overfitting to the 58 set. Some of the disagreements are also interesting, could be different diagnosis policies or missing context (e.g. previous MRI/x-ray). Maybe I'll find a radiologist to look at some examples. </p>"
    },
    {
      "id": 3518960,
      "authorName": "Salem Ali",
      "votes": 1,
      "postDate": "2026-08-31T22:26:29.457000",
      "content": "<p>I completely agree with your point. i came to the exact same conclusion: due to the inherent noise in the data and the fact that radiologists simply don't mention everything in their reports, the report-&gt;label mapping can never truly equal the image-&gt;label reality. Just like in your experience, i found that no matter how much I tried to refine and tweak the labels, it had no meaningful impact on the LB.\n​Since the image model (Student) manages to bypass these textual gaps and detect the true visual features, have you tried using the model's high-confidence predictions to relabel the noisy data and train iteratively? or do you avoid this to prevent a confirmation bias</p>"
    },
    {
      "id": 3519005,
      "authorName": "Tucker Arrants",
      "votes": 1,
      "postDate": "2026-09-01T00:33:33.983000",
      "content": "<p>Yes, its a good idea and worth trying. But you certainly can't validate against labels you repair or your CV will improve by definition. Would be like pseudo labeling then validating against your pseudo labels. </p>"
    },
    {
      "id": 3521265,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-09-05T10:56:37.857000",
      "content": ""
    },
    {
      "id": 3521401,
      "authorName": "Tucker Arrants",
      "votes": 1,
      "postDate": "2026-09-05T18:39:12.240000",
      "content": "<p>I use them in training, but not for validation. They make up around 1.3% of loss mass so not going to make much difference with or without, but they are not useful for CV so might as well train with them so they aren't completely wasted.</p>"
    },
    {
      "id": 3521514,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-09-06T06:54:14.960000",
      "content": ""
    },
    {
      "id": 3515988,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-08-22T16:55:59.833000",
      "content": ""
    },
    {
      "id": 3521095,
      "authorName": "Charles Savas",
      "votes": -6,
      "postDate": "2026-09-04T20:53:32.327000",
      "content": "<p>Hi Tucker — we're the team at 0.943 (charlessavas). Your posts on label handling match what we've measured: we probed public per-target AUC by column ablation, priced the report-to-image label ceiling per target, and built a one-hour harness that A/B tests any external model against our 8-signal ensemble.</p>\n<p>Our lineup is DINOv2-MIL-heavy; yours sounds highly decorrelated from it. Our one external arm (correlation 0.83) was worth +0.004 public across four submissions.</p>\n<p>Proposal: compare model and validation evidence first. If the families look independent, merge through Kaggle; then measure prediction correlation inside the team before changing either pipeline. A combination plausibly gives us a path through the current 0.949 top-10 cutoff. We'd bring the probe map, the external-integration harness, and a documented do-not-retry log of about 40 dead ends. Interested?</p>"
    },
    {
      "id": 3513077,
      "authorName": "Tom Aindow",
      "votes": 5,
      "postDate": "2026-08-15T06:46:10.823000",
      "content": "<p>Single model 0.915, DinoV2 based. </p>\n<p>Highly doubt the top results are ensembles, look at the <a href=\"https://www.kaggle.com/code/ryanholbrook/rsna-knee-abnormalities-efficiency-lb\" target=\"_blank\">efficiency lb</a>, you can see many high scoring submissions also have decent run times, i.e. they're not running 20 different models.</p>\n<p>Public notebooks just ensemble early on to try and get a high leaderboard score and exposure/kudos on their notebooks. Waste of time IMO.</p>"
    },
    {
      "id": 3513424,
      "authorName": "Berat Kirbiyik",
      "votes": 0,
      "postDate": "2026-08-16T12:51:42.463000",
      "content": "<p>Single model 0.92 is impressive. Would you mind sharing the input geometry you landed on — resolution, physical crop in mm, and slices per series?</p>\n<p>We're at 0.826 with 192px / 9 slices / 4 sequence slots and trying to work out whether our ceiling is the input pipeline or the model itself.</p>"
    },
    {
      "id": 3513433,
      "authorName": "Tim Krige",
      "votes": 2,
      "postDate": "2026-08-16T13:27:02.583000",
      "content": "<p>Sharing is quite revealing on my methods. What I can strongly suggest - look at the images your model sees. Does it look detailed enough? Do you think 192px is sufficient to see mm level details? </p>\n<p>I find playing with the pipeline is crucial.  If I can it see the details, how could a model?</p>\n<p>Do research on the labels. Learn what the model needs to see. Llms are very wrong a lot of the time here without careful guidance.</p>"
    },
    {
      "id": 3513877,
      "authorName": "Tom Aindow",
      "votes": 3,
      "postDate": "2026-08-17T20:55:27.327000",
      "content": "<p>Sure, I'm using 392x with 150mm center crop (0.383 mm/px) and I randomly sample a bag of 32 slices per study across all studies during training. I'm still not quite sure what does and doesn't work here. I suspect that labels may actually be one of the most important parts.</p>"
    },
    {
      "id": 3513030,
      "authorName": "Chikuwabu",
      "votes": 5,
      "postDate": "2026-08-15T03:07:13.947000",
      "content": "<p>A single model actually works much better than you'd expect.</p>"
    },
    {
      "id": 3521524,
      "authorName": "Devguru Tiwari",
      "votes": 0,
      "postDate": "2026-09-06T08:11:11.087000",
      "content": "<p>Could you tell me the labels you are using. Since, i dont have any personal labels and using public available labels. Also, any hint on how to improve methodology will be helpful. Thank you in adv man.</p>"
    },
    {
      "id": 3513182,
      "authorName": "Tim Krige",
      "votes": 3,
      "postDate": "2026-08-15T14:37:39.307000",
      "content": "<p>I am on a single model. OOF AUC = LB AUC within noise, 0.92. </p>"
    },
    {
      "id": 3526158,
      "authorName": "Shiv Satyam",
      "votes": 0,
      "postDate": "2026-09-20T14:42:00.697000",
      "content": "<p>I got 91.0 using a Resnet34 backbone coupled with a mundane linear head.</p>"
    },
    {
      "id": 3525010,
      "authorName": "kmn",
      "votes": 0,
      "postDate": "2026-09-16T15:28:32.040000",
      "content": "<p>I'm at 0.930 with a single model (CoAtNet, 1h22min), and I haven't tried ensembles yet.</p>"
    },
    {
      "id": 3524699,
      "authorName": "Raymond Yuen",
      "votes": 0,
      "postDate": "2026-09-15T16:41:09.617000",
      "content": "<p>0.938 Lb. full train CoAtNet with 288*288. For now labels are still my greatest lever, especially pseudo labeling and blending different labels.</p>\n<p>Update: Now at 0.94 lb. still label improvements</p>\n<p>Update: Now 0.943. still label :D</p>"
    },
    {
      "id": 3515698,
      "authorName": "diet1236364",
      "votes": 0,
      "postDate": "2026-08-22T09:23:36.857000",
      "content": "<p>single(5-fold) model 0.929 @ 224px</p>\n<p>LLM output labels only</p>"
    },
    {
      "id": 3520188,
      "authorName": "Saicharan Ramineni",
      "votes": 0,
      "postDate": "2026-09-02T23:36:04.997000",
      "content": "<p>What LLM did you use?</p>"
    },
    {
      "id": 3520505,
      "authorName": "diet1236364",
      "votes": 0,
      "postDate": "2026-09-03T15:07:19.433000",
      "content": "<p>im using \"astra\" lol</p>"
    },
    {
      "id": 3513661,
      "authorName": "Aic",
      "votes": 0,
      "postDate": "2026-08-17T09:09:42.770000",
      "content": "<p>I’d like to ask how long it usually takes to train a model using DINOv2. Also, roughly how much computational resources are required for this competition?</p>"
    },
    {
      "id": 3513671,
      "authorName": "Tim Krige",
      "votes": 2,
      "postDate": "2026-08-17T09:44:32.100000",
      "content": "<p>This is a hard question to answer. Dino has a variety of sizes, small to huge. Within a size there are layers, with larger sized having more layers. In some competitions there have been amazing results from the pretrained dino weights and only finetuned heads. Others require unfreezing a few of the last layers, and still other competitions benefit from a full backbone unfreeze. The more of the backbone you unfreeze the more VRAM and GPU speed  you need. I am finding strong correlations between number of layers unfrozen and score, but I only have limited vram so I am typically training the smallest dino's.</p>\n<p>I am on a 3070 and a 5070ti for most of my training. Models take a day or so to train 5 folds. </p>\n<p>A good technique on a tiny model is better than just throwing data at a large model and hoping compute can solve it, typically. You would also learn more like this.</p>\n<p>I hope this helps. </p>"
    },
    {
      "id": 3521115,
      "authorName": "",
      "votes": -3,
      "postDate": "2026-09-04T22:42:46.270000",
      "content": ""
    },
    {
      "id": 3521218,
      "authorName": "SpeedSci",
      "votes": 0,
      "postDate": "2026-09-05T08:15:52.937000",
      "content": "<p>I think different architectures can still complement each other, so I’m not fully convinced that ensembling is just averaging the same label noise. Right now, the most promising direction to me feels like: good labels + a strong single model + a few truly complementary models blended together.</p>"
    },
    {
      "id": 3521436,
      "authorName": "",
      "votes": -1,
      "postDate": "2026-09-05T22:06:23.003000",
      "content": ""
    },
    {
      "id": 3521219,
      "authorName": "SpeedSci",
      "votes": 0,
      "postDate": "2026-09-05T08:20:15.810000",
      "content": "<p>I’m a bit confused about something and wanted to ask: in my experiments, v10 labels + DINOv2 outperformed v9 labels + DINOv2 — but when I switched to ConvNeXt, v9 labels + ConvNeXt actually did better than v10 labels + ConvNeXt.</p>\n<p>It seems like different architectures have different \"ability\" to benefit from label improvements. Has anyone else run into something similar?</p>"
    },
    {
      "id": 3521431,
      "authorName": "",
      "votes": 0,
      "postDate": "2026-09-05T21:55:27.987000",
      "content": ""
    }
  ],
  "index": {
    "id": "735304",
    "title": "Best single-model score",
    "authorName": "",
    "commentCount": "87",
    "votes": "60",
    "postDate": "2026-08-14 23:38:09.839000"
  },
  "competition": "rsna-knee-abnormality-detection"
}