{
  "topic": {
    "id": 735154,
    "title": "Scaling the encoder bought us nothing (+0.0011)",
    "authorName": "stevenleehans",
    "commentCount": 3,
    "votes": 14,
    "postDate": "2026-08-14T08:05:01.209000"
  },
  "comments": [
    {
      "id": 3513356,
      "authorName": "Dung Duong (Tony) Huynh",
      "votes": 0,
      "postDate": "2026-08-16T07:58:31.900000",
      "content": "<p>Hi OG, thanks for sharing this! May I ask how you reduced each series to 9 slices? Did you randomly sample the slices, or use interpolation? Thank you!</p>"
    },
    {
      "id": 3513374,
      "authorName": "stevenleehans",
      "votes": 1,
      "postDate": "2026-08-16T08:59:25.113000",
      "content": "<p>Hi Tony — neither, actually. No random sampling and no interpolation; it's a deterministic pick of 9 real slices, in three steps.</p>\n<ol>\n<li><p>Sort the stack physically first. This one matters more than the sampling. The filenames in this dataset are SOP Instance UIDs, i.e. random — sorted filename order matches true anatomical order only about 5% of the time on this corpus. So before anything else I re-sort each series by projecting ImagePositionPatient onto the normal of ImageOrientationPatient, with fallbacks to SliceLocation → InstanceNumber → filename. That resolves by geometry on 100% of series here.</p></li>\n<li><p>Place 3 anchors across a window of the sorted stack, evenly spaced with linspace (not random). I clip the ends rather than use the full stack, since the outermost slices are mostly off-joint.</p></li>\n<li><p>Take the 3 physically adjacent slices around each anchor — so 3 × 3 = 9.</p></li>\n</ol>\n<p>The reason for the 3+3+3 layout rather than 9 evenly spread slices: each group of 3 becomes the R/G/B channels of one encoder input. Adjacent slices make that triplet a genuine ~10 mm depth neighbourhood at this dataset's slice gaps. Nine evenly spread slices would cover more of the joint but hand the encoder three views ~20 mm apart, which is a slab, not a 2.5D triplet.</p>\n<p>Two caveats worth knowing:</p>\n<p>Interpolation isn't used because MRI slice gaps here vary a lot; I'd rather feed real slices and fix scale in-plane instead (I crop to a fixed physical extent in mm before resizing, since field of view ranges over 71 distinct values with a median of 160 mm — so a fixed pixel size alone does not fix physical scale).\nOn my CV, 3 slices actually beat 9 by a margin ~4× my noise floor. But that result confounds two things — with one anchor the code takes the window centre, with three it spreads them to the window edges — so I can't yet tell you whether it's the slice count or the slice position doing the work. Don't take 9 as tuned.</p>"
    },
    {
      "id": 3513756,
      "authorName": "Dung Duong (Tony) Huynh",
      "votes": 0,
      "postDate": "2026-08-17T14:45:57.777000",
      "content": "<p>Thanks for sharing my friend!</p>"
    }
  ],
  "index": {
    "id": "735154",
    "title": "Scaling the encoder bought us nothing (+0.0011)",
    "authorName": "",
    "commentCount": "3",
    "votes": "14",
    "postDate": "2026-08-14 08:05:01.209000"
  },
  "competition": "rsna-knee-abnormality-detection"
}