{
  "id": 192284,
  "title": "Watch out for jpg compression",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/192284",
  "author_name": "Theo Viel",
  "post_date": "2020-10-20T21:55:47.821000",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>If you've been training with jpg images, keep in mind that they have been compressed.</p>\n<p>If I have an array of images loaded from dicoms, the following thing will happen :</p>\n<pre><code>cv2.imwrite(path + \".jpg\", img_dicom)\nimg_jpg = cv2.imread(path + \".jpg\")\n\n(img_jpg == img_dicom).all()\n&gt; False\n</code></pre>\n<p>Whereas the same code with a <code>.png</code> extension will return True. <br>\nI am not sure whether doing inference on uncompressed images will hurt results, but you can expect some variation. </p>\n<p>For instance, using a pretrained model for feature extraction :</p>\n<pre><code>(fts_jpg - fts_dcm).max()\n&gt; 0.2480\n(fts_jpg - fts_dcm).mean()\n&gt; 0.0028 \n</code></pre>\n<p>Which can be an issue when training second level models.</p>",
  "messages": [
    {
      "id": 1055515,
      "postDate": "2020-10-20T21:55:47.823Z",
      "content": "<p>If you've been training with jpg images, keep in mind that they have been compressed.</p>\n<p>If I have an array of images loaded from dicoms, the following thing will happen :</p>\n<pre><code>cv2.imwrite(path + \".jpg\", img_dicom)\nimg_jpg = cv2.imread(path + \".jpg\")\n\n(img_jpg == img_dicom).all()\n&gt; False\n</code></pre>\n<p>Whereas the same code with a <code>.png</code> extension will return True. <br>\nI am not sure whether doing inference on uncompressed images will hurt results, but you can expect some variation. </p>\n<p>For instance, using a pretrained model for feature extraction :</p>\n<pre><code>(fts_jpg - fts_dcm).max()\n&gt; 0.2480\n(fts_jpg - fts_dcm).mean()\n&gt; 0.0028 \n</code></pre>\n<p>Which can be an issue when training second level models.</p>",
      "rawMarkdown": "If you've been training with jpg images, keep in mind that they have been compressed.\n\nIf I have an array of images loaded from dicoms, the following thing will happen :\n\n```\ncv2.imwrite(path + \".jpg\", img_dicom)\nimg_jpg = cv2.imread(path + \".jpg\")\n\n(img_jpg == img_dicom).all()\n> False\n```\nWhereas the same code with a `.png` extension will return True. \nI am not sure whether doing inference on uncompressed images will hurt results, but you can expect some variation. \n\nFor instance, using a pretrained model for feature extraction :\n\n```\n(fts_jpg - fts_dcm).max()\n> 0.2480\n(fts_jpg - fts_dcm).mean()\n> 0.0028 \n```\n\nWhich can be an issue when training second level models.\n",
      "votes": 13
    },
    {
      "id": 1057452,
      "postDate": "2020-10-22T16:45:20.157Z",
      "content": "<p>I actually thought of this and I thought maybe adding a slight random noise as a data-augmentation method would solve the problem? I haven't tested this but I think it might work.</p>",
      "rawMarkdown": "I actually thought of this and I thought maybe adding a slight random noise as a data-augmentation method would solve the problem? I haven't tested this but I think it might work."
    },
    {
      "id": 1055938,
      "postDate": "2020-10-21T09:00:27.180Z",
      "content": "<p>Thanks for you finding!</p>",
      "rawMarkdown": "Thanks for you finding!"
    },
    {
      "id": 1055910,
      "postDate": "2020-10-21T08:08:09.723Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing\n"
    },
    {
      "id": 1055661,
      "postDate": "2020-10-21T01:44:40.087Z",
      "content": "<p>Thanks for pointing that out!</p>",
      "rawMarkdown": "Thanks for pointing that out!"
    }
  ],
  "comments": [
    {
      "id": 1057452,
      "author_name": "DarkCube",
      "author_url": "",
      "post_date": "2020-10-22T16:45:20.157000",
      "content": "<p>I actually thought of this and I thought maybe adding a slight random noise as a data-augmentation method would solve the problem? I haven't tested this but I think it might work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1055938,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2020-10-21T09:00:27.180000",
      "content": "<p>Thanks for you finding!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1055910,
      "author_name": "Maciej Gronczynski",
      "author_url": "",
      "post_date": "2020-10-21T08:08:09.723000",
      "content": "<p>thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1055661,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-10-21T01:44:40.087000",
      "content": "<p>Thanks for pointing that out!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1055515": "If you've been training with jpg images, keep in mind that they have been compressed.\n\nIf I have an array of images loaded from dicoms, the following thing will happen :\n\n```\ncv2.imwrite(path + \".jpg\", img_dicom)\nimg_jpg = cv2.imread(path + \".jpg\")\n\n(img_jpg == img_dicom).all()\n> False\n```\nWhereas the same code with a `.png` extension will return True. \nI am not sure whether doing inference on uncompressed images will hurt results, but you can expect some variation. \n\nFor instance, using a pretrained model for feature extraction :\n\n```\n(fts_jpg - fts_dcm).max()\n> 0.2480\n(fts_jpg - fts_dcm).mean()\n> 0.0028 \n```\n\nWhich can be an issue when training second level models.\n",
    "1057452": "I actually thought of this and I thought maybe adding a slight random noise as a data-augmentation method would solve the problem? I haven't tested this but I think it might work.",
    "1055938": "Thanks for you finding!",
    "1055910": "thanks for sharing\n",
    "1055661": "Thanks for pointing that out!"
  }
}