{
  "id": 145618,
  "title": "Simple White Background Trim Function",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/145618",
  "author_name": "Shangqiu Li",
  "post_date": "2020-04-23T21:33:22.001000",
  "votes": 24,
  "comment_count": 15,
  "views": 0,
  "content": "<p>```\nfrom PIL import Image, ImageChops</p>\n\n<p>def trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n```\nAbove code is a simple function for trimming white background(I found <a href=\"https://stackoverflow.com/questions/10615901/trim-whitespace-using-pil\">here</a>) It can reduce fair amount of image size without cutting off any useful information for some samples.</p>\n\n<p>Before:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fe0ef7274e3234ae040f835b7318768c0%2FScreen%20Shot%202020-04-23%20at%202.32.19%20PM.png?generation=1587677563643878&amp;alt=media\" alt=\"\"></p>\n\n<p>After:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ff68e2c7ab8d295005d1f53e96ca1978c%2FScreen%20Shot%202020-04-23%20at%202.32.26%20PM.png?generation=1587677578735304&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 818418,
      "postDate": "2020-04-23T21:33:22Z",
      "content": "<p>```\nfrom PIL import Image, ImageChops</p>\n\n<p>def trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n```\nAbove code is a simple function for trimming white background(I found <a href=\"https://stackoverflow.com/questions/10615901/trim-whitespace-using-pil\">here</a>) It can reduce fair amount of image size without cutting off any useful information for some samples.</p>\n\n<p>Before:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fe0ef7274e3234ae040f835b7318768c0%2FScreen%20Shot%202020-04-23%20at%202.32.19%20PM.png?generation=1587677563643878&amp;alt=media\" alt=\"\"></p>\n\n<p>After:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ff68e2c7ab8d295005d1f53e96ca1978c%2FScreen%20Shot%202020-04-23%20at%202.32.26%20PM.png?generation=1587677578735304&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "```\nfrom PIL import Image, ImageChops\n\ndef trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n```\nAbove code is a simple function for trimming white background(I found [here](https://stackoverflow.com/questions/10615901/trim-whitespace-using-pil)) It can reduce fair amount of image size without cutting off any useful information for some samples.\n\nBefore:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fe0ef7274e3234ae040f835b7318768c0%2FScreen%20Shot%202020-04-23%20at%202.32.19%20PM.png?generation=1587677563643878&amp;alt=media)\n\nAfter:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ff68e2c7ab8d295005d1f53e96ca1978c%2FScreen%20Shot%202020-04-23%20at%202.32.26%20PM.png?generation=1587677578735304&amp;alt=media)\n",
      "votes": 24
    },
    {
      "id": 820012,
      "postDate": "2020-04-25T05:00:11.050Z",
      "content": "<p>Another pure numpy solution:\n<code>\ndef remove_border(image, mask=None):\n    borders = np.where(image.sum(2) != 3*255)\n    x_min = np.min(borders[0])\n    x_max = np.max(borders[0]) + 1\n    y_min = np.min(borders[1])\n    y_max = np.max(borders[1]) + 1\n    image = image[x_min:x_max, y_min:y_max]\n    if mask is not None:\n        mask = mask[x_min:x_max, y_min:y_max]\n        return image, mask\n    return image\n</code></p>",
      "rawMarkdown": "Another pure numpy solution:\n```\ndef remove_border(image, mask=None):\n    borders = np.where(image.sum(2) != 3*255)\n    x_min = np.min(borders[0])\n    x_max = np.max(borders[0]) + 1\n    y_min = np.min(borders[1])\n    y_max = np.max(borders[1]) + 1\n    image = image[x_min:x_max, y_min:y_max]\n    if mask is not None:\n        mask = mask[x_min:x_max, y_min:y_max]\n        return image, mask\n    return image\n```",
      "votes": 4,
      "replies": [
        {
          "id": 820707,
          "postDate": "2020-04-25T16:44:54.930Z",
          "content": "<p>Neat! I actually wrote a little more complex numpy one that is faster. 50ms vs 2s; 16s vs 25s(this one's super slow because the picture I tested is very big(30000,30000).</p>",
          "rawMarkdown": "Neat! I actually wrote a little more complex numpy one that is faster. 50ms vs 2s; 16s vs 25s(this one's super slow because the picture I tested is very big(30000,30000)."
        },
        {
          "id": 820949,
          "postDate": "2020-04-25T20:08:20.597Z",
          "content": "<p>This version is around 2x faster for me:\n<code>\ndef crop_white(image: np.ndarray) -&gt; np.ndarray:\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) != 255).nonzero()\n    xs, = (image.min(0).min(1) != 255).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]\n</code></p>",
          "rawMarkdown": "This version is around 2x faster for me:\n```\ndef crop_white(image: np.ndarray) -&gt; np.ndarray:\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) != 255).nonzero()\n    xs, = (image.min(0).min(1) != 255).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]\n```",
          "votes": 8
        },
        {
          "id": 821155,
          "postDate": "2020-04-26T00:27:19.573Z",
          "content": "<p>Thanks for this! However im getting a tensor size error when adding this function in the inference, have you managed to make it work? </p>",
          "rawMarkdown": "Thanks for this! However im getting a tensor size error when adding this function in the inference, have you managed to make it work? "
        },
        {
          "id": 822247,
          "postDate": "2020-04-26T19:27:23.880Z",
          "content": "<p>Above code is in numpy, which will not work directly because pytorch assume what you feed in is tensor. You will need to transform it into a tensor before feeding it into your network.</p>",
          "rawMarkdown": "Above code is in numpy, which will not work directly because pytorch assume what you feed in is tensor. You will need to transform it into a tensor before feeding it into your network."
        }
      ]
    },
    {
      "id": 865541,
      "postDate": "2020-05-28T17:30:53.367Z",
      "content": "<p>This is a very useful pre-processing technique.</p>",
      "rawMarkdown": "This is a very useful pre-processing technique."
    },
    {
      "id": 819658,
      "postDate": "2020-04-24T18:51:36.473Z",
      "content": "<p>If you code your custom one that would be 100's of time faster, after that you will have to add some code to rearrange your pic but at last, you will still get it 10's of times faster (most likely under a second), and on my pc that algorithm takes 21 seconds so, you will want it faster.</p>",
      "rawMarkdown": "If you code your custom one that would be 100's of time faster, after that you will have to add some code to rearrange your pic but at last, you will still get it 10's of times faster (most likely under a second), and on my pc that algorithm takes 21 seconds so, you will want it faster."
    },
    {
      "id": 819563,
      "postDate": "2020-04-24T17:10:24.837Z",
      "content": "<p>For Histopathology problems like these, the industry standard method to differentiate tissue area from non-tissue area is Otsu filtering.</p>",
      "rawMarkdown": "For Histopathology problems like these, the industry standard method to differentiate tissue area from non-tissue area is Otsu filtering.",
      "replies": [
        {
          "id": 819665,
          "postDate": "2020-04-24T18:57:07.650Z",
          "content": "<p>I'm not an expert of medical imaging, but for this, it seems like the background is simply white, so I just thought this would be good, instead of using a threshold.</p>",
          "rawMarkdown": "I'm not an expert of medical imaging, but for this, it seems like the background is simply white, so I just thought this would be good, instead of using a threshold."
        },
        {
          "id": 819688,
          "postDate": "2020-04-24T19:11:09.257Z",
          "content": "<p>your logic is correct though. A simple thing would be to create slices of the big images <a href=\"https://www.kaggle.com/imrandude/wsi-extract-patches-pytorch\">as shown here</a> (shameless plug 😄) and filter out all the pixels with value 255.</p>",
          "rawMarkdown": "your logic is correct though. A simple thing would be to create slices of the big images [as shown here](https://www.kaggle.com/imrandude/wsi-extract-patches-pytorch) (shameless plug 😄) and filter out all the pixels with value 255."
        }
      ]
    },
    {
      "id": 818716,
      "postDate": "2020-04-24T04:34:31.717Z",
      "content": "<p>Thank you for this, however I just tried your code and it seems to give me an error when training, its like its not cropping some images properly. I tried to fix it but no success :(</p>",
      "rawMarkdown": "Thank you for this, however I just tried your code and it seems to give me an error when training, its like its not cropping some images properly. I tried to fix it but no success :(",
      "replies": [
        {
          "id": 819415,
          "postDate": "2020-04-24T15:24:16.133Z",
          "content": "<p>hmm.... That's very interesting. It might not return any stuff if there's no white boarder detected. Try this:\n```\nfrom PIL import Image, ImageChops</p>\n\n<p>def trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n    else:\n        return im\n``` </p>\n\n<p>Also, make sure the image you pass to that function is a PIL image.</p>\n\n<p>If above doesn't work, do you mind posting the traceback and the problem image?</p>",
          "rawMarkdown": "hmm.... That's very interesting. It might not return any stuff if there's no white boarder detected. Try this:\n```\nfrom PIL import Image, ImageChops\n\ndef trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n    else:\n        return im\n``` \n\nAlso, make sure the image you pass to that function is a PIL image.\n\nIf above doesn't work, do you mind posting the traceback and the problem image?",
          "votes": 1
        },
        {
          "id": 819711,
          "postDate": "2020-04-24T19:38:45.423Z",
          "content": "<p>Thanks! This worked, i guess some images dont have a white background or something. My error was a tuple out of index</p>",
          "rawMarkdown": "Thanks! This worked, i guess some images dont have a white background or something. My error was a tuple out of index"
        }
      ]
    },
    {
      "id": 818420,
      "postDate": "2020-04-23T21:38:18.733Z",
      "content": "<p>Great, that's just what I needed, thanks! \nI think a file containing for each picture the bboxes would be interesting.</p>",
      "rawMarkdown": "Great, that's just what I needed, thanks! \nI think a file containing for each picture the bboxes would be interesting.",
      "replies": [
        {
          "id": 818479,
          "postDate": "2020-04-23T23:17:14.013Z",
          "content": "<p>I was thinking of such file would be helpful(no need to compute every batch). I will start working on such file once I settled down with a good baseline.</p>",
          "rawMarkdown": "I was thinking of such file would be helpful(no need to compute every batch). I will start working on such file once I settled down with a good baseline.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 820012,
      "author_name": "hawkey",
      "author_url": "",
      "post_date": "2020-04-25T05:00:11.050000",
      "content": "<p>Another pure numpy solution:\n<code>\ndef remove_border(image, mask=None):\n    borders = np.where(image.sum(2) != 3*255)\n    x_min = np.min(borders[0])\n    x_max = np.max(borders[0]) + 1\n    y_min = np.min(borders[1])\n    y_max = np.max(borders[1]) + 1\n    image = image[x_min:x_max, y_min:y_max]\n    if mask is not None:\n        mask = mask[x_min:x_max, y_min:y_max]\n        return image, mask\n    return image\n</code></p>",
      "votes": 4,
      "replies": [
        {
          "id": 820707,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-25T16:44:54.930000",
          "content": "<p>Neat! I actually wrote a little more complex numpy one that is faster. 50ms vs 2s; 16s vs 25s(this one's super slow because the picture I tested is very big(30000,30000).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 820949,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-04-25T20:08:20.597000",
          "content": "<p>This version is around 2x faster for me:\n<code>\ndef crop_white(image: np.ndarray) -&gt; np.ndarray:\n    assert image.shape[2] == 3\n    assert image.dtype == np.uint8\n    ys, = (image.min((1, 2)) != 255).nonzero()\n    xs, = (image.min(0).min(1) != 255).nonzero()\n    if len(xs) == 0 or len(ys) == 0:\n        return image\n    return image[ys.min():ys.max() + 1, xs.min():xs.max() + 1]\n</code></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 821155,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-04-26T00:27:19.573000",
          "content": "<p>Thanks for this! However im getting a tensor size error when adding this function in the inference, have you managed to make it work? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 822247,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-26T19:27:23.880000",
          "content": "<p>Above code is in numpy, which will not work directly because pytorch assume what you feed in is tensor. You will need to transform it into a tensor before feeding it into your network.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 865541,
      "author_name": "Kurian Benoy",
      "author_url": "",
      "post_date": "2020-05-28T17:30:53.367000",
      "content": "<p>This is a very useful pre-processing technique.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819658,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-04-24T18:51:36.473000",
      "content": "<p>If you code your custom one that would be 100's of time faster, after that you will have to add some code to rearrange your pic but at last, you will still get it 10's of times faster (most likely under a second), and on my pc that algorithm takes 21 seconds so, you will want it faster.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819563,
      "author_name": "Sheik Mohamed Imran",
      "author_url": "",
      "post_date": "2020-04-24T17:10:24.837000",
      "content": "<p>For Histopathology problems like these, the industry standard method to differentiate tissue area from non-tissue area is Otsu filtering.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 819665,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-24T18:57:07.650000",
          "content": "<p>I'm not an expert of medical imaging, but for this, it seems like the background is simply white, so I just thought this would be good, instead of using a threshold.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 819688,
          "author_name": "Sheik Mohamed Imran",
          "author_url": "",
          "post_date": "2020-04-24T19:11:09.257000",
          "content": "<p>your logic is correct though. A simple thing would be to create slices of the big images <a href=\"https://www.kaggle.com/imrandude/wsi-extract-patches-pytorch\">as shown here</a> (shameless plug 😄) and filter out all the pixels with value 255.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 818716,
      "author_name": "Yann Majewski",
      "author_url": "",
      "post_date": "2020-04-24T04:34:31.717000",
      "content": "<p>Thank you for this, however I just tried your code and it seems to give me an error when training, its like its not cropping some images properly. I tried to fix it but no success :(</p>",
      "votes": 0,
      "replies": [
        {
          "id": 819415,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-24T15:24:16.133000",
          "content": "<p>hmm.... That's very interesting. It might not return any stuff if there's no white boarder detected. Try this:\n```\nfrom PIL import Image, ImageChops</p>\n\n<p>def trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n    else:\n        return im\n``` </p>\n\n<p>Also, make sure the image you pass to that function is a PIL image.</p>\n\n<p>If above doesn't work, do you mind posting the traceback and the problem image?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 819711,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2020-04-24T19:38:45.423000",
          "content": "<p>Thanks! This worked, i guess some images dont have a white background or something. My error was a tuple out of index</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 818420,
      "author_name": "Adil Zouitine",
      "author_url": "",
      "post_date": "2020-04-23T21:38:18.733000",
      "content": "<p>Great, that's just what I needed, thanks! \nI think a file containing for each picture the bboxes would be interesting.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 818479,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-23T23:17:14.013000",
          "content": "<p>I was thinking of such file would be helpful(no need to compute every batch). I will start working on such file once I settled down with a good baseline.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "818418": "```\nfrom PIL import Image, ImageChops\n\ndef trim(im):\n    bg = Image.new(im.mode, im.size, im.getpixel((0,0)))\n    diff = ImageChops.difference(im, bg)\n    diff = ImageChops.add(diff, diff, 2.0, -100)\n    bbox = diff.getbbox()\n    if bbox:\n        return im.crop(bbox)\n```\nAbove code is a simple function for trimming white background(I found [here](https://stackoverflow.com/questions/10615901/trim-whitespace-using-pil)) It can reduce fair amount of image size without cutting off any useful information for some samples.\n\nBefore:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fe0ef7274e3234ae040f835b7318768c0%2FScreen%20Shot%202020-04-23%20at%202.32.19%20PM.png?generation=1587677563643878&amp;alt=media)\n\nAfter:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ff68e2c7ab8d295005d1f53e96ca1978c%2FScreen%20Shot%202020-04-23%20at%202.32.26%20PM.png?generation=1587677578735304&amp;alt=media)\n",
    "820012": "Another pure numpy solution:\n```\ndef remove_border(image, mask=None):\n    borders = np.where(image.sum(2) != 3*255)\n    x_min = np.min(borders[0])\n    x_max = np.max(borders[0]) + 1\n    y_min = np.min(borders[1])\n    y_max = np.max(borders[1]) + 1\n    image = image[x_min:x_max, y_min:y_max]\n    if mask is not None:\n        mask = mask[x_min:x_max, y_min:y_max]\n        return image, mask\n    return image\n```",
    "865541": "This is a very useful pre-processing technique.",
    "819658": "If you code your custom one that would be 100's of time faster, after that you will have to add some code to rearrange your pic but at last, you will still get it 10's of times faster (most likely under a second), and on my pc that algorithm takes 21 seconds so, you will want it faster.",
    "819563": "For Histopathology problems like these, the industry standard method to differentiate tissue area from non-tissue area is Otsu filtering.",
    "818716": "Thank you for this, however I just tried your code and it seems to give me an error when training, its like its not cropping some images properly. I tried to fix it but no success :(",
    "818420": "Great, that's just what I needed, thanks! \nI think a file containing for each picture the bboxes would be interesting."
  }
}