{
  "id": 420644,
  "title": "Many Many Duplicate Images [v2]. Contradicting Dup Labels in Training.",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/420644",
  "author_name": "SSS",
  "post_date": "2023-07-01T20:56:36.004000",
  "votes": 38,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I have recently posted few cross duplicated train-val examples <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419994\" target=\"_blank\">here</a>.</p>\n<p>Now I have found something more interesting. There are two duplicates from the train dataset with contradicting labels! <strong>It confuses my models!</strong> Look at the score. Here you go:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F45b908623417e42d1f25f7bcaf97f54f%2Ftrain1.png?generation=1688244592481433&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fbe70b581b3b99ec3285d00cbdfd64e36%2Ftrain2.png?generation=1688244600589058&amp;alt=media\" alt=\"\"></p>\n<p>I guess some of the competitors have already identified the problem, fixed or removed those examples from the training. I believe it might make sense to run similarity test for the entire dataset to identify such examples for dealing with them later on. </p>\n<p>The simpliest idea I can come up with right away is to find the mse between all images.</p>\n<pre><code> numpy  np\n\n ():\n     np.mean((imageA - imageB) ** )\n</code></pre>\n<h1><strong>Update</strong>:</h1>\n<h2>Here is a simple vectorized KNN.</h2>\n<pre><code>\n\n numpy  np\na = image4_loaded.reshape((image4_loaded.shape[], -))\n\n\n\n\na_squared = np.(np.square(a),axis=)\nmul = np.dot(a, a.T)\ndists = np.sqrt(a_squared[:,np.newaxis] + a_squared-*mul)\n\n\ndists = np.nan_to_num(dists)\nnp.fill_diagonal(dists, )\n\nK = \nclosest_idx = dists.argsort(axis=)[:, :K]\n</code></pre>\n<p>Examples with dist &lt; 25:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa02b5d8fc3b0108ea3857ea39a4038a3%2F3.png?generation=1688248348135714&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc23d92860e03a9a206cd70c9d8d49a79%2F4.png?generation=1688248361689141&amp;alt=media\" alt=\"\"></p>\n<p>You are welcome to share your techniques. I would greatly appreciate that personally, so the community, I believe.</p>",
  "messages": [
    {
      "id": 2326057,
      "postDate": "2023-07-01T20:56:36.003Z",
      "content": "<p>Hi,</p>\n<p>I have recently posted few cross duplicated train-val examples <a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419994\" target=\"_blank\">here</a>.</p>\n<p>Now I have found something more interesting. There are two duplicates from the train dataset with contradicting labels! <strong>It confuses my models!</strong> Look at the score. Here you go:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F45b908623417e42d1f25f7bcaf97f54f%2Ftrain1.png?generation=1688244592481433&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fbe70b581b3b99ec3285d00cbdfd64e36%2Ftrain2.png?generation=1688244600589058&amp;alt=media\" alt=\"\"></p>\n<p>I guess some of the competitors have already identified the problem, fixed or removed those examples from the training. I believe it might make sense to run similarity test for the entire dataset to identify such examples for dealing with them later on. </p>\n<p>The simpliest idea I can come up with right away is to find the mse between all images.</p>\n<pre><code> numpy  np\n\n ():\n     np.mean((imageA - imageB) ** )\n</code></pre>\n<h1><strong>Update</strong>:</h1>\n<h2>Here is a simple vectorized KNN.</h2>\n<pre><code>\n\n numpy  np\na = image4_loaded.reshape((image4_loaded.shape[], -))\n\n\n\n\na_squared = np.(np.square(a),axis=)\nmul = np.dot(a, a.T)\ndists = np.sqrt(a_squared[:,np.newaxis] + a_squared-*mul)\n\n\ndists = np.nan_to_num(dists)\nnp.fill_diagonal(dists, )\n\nK = \nclosest_idx = dists.argsort(axis=)[:, :K]\n</code></pre>\n<p>Examples with dist &lt; 25:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa02b5d8fc3b0108ea3857ea39a4038a3%2F3.png?generation=1688248348135714&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc23d92860e03a9a206cd70c9d8d49a79%2F4.png?generation=1688248361689141&amp;alt=media\" alt=\"\"></p>\n<p>You are welcome to share your techniques. I would greatly appreciate that personally, so the community, I believe.</p>",
      "rawMarkdown": "Hi,\n\nI have recently posted few cross duplicated train-val examples [here](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419994).\n\nNow I have found something more interesting. There are two duplicates from the train dataset with contradicting labels! **It confuses my models!** Look at the score. Here you go:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F45b908623417e42d1f25f7bcaf97f54f%2Ftrain1.png?generation=1688244592481433&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fbe70b581b3b99ec3285d00cbdfd64e36%2Ftrain2.png?generation=1688244600589058&alt=media)\n\nI guess some of the competitors have already identified the problem, fixed or removed those examples from the training. I believe it might make sense to run similarity test for the entire dataset to identify such examples for dealing with them later on. \n\n\nThe simpliest idea I can come up with right away is to find the mse between all images.\n```python\nimport numpy as np\n\ndef mean_squared_error(imageA, imageB):\n    return np.mean((imageA - imageB) ** 2)\n```\n\n#**Update**:\n\n## Here is a simple vectorized KNN.\n\n```python\n#KNN\n\nimport numpy as np\na = image4_loaded.reshape((image4_loaded.shape[0], -1))\n\n# Notice that (A - B)^2 = A^2 + B^2 - 2*A*B\n# Implement A^2 + B^2 - 2*A*B\n\na_squared = np.sum(np.square(a),axis=1)\nmul = np.dot(a, a.T)\ndists = np.sqrt(a_squared[:,np.newaxis] + a_squared-2*mul)\n\n# Fills nans with 0 and diagonal with 999.\ndists = np.nan_to_num(dists)\nnp.fill_diagonal(dists, 999)\n\nK = 3\nclosest_idx = dists.argsort(axis=1)[:, :K]\n```\n\nExamples with dist < 25:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa02b5d8fc3b0108ea3857ea39a4038a3%2F3.png?generation=1688248348135714&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc23d92860e03a9a206cd70c9d8d49a79%2F4.png?generation=1688248361689141&alt=media)\n\n\nYou are welcome to share your techniques. I would greatly appreciate that personally, so the community, I believe.",
      "votes": 37
    },
    {
      "id": 2326809,
      "postDate": "2023-07-02T12:25:53.300Z",
      "content": "<p>As I shared before, what I do is that I find geographic duplicates and look at the ones where the time difference is smaller than 1hour.  it does work well and yes removing them from the training set does improve the score. Nice catch on the contradicting masks tho !</p>",
      "rawMarkdown": "As I shared before, what I do is that I find geographic duplicates and look at the ones where the time difference is smaller than 1hour.  it does work well and yes removing them from the training set does improve the score. Nice catch on the contradicting masks tho !",
      "votes": 18,
      "replies": [
        {
          "id": 2327535,
          "postDate": "2023-07-03T03:10:47.297Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> , what do you mean by geographic duplicates? Which values from the metadata define Geographic. From what I see,  central_meridian can be used, but its just 1 value, are there other geo columns as well? Thanks. </p>",
          "rawMarkdown": "Hi @janmpia @sergiosaharovskiy , what do you mean by geographic duplicates? Which values from the metadata define Geographic. From what I see,  central_meridian can be used, but its just 1 value, are there other geo columns as well? Thanks. ",
          "votes": 2,
          "replies": [
            {
              "id": 2327584,
              "postDate": "2023-07-03T03:56:11.423Z",
              "content": "<p>Row_min, row_size, col_min, col_size are the values which define the geographic space of the frames.</p>",
              "rawMarkdown": "Row_min, row_size, col_min, col_size are the values which define the geographic space of the frames.",
              "votes": 2
            },
            {
              "id": 2327671,
              "postDate": "2023-07-03T05:24:25.043Z",
              "content": "<p>Thanks for the info! </p>",
              "rawMarkdown": "Thanks for the info! "
            },
            {
              "id": 2328999,
              "postDate": "2023-07-04T03:32:05.430Z",
              "content": "<p><a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> approximately how many images did you remove from training set?</p>",
              "rawMarkdown": "@janmpia approximately how many images did you remove from training set?"
            }
          ]
        },
        {
          "id": 2345905,
          "postDate": "2023-07-15T19:40:40.093Z",
          "content": "<p>For me the score only increases with 5 fold models not with single models and only on validation set (I keep validation set as holdout for 5 fold)</p>",
          "rawMarkdown": "For me the score only increases with 5 fold models not with single models and only on validation set (I keep validation set as holdout for 5 fold)"
        }
      ]
    },
    {
      "id": 2342469,
      "postDate": "2023-07-12T21:16:13.360Z",
      "content": "<p>If you want another fast way of finding duplicates, take a look at <a href=\"https://github.com/idealo/imagededup\" target=\"_blank\">https://github.com/idealo/imagededup</a></p>\n<p>The <code>Phash</code> method is very fast</p>",
      "rawMarkdown": "If you want another fast way of finding duplicates, take a look at https://github.com/idealo/imagededup\n\nThe `Phash` method is very fast",
      "votes": 13,
      "replies": [
        {
          "id": 2350865,
          "postDate": "2023-07-19T15:33:15.607Z",
          "content": "<p>I tried this method, <code>2383</code> samples out of 22385 are unique…</p>",
          "rawMarkdown": "I tried this method, `2383` samples out of 22385 are unique...",
          "replies": [
            {
              "id": 2353197,
              "postDate": "2023-07-21T14:00:35.390Z",
              "content": "<p>I am surprised because I used this same technique with parameters by default and I found 18287 unique samples. I haven't tested this new models on the LB yet but it looks like my training is more stable. My CV was constantly 0.015 above my LB score , probably due to potential data leakage between my folds.</p>",
              "rawMarkdown": "I am surprised because I used this same technique with parameters by default and I found 18287 unique samples. I haven't tested this new models on the LB yet but it looks like my training is more stable. My CV was constantly 0.015 above my LB score , probably due to potential data leakage between my folds.",
              "votes": 2
            },
            {
              "id": 2353299,
              "postDate": "2023-07-21T15:18:04.743Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/frlemarchand\" target=\"_blank\">@frlemarchand</a> , I found it was because I used the normalized image array. I use the default distance i.e. <code>max_distance_threshold: int = 10</code> and got 18313 unique samples after combing the train and valid.</p>",
              "rawMarkdown": "Thanks @frlemarchand , I found it was because I used the normalized image array. I use the default distance i.e. `max_distance_threshold: int = 10` and got 18313 unique samples after combing the train and valid.",
              "votes": 1
            },
            {
              "id": 2353851,
              "postDate": "2023-07-22T04:37:00.693Z",
              "content": "<p>I tried using it earlier, but it seems it doesn't work with npy files, right? So I need to convert npy to png/jpeg for this to work. </p>",
              "rawMarkdown": "I tried using it earlier, but it seems it doesn't work with npy files, right? So I need to convert npy to png/jpeg for this to work. "
            },
            {
              "id": 2354180,
              "postDate": "2023-07-22T09:24:23.837Z",
              "content": "<p>Yes, if you're using the .npy files from the <a href=\"https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color\" target=\"_blank\">Contrails Images (Ash Color) dataset</a>, you can turn them into .jpg by extracting the image with <code>img = np.load(path+np_file)[:,:,:-1]</code> and save them with <code>matplotlib.image.imsave()</code>. </p>",
              "rawMarkdown": "Yes, if you're using the .npy files from the [Contrails Images (Ash Color) dataset](https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color), you can turn them into .jpg by extracting the image with `img = np.load(path+np_file)[:,:,:-1]` and save them with `matplotlib.image.imsave()`. "
            },
            {
              "id": 2354355,
              "postDate": "2023-07-22T11:23:02.600Z",
              "content": "<p>You can use a custom model, <a href=\"https://github.com/idealo/imagededup/blob/master/examples/use_custom_model.ipynb\" target=\"_blank\">https://github.com/idealo/imagededup/blob/master/examples/use_custom_model.ipynb</a>. Maybe by changing the forward pass to accept a batch of npy files and convert those files from npy to image there, avoid the need of saving those npy files to jpg as a dataset to later use. Or you can use another library, like Imagehash, here is a good implementation, <a href=\"https://www.kaggle.com/code/schulta/petfinder-identify-duplicates-and-share-findings/notebook\" target=\"_blank\">https://www.kaggle.com/code/schulta/petfinder-identify-duplicates-and-share-findings/notebook</a>.</p>",
              "rawMarkdown": "You can use a custom model, https://github.com/idealo/imagededup/blob/master/examples/use_custom_model.ipynb. Maybe by changing the forward pass to accept a batch of npy files and convert those files from npy to image there, avoid the need of saving those npy files to jpg as a dataset to later use. Or you can use another library, like Imagehash, here is a good implementation, https://www.kaggle.com/code/schulta/petfinder-identify-duplicates-and-share-findings/notebook."
            }
          ]
        },
        {
          "id": 2375728,
          "postDate": "2023-08-05T21:07:35.027Z",
          "content": "<p>Thank you for sharing this! I agree that Phash method is significantly faster. Some instances tend to have a lot of duplicates:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1881746%2F6554820174b39b99f37126d3bfcad5e8%2FScreen%20Shot%202023-08-05%20at%202.04.40%20PM.png?generation=1691269691355600&amp;alt=media\" alt=\"\"></p>\n<p>For the training data, there is a drop of 2k records:<br>\nBefore dedup: 20529<br>\nAfter dedup: 18180</p>\n<p>For the validation data, there is drop of about 40 records:<br>\nBefore dedup: 1856<br>\nAfter dedup: 1811</p>\n<p>I conducted the dedup across the training and validation datasets, and I have gone ahead and created a dataset with the updated train_df.csv and valid_df.csv files.<br>\nHere is the link to the dataset - <a href=\"https://www.kaggle.com/datasets/darshann25/contrails-images-ash-color-deduped\" target=\"_blank\">https://www.kaggle.com/datasets/darshann25/contrails-images-ash-color-deduped</a></p>\n<p>While using it, ensure to only update the train and test paths, and continue reading the npy files from the original ash color dataset. See follows:</p>\n<pre><code>contrails = \ntrain_path = \nvalid_path =  \n</code></pre>",
          "rawMarkdown": "Thank you for sharing this! I agree that Phash method is significantly faster. Some instances tend to have a lot of duplicates:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1881746%2F6554820174b39b99f37126d3bfcad5e8%2FScreen%20Shot%202023-08-05%20at%202.04.40%20PM.png?generation=1691269691355600&alt=media)\n\nFor the training data, there is a drop of 2k records:\nBefore dedup: 20529\nAfter dedup: 18180\n\nFor the validation data, there is drop of about 40 records:\nBefore dedup: 1856\nAfter dedup: 1811\n\nI conducted the dedup across the training and validation datasets, and I have gone ahead and created a dataset with the updated train_df.csv and valid_df.csv files.\nHere is the link to the dataset - https://www.kaggle.com/datasets/darshann25/contrails-images-ash-color-deduped\n\nWhile using it, ensure to only update the train and test paths, and continue reading the npy files from the original ash color dataset. See follows:\n\n```python\ncontrails = '/kaggle/input/contrails-images-ash-color/contrails'\ntrain_path = '/kaggle/input/contrails-images-ash-color-deduped/dedup_train_df.csv'\nvalid_path = '/kaggle/input/contrails-images-ash-color-deduped/dedup_valid_df.csv' \n```\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2353235,
      "postDate": "2023-07-21T14:19:00.653Z",
      "content": "<p>I use Euclidean Distance between 2 images to detect the duplicates:</p>\n<pre><code> numpy  np\n\n ():\n     np.sqrt(np.((imageA - imageB) ** ))\n</code></pre>\n<p>'6002874852032119389.npy', '2520175669696585553.npy' have very similar images with distance=30.89</p>\n<p>the number of duplicates v.s. maximum distance is as the attached image. I need to remove half of the images to resolve this issue😂<br>\n（I'm not sure why I can't paste image directly in the textarea）</p>",
      "rawMarkdown": "I use Euclidean Distance between 2 images to detect the duplicates:\n\n```Python\nimport numpy as np\n\ndef distance(imageA, imageB):\n    return np.sqrt(np.sum((imageA - imageB) ** 2))\n```\n'6002874852032119389.npy', '2520175669696585553.npy' have very similar images with distance=30.89\n\nthe number of duplicates v.s. maximum distance is as the attached image. I need to remove half of the images to resolve this issue😂\n（I'm not sure why I can't paste image directly in the textarea）\n",
      "votes": 1
    },
    {
      "id": 2342360,
      "postDate": "2023-07-12T17:57:04.210Z",
      "content": "<p>I am struggling to understand the 'Vectorized KNN' code. Could you please refer me to any articles that can help me better understand it?</p>",
      "rawMarkdown": "I am struggling to understand the 'Vectorized KNN' code. Could you please refer me to any articles that can help me better understand it?",
      "votes": 1,
      "replies": [
        {
          "id": 2342421,
          "postDate": "2023-07-12T19:32:06.617Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sahilsg\" target=\"_blank\">@sahilsg</a>, you can definitely google it.  But I just expanded the equation <strong>$${(a - b)^2} = {a^2} - {2 * a * b} + {b^2}$$</strong><br>\n which was L2 distance I used in KNN. Then I applied matrix operations with numpy.</p>",
          "rawMarkdown": "Hi @sahilsg, you can definitely google it.  But I just expanded the equation **$${(a - b)^2} = {a^2} - {2 * a * b} + {b^2}$$**\n which was L2 distance I used in KNN. Then I applied matrix operations with numpy.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2326809,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-07-02T12:25:53.300000",
      "content": "<p>As I shared before, what I do is that I find geographic duplicates and look at the ones where the time difference is smaller than 1hour.  it does work well and yes removing them from the training set does improve the score. Nice catch on the contradicting masks tho !</p>",
      "votes": 18,
      "replies": [
        {
          "id": 2327535,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2023-07-03T03:10:47.297000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> <a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> , what do you mean by geographic duplicates? Which values from the metadata define Geographic. From what I see,  central_meridian can be used, but its just 1 value, are there other geo columns as well? Thanks. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2327584,
              "author_name": "Ari",
              "author_url": "",
              "post_date": "2023-07-03T03:56:11.423000",
              "content": "<p>Row_min, row_size, col_min, col_size are the values which define the geographic space of the frames.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2327671,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2023-07-03T05:24:25.043000",
              "content": "<p>Thanks for the info! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2328999,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2023-07-04T03:32:05.430000",
              "content": "<p><a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> approximately how many images did you remove from training set?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2345905,
          "author_name": "Ari",
          "author_url": "",
          "post_date": "2023-07-15T19:40:40.093000",
          "content": "<p>For me the score only increases with 5 fold models not with single models and only on validation set (I keep validation set as holdout for 5 fold)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2342469,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2023-07-12T21:16:13.360000",
      "content": "<p>If you want another fast way of finding duplicates, take a look at <a href=\"https://github.com/idealo/imagededup\" target=\"_blank\">https://github.com/idealo/imagededup</a></p>\n<p>The <code>Phash</code> method is very fast</p>",
      "votes": 13,
      "replies": [
        {
          "id": 2350865,
          "author_name": "william.wu",
          "author_url": "",
          "post_date": "2023-07-19T15:33:15.607000",
          "content": "<p>I tried this method, <code>2383</code> samples out of 22385 are unique…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2353197,
              "author_name": "Francois Lemarchand",
              "author_url": "",
              "post_date": "2023-07-21T14:00:35.390000",
              "content": "<p>I am surprised because I used this same technique with parameters by default and I found 18287 unique samples. I haven't tested this new models on the LB yet but it looks like my training is more stable. My CV was constantly 0.015 above my LB score , probably due to potential data leakage between my folds.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2353299,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-07-21T15:18:04.743000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/frlemarchand\" target=\"_blank\">@frlemarchand</a> , I found it was because I used the normalized image array. I use the default distance i.e. <code>max_distance_threshold: int = 10</code> and got 18313 unique samples after combing the train and valid.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2353851,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2023-07-22T04:37:00.693000",
              "content": "<p>I tried using it earlier, but it seems it doesn't work with npy files, right? So I need to convert npy to png/jpeg for this to work. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2354180,
              "author_name": "Francois Lemarchand",
              "author_url": "",
              "post_date": "2023-07-22T09:24:23.837000",
              "content": "<p>Yes, if you're using the .npy files from the <a href=\"https://www.kaggle.com/datasets/shashwatraman/contrails-images-ash-color\" target=\"_blank\">Contrails Images (Ash Color) dataset</a>, you can turn them into .jpg by extracting the image with <code>img = np.load(path+np_file)[:,:,:-1]</code> and save them with <code>matplotlib.image.imsave()</code>. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2354355,
              "author_name": "Maximiliano Diaz Battan",
              "author_url": "",
              "post_date": "2023-07-22T11:23:02.600000",
              "content": "<p>You can use a custom model, <a href=\"https://github.com/idealo/imagededup/blob/master/examples/use_custom_model.ipynb\" target=\"_blank\">https://github.com/idealo/imagededup/blob/master/examples/use_custom_model.ipynb</a>. Maybe by changing the forward pass to accept a batch of npy files and convert those files from npy to image there, avoid the need of saving those npy files to jpg as a dataset to later use. Or you can use another library, like Imagehash, here is a good implementation, <a href=\"https://www.kaggle.com/code/schulta/petfinder-identify-duplicates-and-share-findings/notebook\" target=\"_blank\">https://www.kaggle.com/code/schulta/petfinder-identify-duplicates-and-share-findings/notebook</a>.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2375728,
          "author_name": "Darshan Patel",
          "author_url": "",
          "post_date": "2023-08-05T21:07:35.027000",
          "content": "<p>Thank you for sharing this! I agree that Phash method is significantly faster. Some instances tend to have a lot of duplicates:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1881746%2F6554820174b39b99f37126d3bfcad5e8%2FScreen%20Shot%202023-08-05%20at%202.04.40%20PM.png?generation=1691269691355600&amp;alt=media\" alt=\"\"></p>\n<p>For the training data, there is a drop of 2k records:<br>\nBefore dedup: 20529<br>\nAfter dedup: 18180</p>\n<p>For the validation data, there is drop of about 40 records:<br>\nBefore dedup: 1856<br>\nAfter dedup: 1811</p>\n<p>I conducted the dedup across the training and validation datasets, and I have gone ahead and created a dataset with the updated train_df.csv and valid_df.csv files.<br>\nHere is the link to the dataset - <a href=\"https://www.kaggle.com/datasets/darshann25/contrails-images-ash-color-deduped\" target=\"_blank\">https://www.kaggle.com/datasets/darshann25/contrails-images-ash-color-deduped</a></p>\n<p>While using it, ensure to only update the train and test paths, and continue reading the npy files from the original ash color dataset. See follows:</p>\n<pre><code>contrails = \ntrain_path = \nvalid_path =  \n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2353235,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2023-07-21T14:19:00.653000",
      "content": "<p>I use Euclidean Distance between 2 images to detect the duplicates:</p>\n<pre><code> numpy  np\n\n ():\n     np.sqrt(np.((imageA - imageB) ** ))\n</code></pre>\n<p>'6002874852032119389.npy', '2520175669696585553.npy' have very similar images with distance=30.89</p>\n<p>the number of duplicates v.s. maximum distance is as the attached image. I need to remove half of the images to resolve this issue😂<br>\n（I'm not sure why I can't paste image directly in the textarea）</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2342360,
      "author_name": "sahilsg",
      "author_url": "",
      "post_date": "2023-07-12T17:57:04.210000",
      "content": "<p>I am struggling to understand the 'Vectorized KNN' code. Could you please refer me to any articles that can help me better understand it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2342421,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2023-07-12T19:32:06.617000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sahilsg\" target=\"_blank\">@sahilsg</a>, you can definitely google it.  But I just expanded the equation <strong>$${(a - b)^2} = {a^2} - {2 * a * b} + {b^2}$$</strong><br>\n which was L2 distance I used in KNN. Then I applied matrix operations with numpy.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2326057": "Hi,\n\nI have recently posted few cross duplicated train-val examples [here](https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419994).\n\nNow I have found something more interesting. There are two duplicates from the train dataset with contradicting labels! **It confuses my models!** Look at the score. Here you go:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F45b908623417e42d1f25f7bcaf97f54f%2Ftrain1.png?generation=1688244592481433&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fbe70b581b3b99ec3285d00cbdfd64e36%2Ftrain2.png?generation=1688244600589058&alt=media)\n\nI guess some of the competitors have already identified the problem, fixed or removed those examples from the training. I believe it might make sense to run similarity test for the entire dataset to identify such examples for dealing with them later on. \n\n\nThe simpliest idea I can come up with right away is to find the mse between all images.\n```python\nimport numpy as np\n\ndef mean_squared_error(imageA, imageB):\n    return np.mean((imageA - imageB) ** 2)\n```\n\n#**Update**:\n\n## Here is a simple vectorized KNN.\n\n```python\n#KNN\n\nimport numpy as np\na = image4_loaded.reshape((image4_loaded.shape[0], -1))\n\n# Notice that (A - B)^2 = A^2 + B^2 - 2*A*B\n# Implement A^2 + B^2 - 2*A*B\n\na_squared = np.sum(np.square(a),axis=1)\nmul = np.dot(a, a.T)\ndists = np.sqrt(a_squared[:,np.newaxis] + a_squared-2*mul)\n\n# Fills nans with 0 and diagonal with 999.\ndists = np.nan_to_num(dists)\nnp.fill_diagonal(dists, 999)\n\nK = 3\nclosest_idx = dists.argsort(axis=1)[:, :K]\n```\n\nExamples with dist < 25:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fa02b5d8fc3b0108ea3857ea39a4038a3%2F3.png?generation=1688248348135714&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fc23d92860e03a9a206cd70c9d8d49a79%2F4.png?generation=1688248361689141&alt=media)\n\n\nYou are welcome to share your techniques. I would greatly appreciate that personally, so the community, I believe.",
    "2326809": "As I shared before, what I do is that I find geographic duplicates and look at the ones where the time difference is smaller than 1hour.  it does work well and yes removing them from the training set does improve the score. Nice catch on the contradicting masks tho !",
    "2342469": "If you want another fast way of finding duplicates, take a look at https://github.com/idealo/imagededup\n\nThe `Phash` method is very fast",
    "2353235": "I use Euclidean Distance between 2 images to detect the duplicates:\n\n```Python\nimport numpy as np\n\ndef distance(imageA, imageB):\n    return np.sqrt(np.sum((imageA - imageB) ** 2))\n```\n'6002874852032119389.npy', '2520175669696585553.npy' have very similar images with distance=30.89\n\nthe number of duplicates v.s. maximum distance is as the attached image. I need to remove half of the images to resolve this issue😂\n（I'm not sure why I can't paste image directly in the textarea）\n",
    "2342360": "I am struggling to understand the 'Vectorized KNN' code. Could you please refer me to any articles that can help me better understand it?"
  }
}