{
  "id": 419816,
  "title": "Are 3D models worth?",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419816",
  "author_name": "MPWARE",
  "post_date": "2023-06-27T18:06:55.530000",
  "votes": 27,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I would like to open the discussion on 3D models. According to the Contrails paper, their best result is for a 3D backbone based on I3D. They keep time dimension in 3D convolutions and pick the frame with the annotated masks to push in the decoder and make a prediction.</p>\n<p>Top winners of the just completed <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection\" target=\"_blank\">vesuvius</a> competition also use some models with 3D backbones. The problem is slightly different but the idea is the same.</p>\n<p>I've spent a few days on testing a lot of 3D backbones coming from the great <a href=\"https://github.com/open-mmlab/mmaction2\" target=\"_blank\">MMAction2</a> implementations such as:</p>\n<ul>\n<li>Vanilla Resnet3D </li>\n<li>Resnet3D - irCSN 50/152</li>\n<li>Resnet(2+1)D</li>\n<li>MViTv2<br>\nand more …</li>\n</ul>\n<p>It works but never better than a basic 2D backbone as we can see in the NoteBooks/Code section.<br>\nWhatever pooling or extracting the correct slice to feed the decoder, the results look not matching the paper. </p>\n<p>Anyone with similar disapointing results on 3D modeling?</p>",
  "messages": [
    {
      "id": 2320356,
      "postDate": "2023-06-27T18:06:55.530Z",
      "content": "<p>Hi,</p>\n<p>I would like to open the discussion on 3D models. According to the Contrails paper, their best result is for a 3D backbone based on I3D. They keep time dimension in 3D convolutions and pick the frame with the annotated masks to push in the decoder and make a prediction.</p>\n<p>Top winners of the just completed <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection\" target=\"_blank\">vesuvius</a> competition also use some models with 3D backbones. The problem is slightly different but the idea is the same.</p>\n<p>I've spent a few days on testing a lot of 3D backbones coming from the great <a href=\"https://github.com/open-mmlab/mmaction2\" target=\"_blank\">MMAction2</a> implementations such as:</p>\n<ul>\n<li>Vanilla Resnet3D </li>\n<li>Resnet3D - irCSN 50/152</li>\n<li>Resnet(2+1)D</li>\n<li>MViTv2<br>\nand more …</li>\n</ul>\n<p>It works but never better than a basic 2D backbone as we can see in the NoteBooks/Code section.<br>\nWhatever pooling or extracting the correct slice to feed the decoder, the results look not matching the paper. </p>\n<p>Anyone with similar disapointing results on 3D modeling?</p>",
      "rawMarkdown": "Hi,\n\nI would like to open the discussion on 3D models. According to the Contrails paper, their best result is for a 3D backbone based on I3D. They keep time dimension in 3D convolutions and pick the frame with the annotated masks to push in the decoder and make a prediction.\n\nTop winners of the just completed [vesuvius](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection) competition also use some models with 3D backbones. The problem is slightly different but the idea is the same.\n\nI've spent a few days on testing a lot of 3D backbones coming from the great [MMAction2](https://github.com/open-mmlab/mmaction2) implementations such as:\n- Vanilla Resnet3D \n- Resnet3D - irCSN 50/152\n- Resnet(2+1)D\n- MViTv2\nand more ...\n\nIt works but never better than a basic 2D backbone as we can see in the NoteBooks/Code section.\nWhatever pooling or extracting the correct slice to feed the decoder, the results look not matching the paper. \n\nAnyone with similar disapointing results on 3D modeling?",
      "votes": 27
    },
    {
      "id": 2320611,
      "postDate": "2023-06-28T00:41:53.817Z",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> your AUCPR results already bit the paper. I don't think that there is any practical sense to \"match\" something like  \"we trained on 16 Tensor Processing Unit V3 (TPUv3) chips\" with batch size 256. I have checked on 3D as well and it slowed me down for both training and LB progression. I found many examples where contrails were present on other frames but they never appeared on the labeled frame. My rationale is that it can add more noise to 3d model rather than benefit… (it is my humble opinion though)</p>",
      "rawMarkdown": "@mpware your AUCPR results already bit the paper. I don't think that there is any practical sense to \"match\" something like  \"we trained on 16 Tensor Processing Unit V3 (TPUv3) chips\" with batch size 256. I have checked on 3D as well and it slowed me down for both training and LB progression. I found many examples where contrails were present on other frames but they never appeared on the labeled frame. My rationale is that it can add more noise to 3d model rather than benefit... (it is my humble opinion though)",
      "votes": 6,
      "replies": [
        {
          "id": 2320888,
          "postDate": "2023-06-28T06:20:41.013Z",
          "content": "<blockquote>\n  <p>your AUCPR results already bit the paper</p>\n</blockquote>\n<p>That's true and my 3D model also beats their ResNet3D-101 AUCPR. I was hoping some predictive power for frames close to the one annotated.</p>",
          "rawMarkdown": "> your AUCPR results already bit the paper\n\nThat's true and my 3D model also beats their ResNet3D-101 AUCPR. I was hoping some predictive power for frames close to the one annotated.",
          "replies": [
            {
              "id": 2321588,
              "postDate": "2023-06-28T17:24:17.580Z",
              "content": "<p>just to add to the discussion: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe8cfa227e41ced0a643622e02ffdfbfa%2F4b03dfa0-73fc-4609-9c07-37ff50c91f2f.png?generation=1687973048856614&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "just to add to the discussion: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe8cfa227e41ced0a643622e02ffdfbfa%2F4b03dfa0-73fc-4609-9c07-37ff50c91f2f.png?generation=1687973048856614&alt=media)",
              "votes": 3
            },
            {
              "id": 2321620,
              "postDate": "2023-06-28T17:43:34.747Z",
              "content": "<p>I see this one does not help the model. Thanks for showing it.</p>",
              "rawMarkdown": "I see this one does not help the model. Thanks for showing it.",
              "votes": 1
            },
            {
              "id": 2367995,
              "postDate": "2023-08-01T01:06:21.503Z",
              "content": "<p>well in fact, the absense of a contrail in the frame 3 should help in theory. I remember that in the paper or in the labelling rules they were saying that one factor to determine if a line was a contrail or not was to see it moving or appearing. So in that case that should help to better distinguish contrails from other static shadows/clouds/things that shouldn't be labeled as contrails.</p>",
              "rawMarkdown": "well in fact, the absense of a contrail in the frame 3 should help in theory. I remember that in the paper or in the labelling rules they were saying that one factor to determine if a line was a contrail or not was to see it moving or appearing. So in that case that should help to better distinguish contrails from other static shadows/clouds/things that shouldn't be labeled as contrails."
            }
          ]
        }
      ]
    },
    {
      "id": 2321461,
      "postDate": "2023-06-28T15:37:01.570Z",
      "content": "<p>I also try to 2.5D and 3D models, accding to the ink competition, using 4-8 frames….no models beat my 2D models</p>",
      "rawMarkdown": "I also try to 2.5D and 3D models, accding to the ink competition, using 4-8 frames....no models beat my 2D models",
      "votes": 1
    },
    {
      "id": 2321319,
      "postDate": "2023-06-28T13:06:24.437Z",
      "content": "<p>Did you use MMACTION framework or only the backbones ?</p>",
      "rawMarkdown": "Did you use MMACTION framework or only the backbones ?",
      "votes": 1,
      "replies": [
        {
          "id": 2321618,
          "postDate": "2023-06-28T17:42:11.067Z",
          "content": "<p>Only the backbones by inheritance to get a custom output. <br>\nFor example for <a href=\"https://github.com/open-mmlab/mmaction2/blob/c331b65dadcd8090ae92f5528e07d4606c5b7c47/mmaction/models/backbones/resnet2plus1d.py#L33\" target=\"_blank\">R(2+1)d</a>:</p>\n<pre><code>\n ():\n    </code></pre>\n<p>…</p>\n<pre><code> ():\n        \n\n        x = self.conv1(x)\n        x = self.maxpool(x)\n        \n        out = [x]\n         layer_name  self.res_layers:\n            res_layer = (self, layer_name)\n            \n            x = res_layer(x)\n            out.append(x)\n\n         out\n</code></pre>",
          "rawMarkdown": "Only the backbones by inheritance to get a custom output. \nFor example for [R(2+1)d](https://github.com/open-mmlab/mmaction2/blob/c331b65dadcd8090ae92f5528e07d4606c5b7c47/mmaction/models/backbones/resnet2plus1d.py#L33):\n\n\n```python\n@MODELS.register_module()\nclass ResNet2Plus1dCustom(ResNet3dCustom):\n    \"\"\"ResNet (2+1)d backbone.\n```\n\n...\n\n\n```python\ndef forward(self, x):\n        \"\"\"Defines the computation performed at every call.\n\n        Args:\n            x (torch.Tensor): The input data.\n\n        Returns:\n            torch.Tensor: The feature of the input\n            samples extracted by the backbone.\n        \"\"\"\n\n        x = self.conv1(x)\n        x = self.maxpool(x)\n        # Keep everything to be returned in out array\n        out = [x]\n        for layer_name in self.res_layers:\n            res_layer = getattr(self, layer_name)\n            # no pool2 in R(2+1)d\n            x = res_layer(x)\n            out.append(x)\n\n        return out\n```\n",
          "votes": 1,
          "replies": [
            {
              "id": 2348873,
              "postDate": "2023-07-18T04:10:19.603Z",
              "content": "<p>when I use mmaction's r2+1D ,an error is reported: KeyError: 'Cannot find None in registry under scope name mmengine'. I found the place where the error occurred: /mmaction/lib/python3.8/site-packages/mmcv/cnn/bricks/conv.py, I think that conv2+1d was not added to MODELS.register_module. But I don't know how to modify it, can you tell me what should I do, thanks for your help. If this is about contest privacy, please skip this comment and thank you as well</p>",
              "rawMarkdown": "when I use mmaction's r2+1D ,an error is reported: KeyError: 'Cannot find None in registry under scope name mmengine'. I found the place where the error occurred: /mmaction/lib/python3.8/site-packages/mmcv/cnn/bricks/conv.py, I think that conv2+1d was not added to MODELS.register_module. But I don't know how to modify it, can you tell me what should I do, thanks for your help. If this is about contest privacy, please skip this comment and thank you as well"
            },
            {
              "id": 2349006,
              "postDate": "2023-07-18T06:08:41.703Z",
              "content": "<p>When you instantiate do you have this?</p>\n<pre><code>cfg = (\n     =,\n     depth=model_depth,\n     pretrained=,\n     pretrained2d=,\n     norm_eval=,\n     conv_cfg=(=),\n     \n     conv1_kernel=(, , ),\n     conv1_stride_t=,\n     ...\n</code></pre>\n<p>I had to comment <code>SyncBN</code> because I've a single GPU.</p>",
              "rawMarkdown": "When you instantiate do you have this?\n\n```python\ncfg = dict(\n     type='ResNet2Plus1d',\n     depth=model_depth,\n     pretrained=None,\n     pretrained2d=False,\n     norm_eval=False,\n     conv_cfg=dict(type='Conv2plus1d'),\n     # norm_cfg=dict(type='SyncBN', requires_grad=True, eps=1e-3),\n     conv1_kernel=(3, 7, 7),\n     conv1_stride_t=1,\n     ...\n```\nI had to comment `SyncBN` because I've a single GPU."
            },
            {
              "id": 2350048,
              "postDate": "2023-07-18T23:37:08.583Z",
              "content": "<p>Thanks for your reply, I will check my code again based on your reply</p>",
              "rawMarkdown": "Thanks for your reply, I will check my code again based on your reply"
            }
          ]
        }
      ]
    },
    {
      "id": 2320637,
      "postDate": "2023-06-28T01:46:19.810Z",
      "content": "<p>I tried x3d on the backbone, which also significantly lowered LB.<br>\nAnd the tendency for the 3D models to be unfavorable was evident from the early stages of the study. Did you see the same trend?</p>",
      "rawMarkdown": "I tried x3d on the backbone, which also significantly lowered LB.\nAnd the tendency for the 3D models to be unfavorable was evident from the early stages of the study. Did you see the same trend?",
      "votes": 2,
      "replies": [
        {
          "id": 2320863,
          "postDate": "2023-06-28T06:09:02.733Z",
          "content": "<p>I'm just monitoring my CV and it's lower, it's around 0.66 with 3D backbone (frames used = 4,3,2,1) instead of 0.69 with 2D.</p>",
          "rawMarkdown": "I'm just monitoring my CV and it's lower, it's around 0.66 with 3D backbone (frames used = 4,3,2,1) instead of 0.69 with 2D.",
          "votes": 1,
          "replies": [
            {
              "id": 2321383,
              "postDate": "2023-06-28T14:14:56.420Z",
              "content": "<p>Thank you for the reply.<br>\nAre you using the same decoder as in the paper for these 3D backbones?</p>",
              "rawMarkdown": "Thank you for the reply.\nAre you using the same decoder as in the paper for these 3D backbones?",
              "votes": 1
            },
            {
              "id": 2321396,
              "postDate": "2023-06-28T14:28:22.533Z",
              "content": "<p>No but I've tried several and results depends more on encoder.</p>",
              "rawMarkdown": "No but I've tried several and results depends more on encoder.",
              "votes": 1
            },
            {
              "id": 2321957,
              "postDate": "2023-06-29T02:40:45.547Z",
              "content": "<blockquote>\n  <p>No but I've tried several and results depends more on encoder.</p>\n</blockquote>\n<p>Then I guess my decoder is wrong.<br>\nI trained with x3d backbone using frames1-4 as you did, but I can't get to 0.6 at all.<br>\nAny hints you can give me about the decoder would be appreciated.</p>",
              "rawMarkdown": ">No but I've tried several and results depends more on encoder.\n\nThen I guess my decoder is wrong.\nI trained with x3d backbone using frames1-4 as you did, but I can't get to 0.6 at all.\nAny hints you can give me about the decoder would be appreciated."
            }
          ]
        }
      ]
    },
    {
      "id": 2320486,
      "postDate": "2023-06-27T20:28:59.640Z",
      "content": "<p>I've tried a few experiments with 2.5D (just using a 2D model and simply adding temporal context as extra channels so images will have 6 or 9 channels instead of the regular 3) but I haven't found anything promising. I was thinking of dabbling with 3D if I rent out a GPU but that's about it right now.</p>",
      "rawMarkdown": "I've tried a few experiments with 2.5D (just using a 2D model and simply adding temporal context as extra channels so images will have 6 or 9 channels instead of the regular 3) but I haven't found anything promising. I was thinking of dabbling with 3D if I rent out a GPU but that's about it right now.",
      "votes": 2
    },
    {
      "id": 2367465,
      "postDate": "2023-07-31T15:07:58.307Z",
      "content": "<p>hi :),<br>\n   i am unable to understand  on picking the frame with the annotated masks . if the input is of shape 3,256,256,8 after the encoder layer the shape would be something like this right 64,32,32,2 given 4th frame is annotated how to pick it .<br>\ni am a begginer in computer vision do correct me if i am wrong somewhere .<br>\ndo suggest me if there is good source to learn about 3d models.<br>\nthank you :)</p>",
      "rawMarkdown": "hi :),\n   i am unable to understand  on picking the frame with the annotated masks . if the input is of shape 3,256,256,8 after the encoder layer the shape would be something like this right 64,32,32,2 given 4th frame is annotated how to pick it .\ni am a begginer in computer vision do correct me if i am wrong somewhere .\ndo suggest me if there is good source to learn about 3d models.\nthank you :)"
    }
  ],
  "comments": [
    {
      "id": 2320611,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2023-06-28T00:41:53.817000",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> your AUCPR results already bit the paper. I don't think that there is any practical sense to \"match\" something like  \"we trained on 16 Tensor Processing Unit V3 (TPUv3) chips\" with batch size 256. I have checked on 3D as well and it slowed me down for both training and LB progression. I found many examples where contrails were present on other frames but they never appeared on the labeled frame. My rationale is that it can add more noise to 3d model rather than benefit… (it is my humble opinion though)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2320888,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2023-06-28T06:20:41.013000",
          "content": "<blockquote>\n  <p>your AUCPR results already bit the paper</p>\n</blockquote>\n<p>That's true and my 3D model also beats their ResNet3D-101 AUCPR. I was hoping some predictive power for frames close to the one annotated.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2321588,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2023-06-28T17:24:17.580000",
              "content": "<p>just to add to the discussion: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2Fe8cfa227e41ced0a643622e02ffdfbfa%2F4b03dfa0-73fc-4609-9c07-37ff50c91f2f.png?generation=1687973048856614&amp;alt=media\" alt=\"\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2321620,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-06-28T17:43:34.747000",
              "content": "<p>I see this one does not help the model. Thanks for showing it.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2367995,
              "author_name": "Enric Domingo",
              "author_url": "",
              "post_date": "2023-08-01T01:06:21.503000",
              "content": "<p>well in fact, the absense of a contrail in the frame 3 should help in theory. I remember that in the paper or in the labelling rules they were saying that one factor to determine if a line was a contrail or not was to see it moving or appearing. So in that case that should help to better distinguish contrails from other static shadows/clouds/things that shouldn't be labeled as contrails.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2321461,
      "author_name": "lyu",
      "author_url": "",
      "post_date": "2023-06-28T15:37:01.570000",
      "content": "<p>I also try to 2.5D and 3D models, accding to the ink competition, using 4-8 frames….no models beat my 2D models</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2321319,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2023-06-28T13:06:24.437000",
      "content": "<p>Did you use MMACTION framework or only the backbones ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2321618,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2023-06-28T17:42:11.067000",
          "content": "<p>Only the backbones by inheritance to get a custom output. <br>\nFor example for <a href=\"https://github.com/open-mmlab/mmaction2/blob/c331b65dadcd8090ae92f5528e07d4606c5b7c47/mmaction/models/backbones/resnet2plus1d.py#L33\" target=\"_blank\">R(2+1)d</a>:</p>\n<pre><code>\n ():\n    </code></pre>\n<p>…</p>\n<pre><code> ():\n        \n\n        x = self.conv1(x)\n        x = self.maxpool(x)\n        \n        out = [x]\n         layer_name  self.res_layers:\n            res_layer = (self, layer_name)\n            \n            x = res_layer(x)\n            out.append(x)\n\n         out\n</code></pre>",
          "votes": 1,
          "replies": [
            {
              "id": 2348873,
              "author_name": "Yu Fujiang",
              "author_url": "",
              "post_date": "2023-07-18T04:10:19.603000",
              "content": "<p>when I use mmaction's r2+1D ,an error is reported: KeyError: 'Cannot find None in registry under scope name mmengine'. I found the place where the error occurred: /mmaction/lib/python3.8/site-packages/mmcv/cnn/bricks/conv.py, I think that conv2+1d was not added to MODELS.register_module. But I don't know how to modify it, can you tell me what should I do, thanks for your help. If this is about contest privacy, please skip this comment and thank you as well</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2349006,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-07-18T06:08:41.703000",
              "content": "<p>When you instantiate do you have this?</p>\n<pre><code>cfg = (\n     =,\n     depth=model_depth,\n     pretrained=,\n     pretrained2d=,\n     norm_eval=,\n     conv_cfg=(=),\n     \n     conv1_kernel=(, , ),\n     conv1_stride_t=,\n     ...\n</code></pre>\n<p>I had to comment <code>SyncBN</code> because I've a single GPU.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2350048,
              "author_name": "Yu Fujiang",
              "author_url": "",
              "post_date": "2023-07-18T23:37:08.583000",
              "content": "<p>Thanks for your reply, I will check my code again based on your reply</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2320637,
      "author_name": "Shota Nishiyama",
      "author_url": "",
      "post_date": "2023-06-28T01:46:19.810000",
      "content": "<p>I tried x3d on the backbone, which also significantly lowered LB.<br>\nAnd the tendency for the 3D models to be unfavorable was evident from the early stages of the study. Did you see the same trend?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2320863,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2023-06-28T06:09:02.733000",
          "content": "<p>I'm just monitoring my CV and it's lower, it's around 0.66 with 3D backbone (frames used = 4,3,2,1) instead of 0.69 with 2D.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2321383,
              "author_name": "Shota Nishiyama",
              "author_url": "",
              "post_date": "2023-06-28T14:14:56.420000",
              "content": "<p>Thank you for the reply.<br>\nAre you using the same decoder as in the paper for these 3D backbones?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2321396,
              "author_name": "MPWARE",
              "author_url": "",
              "post_date": "2023-06-28T14:28:22.533000",
              "content": "<p>No but I've tried several and results depends more on encoder.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2321957,
              "author_name": "Shota Nishiyama",
              "author_url": "",
              "post_date": "2023-06-29T02:40:45.547000",
              "content": "<blockquote>\n  <p>No but I've tried several and results depends more on encoder.</p>\n</blockquote>\n<p>Then I guess my decoder is wrong.<br>\nI trained with x3d backbone using frames1-4 as you did, but I can't get to 0.6 at all.<br>\nAny hints you can give me about the decoder would be appreciated.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2320486,
      "author_name": "Ari",
      "author_url": "",
      "post_date": "2023-06-27T20:28:59.640000",
      "content": "<p>I've tried a few experiments with 2.5D (just using a 2D model and simply adding temporal context as extra channels so images will have 6 or 9 channels instead of the regular 3) but I haven't found anything promising. I was thinking of dabbling with 3D if I rent out a GPU but that's about it right now.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2367465,
      "author_name": "Monkey D Donut",
      "author_url": "",
      "post_date": "2023-07-31T15:07:58.307000",
      "content": "<p>hi :),<br>\n   i am unable to understand  on picking the frame with the annotated masks . if the input is of shape 3,256,256,8 after the encoder layer the shape would be something like this right 64,32,32,2 given 4th frame is annotated how to pick it .<br>\ni am a begginer in computer vision do correct me if i am wrong somewhere .<br>\ndo suggest me if there is good source to learn about 3d models.<br>\nthank you :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2320356": "Hi,\n\nI would like to open the discussion on 3D models. According to the Contrails paper, their best result is for a 3D backbone based on I3D. They keep time dimension in 3D convolutions and pick the frame with the annotated masks to push in the decoder and make a prediction.\n\nTop winners of the just completed [vesuvius](https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection) competition also use some models with 3D backbones. The problem is slightly different but the idea is the same.\n\nI've spent a few days on testing a lot of 3D backbones coming from the great [MMAction2](https://github.com/open-mmlab/mmaction2) implementations such as:\n- Vanilla Resnet3D \n- Resnet3D - irCSN 50/152\n- Resnet(2+1)D\n- MViTv2\nand more ...\n\nIt works but never better than a basic 2D backbone as we can see in the NoteBooks/Code section.\nWhatever pooling or extracting the correct slice to feed the decoder, the results look not matching the paper. \n\nAnyone with similar disapointing results on 3D modeling?",
    "2320611": "@mpware your AUCPR results already bit the paper. I don't think that there is any practical sense to \"match\" something like  \"we trained on 16 Tensor Processing Unit V3 (TPUv3) chips\" with batch size 256. I have checked on 3D as well and it slowed me down for both training and LB progression. I found many examples where contrails were present on other frames but they never appeared on the labeled frame. My rationale is that it can add more noise to 3d model rather than benefit... (it is my humble opinion though)",
    "2321461": "I also try to 2.5D and 3D models, accding to the ink competition, using 4-8 frames....no models beat my 2D models",
    "2321319": "Did you use MMACTION framework or only the backbones ?",
    "2320637": "I tried x3d on the backbone, which also significantly lowered LB.\nAnd the tendency for the 3D models to be unfavorable was evident from the early stages of the study. Did you see the same trend?",
    "2320486": "I've tried a few experiments with 2.5D (just using a 2D model and simply adding temporal context as extra channels so images will have 6 or 9 channels instead of the regular 3) but I haven't found anything promising. I was thinking of dabbling with 3D if I rent out a GPU but that's about it right now.",
    "2367465": "hi :),\n   i am unable to understand  on picking the frame with the annotated masks . if the input is of shape 3,256,256,8 after the encoder layer the shape would be something like this right 64,32,32,2 given 4th frame is annotated how to pick it .\ni am a begginer in computer vision do correct me if i am wrong somewhere .\ndo suggest me if there is good source to learn about 3d models.\nthank you :)"
  }
}