{
  "id": 436805,
  "title": "EfficientNet ",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/436805",
  "author_name": "Arman Amedi",
  "post_date": "2023-09-04T07:31:37.584000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Has anyone experimented with EfficientNet, or which CNN model is typically employed in this scenario?</p>",
  "messages": [
    {
      "id": 2425531,
      "postDate": "2023-09-06T02:28:07.733Z",
      "content": "<p>it works for me, whether resnet,efficientnet,convext or Vit transformer in this kaggle compeitions.</p>\n<p>a few comments:</p>\n<ol>\n<li><p>in theory if one architecture can work, other should also work.This is  because the difference in performance in imagenet (image data) for the different architectures are within +/- 10%</p></li>\n<li><p>however, some are easier to work because they are less sensitive to \"bugs\" e.g. if even if you use wrong train hyperparams, wrong input normalisation, forget to load pretrain weights, … some architectures work better.</p></li>\n<li><p>when in doubts, choose the simplest architecture first, usually the plain resnet (like resnet18,34,50). note that resnet only have shortcut and bn trick, it can work even if you forget to normalise input (e.g. you wrongly input 0 to 255 instead of mean +/- std)</p></li>\n<li><p>those architectures that use attention (like trasnformer, SE-resnet, efficientnet), one must use correctly nomalised input. e.g. if the values are too large (0 to 255 instead of -1 to 1), the attention weights become zero and all signed are masked. you get zero output.</p></li>\n<li><p>when you change from 3-channel input to say N-channel, you activation in the initial layers become larger than expected. they can cause problem(4) mentioned above.</p></li>\n</ol>",
      "rawMarkdown": "it works for me, whether resnet,efficientnet,convext or Vit transformer in this kaggle compeitions.\n\na few comments:\n\n1. in theory if one architecture can work, other should also work.This is  because the difference in performance in imagenet (image data) for the different architectures are within +/- 10%\n\n2. however, some are easier to work because they are less sensitive to \"bugs\" e.g. if even if you use wrong train hyperparams, wrong input normalisation, forget to load pretrain weights, ... some architectures work better.\n\n3. when in doubts, choose the simplest architecture first, usually the plain resnet (like resnet18,34,50). note that resnet only have shortcut and bn trick, it can work even if you forget to normalise input (e.g. you wrongly input 0 to 255 instead of mean +/- std)\n\n4. those architectures that use attention (like trasnformer, SE-resnet, efficientnet), one must use correctly nomalised input. e.g. if the values are too large (0 to 255 instead of -1 to 1), the attention weights become zero and all signed are masked. you get zero output.\n\n5. when you change from 3-channel input to say N-channel, you activation in the initial layers become larger than expected. they can cause problem(4) mentioned above.\n",
      "votes": 2
    },
    {
      "id": 2425496,
      "postDate": "2023-09-06T01:29:25.037Z",
      "content": "<p>So far all 2D models I have tried perform equally poorly :)</p>\n<p>It's not the model - it's the data that needs the work.</p>",
      "rawMarkdown": "So far all 2D models I have tried perform equally poorly :)\n\nIt's not the model - it's the data that needs the work."
    },
    {
      "id": 2422706,
      "postDate": "2023-09-04T07:31:37.583Z",
      "content": "<p>Has anyone experimented with EfficientNet, or which CNN model is typically employed in this scenario?</p>",
      "rawMarkdown": "Has anyone experimented with EfficientNet, or which CNN model is typically employed in this scenario?"
    },
    {
      "id": 2422711,
      "postDate": "2023-09-04T07:36:54.033Z",
      "content": "<p>My 2D methods on EfficientNET and Resnet101 achieved similar results</p>",
      "rawMarkdown": "My 2D methods on EfficientNET and Resnet101 achieved similar results",
      "isDeleted": true
    },
    {
      "id": 2425451,
      "postDate": "2023-09-05T23:53:58.403Z",
      "content": "<p>Thanks, for post.</p>",
      "rawMarkdown": "Thanks, for post."
    }
  ],
  "comments": [
    {
      "id": 2425531,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-09-06T02:28:07.733000",
      "content": "<p>it works for me, whether resnet,efficientnet,convext or Vit transformer in this kaggle compeitions.</p>\n<p>a few comments:</p>\n<ol>\n<li><p>in theory if one architecture can work, other should also work.This is  because the difference in performance in imagenet (image data) for the different architectures are within +/- 10%</p></li>\n<li><p>however, some are easier to work because they are less sensitive to \"bugs\" e.g. if even if you use wrong train hyperparams, wrong input normalisation, forget to load pretrain weights, … some architectures work better.</p></li>\n<li><p>when in doubts, choose the simplest architecture first, usually the plain resnet (like resnet18,34,50). note that resnet only have shortcut and bn trick, it can work even if you forget to normalise input (e.g. you wrongly input 0 to 255 instead of mean +/- std)</p></li>\n<li><p>those architectures that use attention (like trasnformer, SE-resnet, efficientnet), one must use correctly nomalised input. e.g. if the values are too large (0 to 255 instead of -1 to 1), the attention weights become zero and all signed are masked. you get zero output.</p></li>\n<li><p>when you change from 3-channel input to say N-channel, you activation in the initial layers become larger than expected. they can cause problem(4) mentioned above.</p></li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2425496,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2023-09-06T01:29:25.037000",
      "content": "<p>So far all 2D models I have tried perform equally poorly :)</p>\n<p>It's not the model - it's the data that needs the work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2422711,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-04T07:36:54.033000",
      "content": "<p>My 2D methods on EfficientNET and Resnet101 achieved similar results</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2425451,
      "author_name": "Kotaro Tokitsu",
      "author_url": "",
      "post_date": "2023-09-05T23:53:58.403000",
      "content": "<p>Thanks, for post.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2425531": "it works for me, whether resnet,efficientnet,convext or Vit transformer in this kaggle compeitions.\n\na few comments:\n\n1. in theory if one architecture can work, other should also work.This is  because the difference in performance in imagenet (image data) for the different architectures are within +/- 10%\n\n2. however, some are easier to work because they are less sensitive to \"bugs\" e.g. if even if you use wrong train hyperparams, wrong input normalisation, forget to load pretrain weights, ... some architectures work better.\n\n3. when in doubts, choose the simplest architecture first, usually the plain resnet (like resnet18,34,50). note that resnet only have shortcut and bn trick, it can work even if you forget to normalise input (e.g. you wrongly input 0 to 255 instead of mean +/- std)\n\n4. those architectures that use attention (like trasnformer, SE-resnet, efficientnet), one must use correctly nomalised input. e.g. if the values are too large (0 to 255 instead of -1 to 1), the attention weights become zero and all signed are masked. you get zero output.\n\n5. when you change from 3-channel input to say N-channel, you activation in the initial layers become larger than expected. they can cause problem(4) mentioned above.\n",
    "2425496": "So far all 2D models I have tried perform equally poorly :)\n\nIt's not the model - it's the data that needs the work.",
    "2422706": "Has anyone experimented with EfficientNet, or which CNN model is typically employed in this scenario?",
    "2422711": "My 2D methods on EfficientNET and Resnet101 achieved similar results",
    "2425451": "Thanks, for post."
  }
}