{
  "id": 380138,
  "title": "Visual Transformer (ViT) vs ResNet50 | My Experimentation Results",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/380138",
  "author_name": "Umong Sain",
  "post_date": "2023-01-22T10:33:33.136000",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,<br>\nIn <a href=\"https://www.kaggle.com/code/umongsain/vision-transformer-from-scratch-pytorch\" target=\"_blank\">this notebook</a>, I have tried to implement ViT from scratch. Later, I have trained both ViT and ResNet50 on a small subset for few epochs. I oversampled the positive images as the dataset is imbalanced.</p>\n<p><strong>What have I learned so far?</strong></p>\n<ul>\n<li>Though ViT shows slightly better results than ResNet50, the difference is avoidable.</li>\n<li>Both models performed poorly on mammogram datasets.</li>\n<li>Changing  the hyperparameters did not affect the result that much.</li>\n</ul>\n<p><strong>Conclusion</strong><br>\nFocus should be put on preprocessing and augmentation so that models can learn more easily.</p>\n<p>Any feedback would be greatly appreciated.</p>",
  "messages": [
    {
      "id": 2110655,
      "postDate": "2023-01-22T10:33:33.137Z",
      "content": "<p>Hello everyone,<br>\nIn <a href=\"https://www.kaggle.com/code/umongsain/vision-transformer-from-scratch-pytorch\" target=\"_blank\">this notebook</a>, I have tried to implement ViT from scratch. Later, I have trained both ViT and ResNet50 on a small subset for few epochs. I oversampled the positive images as the dataset is imbalanced.</p>\n<p><strong>What have I learned so far?</strong></p>\n<ul>\n<li>Though ViT shows slightly better results than ResNet50, the difference is avoidable.</li>\n<li>Both models performed poorly on mammogram datasets.</li>\n<li>Changing  the hyperparameters did not affect the result that much.</li>\n</ul>\n<p><strong>Conclusion</strong><br>\nFocus should be put on preprocessing and augmentation so that models can learn more easily.</p>\n<p>Any feedback would be greatly appreciated.</p>",
      "rawMarkdown": "Hello everyone,\nIn [this notebook](https://www.kaggle.com/code/umongsain/vision-transformer-from-scratch-pytorch), I have tried to implement ViT from scratch. Later, I have trained both ViT and ResNet50 on a small subset for few epochs. I oversampled the positive images as the dataset is imbalanced.\n\n**What have I learned so far?**\n- Though ViT shows slightly better results than ResNet50, the difference is avoidable.\n- Both models performed poorly on mammogram datasets.\n- Changing  the hyperparameters did not affect the result that much.\n\n**Conclusion**\nFocus should be put on preprocessing and augmentation so that models can learn more easily.\n\nAny feedback would be greatly appreciated.",
      "votes": 6
    },
    {
      "id": 2111209,
      "postDate": "2023-01-22T17:26:37.343Z",
      "content": "<p>These results are great, showing that a lot of deep learning problems mostly depend on data not that much on model-centric approaches nowadays. I think we have reached a point in deep learning where model-centric approaches deliever tiny improvements, we are using up to billions of parameters and huge datasets to train our models, still a lot of parameters are redundant and could be pruned out (as well as slight improvements in score) Which leads to the only possible conclusion, using more data-centric approaches, like you mentioned, preprocessing, data cleaning, understanding the data. </p>\n<p>These methods should lead to a way better improvement instead of bloating models with more and more parameters and focusing only on the model as well as its hyperparameters.</p>",
      "rawMarkdown": "These results are great, showing that a lot of deep learning problems mostly depend on data not that much on model-centric approaches nowadays. I think we have reached a point in deep learning where model-centric approaches deliever tiny improvements, we are using up to billions of parameters and huge datasets to train our models, still a lot of parameters are redundant and could be pruned out (as well as slight improvements in score) Which leads to the only possible conclusion, using more data-centric approaches, like you mentioned, preprocessing, data cleaning, understanding the data. \n\nThese methods should lead to a way better improvement instead of bloating models with more and more parameters and focusing only on the model as well as its hyperparameters.",
      "votes": 3
    },
    {
      "id": 2111531,
      "postDate": "2023-01-23T00:19:50.733Z",
      "content": "<p>Broken data -&gt; Broken model..</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Broken data -> Broken model..\n\nThe Devastator.\n"
    }
  ],
  "comments": [
    {
      "id": 2111209,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2023-01-22T17:26:37.343000",
      "content": "<p>These results are great, showing that a lot of deep learning problems mostly depend on data not that much on model-centric approaches nowadays. I think we have reached a point in deep learning where model-centric approaches deliever tiny improvements, we are using up to billions of parameters and huge datasets to train our models, still a lot of parameters are redundant and could be pruned out (as well as slight improvements in score) Which leads to the only possible conclusion, using more data-centric approaches, like you mentioned, preprocessing, data cleaning, understanding the data. </p>\n<p>These methods should lead to a way better improvement instead of bloating models with more and more parameters and focusing only on the model as well as its hyperparameters.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2111531,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2023-01-23T00:19:50.733000",
      "content": "<p>Broken data -&gt; Broken model..</p>\n<p>The Devastator.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2110655": "Hello everyone,\nIn [this notebook](https://www.kaggle.com/code/umongsain/vision-transformer-from-scratch-pytorch), I have tried to implement ViT from scratch. Later, I have trained both ViT and ResNet50 on a small subset for few epochs. I oversampled the positive images as the dataset is imbalanced.\n\n**What have I learned so far?**\n- Though ViT shows slightly better results than ResNet50, the difference is avoidable.\n- Both models performed poorly on mammogram datasets.\n- Changing  the hyperparameters did not affect the result that much.\n\n**Conclusion**\nFocus should be put on preprocessing and augmentation so that models can learn more easily.\n\nAny feedback would be greatly appreciated.",
    "2111209": "These results are great, showing that a lot of deep learning problems mostly depend on data not that much on model-centric approaches nowadays. I think we have reached a point in deep learning where model-centric approaches deliever tiny improvements, we are using up to billions of parameters and huge datasets to train our models, still a lot of parameters are redundant and could be pruned out (as well as slight improvements in score) Which leads to the only possible conclusion, using more data-centric approaches, like you mentioned, preprocessing, data cleaning, understanding the data. \n\nThese methods should lead to a way better improvement instead of bloating models with more and more parameters and focusing only on the model as well as its hyperparameters.",
    "2111531": "Broken data -> Broken model..\n\nThe Devastator.\n"
  }
}