{
  "id": 428543,
  "title": "Computational Resources- Advice?",
  "url": "/competitions/asl-fingerspelling/discussion/428543",
  "author_name": "Kyle Proffitt",
  "post_date": "2023-08-01T19:55:31.869000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Okay, so once again, mostly new user here. I've hit my GPU quota for the first time in trying things with this competition. Well now what?! - even before I hit that, I wondered if I might not get further trying to run some of this offline. I found this post: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/421924#2337500\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/421924#2337500</a> And I attempted to follow some of that, but I think maybe it's never going to work for me? because I'm on a Mac with their M1 chip or whatever, yes GPU, but not Nvidia GPU, so from what I can tell it's not going to work. </p>\n<p>I also sleuthed a little and found recommendation to use colab. I tried that some, and it seems like it might work, except that there's a storage limit, and so I can't download this full dataset to work on it unless I spend some $. And I'm cheap. Perhaps I could break the data into chunks someway so I could at least run smaller experiments, but that will be its own effort of course, instead of just using the code as I already have it… and so far, every time I try to make a small change, I spend an hour troubleshooting to make it work again…</p>\n<p>I wondered before joining this competition what some of the factors are that separate out the top performers- no doubt it's knowledge and experience, but I think I also read/heard somewhere that one of the major limits in a competition is just how many things you can try. More people and more resources allows you to try a lot of experiments-- so is it the case that the top performers are often spending significantly on computing resources to achieve what they do? and/or do they just have pretty nice home setups because they're into this stuff anyway (not lost on me that I have my own set of privileges). </p>\n<p>In any case, I wonder if anyone would chime in to say how they usually approach something like this and if there are any tips on how else to try and run through experiments when you meet the kaggle gpu quota. I did also try turning off gpu, and I think this slowed the model performance to the point it was never going to complete. Glad I get to participate.</p>",
  "messages": [
    {
      "id": 2369544,
      "postDate": "2023-08-01T19:55:31.870Z",
      "content": "<p>Okay, so once again, mostly new user here. I've hit my GPU quota for the first time in trying things with this competition. Well now what?! - even before I hit that, I wondered if I might not get further trying to run some of this offline. I found this post: <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/421924#2337500\" target=\"_blank\">https://www.kaggle.com/competitions/asl-fingerspelling/discussion/421924#2337500</a> And I attempted to follow some of that, but I think maybe it's never going to work for me? because I'm on a Mac with their M1 chip or whatever, yes GPU, but not Nvidia GPU, so from what I can tell it's not going to work. </p>\n<p>I also sleuthed a little and found recommendation to use colab. I tried that some, and it seems like it might work, except that there's a storage limit, and so I can't download this full dataset to work on it unless I spend some $. And I'm cheap. Perhaps I could break the data into chunks someway so I could at least run smaller experiments, but that will be its own effort of course, instead of just using the code as I already have it… and so far, every time I try to make a small change, I spend an hour troubleshooting to make it work again…</p>\n<p>I wondered before joining this competition what some of the factors are that separate out the top performers- no doubt it's knowledge and experience, but I think I also read/heard somewhere that one of the major limits in a competition is just how many things you can try. More people and more resources allows you to try a lot of experiments-- so is it the case that the top performers are often spending significantly on computing resources to achieve what they do? and/or do they just have pretty nice home setups because they're into this stuff anyway (not lost on me that I have my own set of privileges). </p>\n<p>In any case, I wonder if anyone would chime in to say how they usually approach something like this and if there are any tips on how else to try and run through experiments when you meet the kaggle gpu quota. I did also try turning off gpu, and I think this slowed the model performance to the point it was never going to complete. Glad I get to participate.</p>",
      "rawMarkdown": "Okay, so once again, mostly new user here. I've hit my GPU quota for the first time in trying things with this competition. Well now what?! - even before I hit that, I wondered if I might not get further trying to run some of this offline. I found this post: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/421924#2337500 And I attempted to follow some of that, but I think maybe it's never going to work for me? because I'm on a Mac with their M1 chip or whatever, yes GPU, but not Nvidia GPU, so from what I can tell it's not going to work. \n\nI also sleuthed a little and found recommendation to use colab. I tried that some, and it seems like it might work, except that there's a storage limit, and so I can't download this full dataset to work on it unless I spend some $. And I'm cheap. Perhaps I could break the data into chunks someway so I could at least run smaller experiments, but that will be its own effort of course, instead of just using the code as I already have it... and so far, every time I try to make a small change, I spend an hour troubleshooting to make it work again...\n\nI wondered before joining this competition what some of the factors are that separate out the top performers- no doubt it's knowledge and experience, but I think I also read/heard somewhere that one of the major limits in a competition is just how many things you can try. More people and more resources allows you to try a lot of experiments-- so is it the case that the top performers are often spending significantly on computing resources to achieve what they do? and/or do they just have pretty nice home setups because they're into this stuff anyway (not lost on me that I have my own set of privileges). \n\nIn any case, I wonder if anyone would chime in to say how they usually approach something like this and if there are any tips on how else to try and run through experiments when you meet the kaggle gpu quota. I did also try turning off gpu, and I think this slowed the model performance to the point it was never going to complete. Glad I get to participate.",
      "votes": 2
    },
    {
      "id": 2391079,
      "postDate": "2023-08-15T01:23:02.103Z",
      "content": "<p>Hi there! To try and balance out the amount of my quota I use, I've been going back and forth between google colab and here. Whenever I'm on google colab it's mainly testing stuff and whatnot, and when I'm ready to submit I go back here and run the full thing. </p>",
      "rawMarkdown": "Hi there! To try and balance out the amount of my quota I use, I've been going back and forth between google colab and here. Whenever I'm on google colab it's mainly testing stuff and whatnot, and when I'm ready to submit I go back here and run the full thing. ",
      "replies": [
        {
          "id": 2395331,
          "postDate": "2023-08-17T13:08:10.267Z",
          "content": "<p>well as I said, my issue with Colab was that I couldn't figure out how to re-run what I'm doing as-is, like using the entire dataset, because the dataset is too big to be usable within the free limits, and I didn't want to spend money just to try things. There's probably a way to chunk it and run test code, but you know, that would be its own entire project for me. Thanks for commenting.</p>",
          "rawMarkdown": "well as I said, my issue with Colab was that I couldn't figure out how to re-run what I'm doing as-is, like using the entire dataset, because the dataset is too big to be usable within the free limits, and I didn't want to spend money just to try things. There's probably a way to chunk it and run test code, but you know, that would be its own entire project for me. Thanks for commenting."
        }
      ]
    },
    {
      "id": 2369572,
      "postDate": "2023-08-01T20:35:23.920Z",
      "content": "<p>you could use the tpu on kaggle. depending on your model size training can be as fast as 8 seconds per epoch. it is super easy if your code is in keras/tensorflow. just two extra lines. </p>",
      "rawMarkdown": "you could use the tpu on kaggle. depending on your model size training can be as fast as 8 seconds per epoch. it is super easy if your code is in keras/tensorflow. just two extra lines. ",
      "replies": [
        {
          "id": 2369587,
          "postDate": "2023-08-01T21:07:51.557Z",
          "content": "<p>Okay I should try this at least. Probably sorta stop gap in terms of the bigger question of computing resources, but it lets me run some experiments now hopefully. I can probably find those two lines when I start looking around, but if you want to paste them right here, that’s okay too :-)</p>",
          "rawMarkdown": "Okay I should try this at least. Probably sorta stop gap in terms of the bigger question of computing resources, but it lets me run some experiments now hopefully. I can probably find those two lines when I start looking around, but if you want to paste them right here, that’s okay too :-)",
          "replies": [
            {
              "id": 2369737,
              "postDate": "2023-08-02T02:10:06.563Z",
              "content": "<p>well, I'd like to report that I think I'm getting it working. Of course when you turn on the TPU it directs you to the documentation <a href=\"https://www.kaggle.com/docs/tpu\" target=\"_blank\">here</a>. And that tells you to add:</p>\n<h1>detect and init the TPU</h1>\n<p>tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect()</p>\n<h1>instantiate a distribution strategy</h1>\n<p>tpu_strategy = tf.distribute.experimental.TPUStrategy(tpu)</p>\n<p>and then</p>\n<h1>instantiating the model in the strategy scope creates the model on the TPU</h1>\n<p>with tpu_strategy.scope():<br>\n    model = tf.keras.Sequential( … ) # define your model normally<br>\n    model.compile( … )</p>\n<p>my model = was within a def get_model function, so I think I put this line in the right place within that get_model function. </p>\n<p>I also ran into some additional issues, needed to make Jit_compile = False because TPUs don't do that, apparently. Had to throw in a </p>\n<p>!pip install pandas pyarrow</p>\n<p>for some reason where I didn't need this prior to work with the parquet files. And I'm getting a decent number more warnings-- like I'm rebuilding the TFRecord files myself, trying to use different data, and I'm getting warnings about retracing… but it's still running.</p>\n<p>And no, I didn't put the with tpu_strategy.scope() in the right place. I hit errors. But, I eventually put things in the right place. Let's leave it at that. It's running again!</p>\n<p>Oh, no it's not. I lied. There's always another problem. many possibilities about batch size and incompatible portions like CTC loss and other that may not work with TPUs. Oh well.</p>",
              "rawMarkdown": "well, I'd like to report that I think I'm getting it working. Of course when you turn on the TPU it directs you to the documentation [here](https://www.kaggle.com/docs/tpu). And that tells you to add:\n# detect and init the TPU\ntpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect()\n\n# instantiate a distribution strategy\ntpu_strategy = tf.distribute.experimental.TPUStrategy(tpu)\n\nand then\n# instantiating the model in the strategy scope creates the model on the TPU\nwith tpu_strategy.scope():\n    model = tf.keras.Sequential( … ) # define your model normally\n    model.compile( … )\n\nmy model = was within a def get_model function, so I think I put this line in the right place within that get_model function. \n\nI also ran into some additional issues, needed to make Jit_compile = False because TPUs don't do that, apparently. Had to throw in a \n\n!pip install pandas pyarrow\n\nfor some reason where I didn't need this prior to work with the parquet files. And I'm getting a decent number more warnings-- like I'm rebuilding the TFRecord files myself, trying to use different data, and I'm getting warnings about retracing... but it's still running.\n\nAnd no, I didn't put the with tpu_strategy.scope() in the right place. I hit errors. But, I eventually put things in the right place. Let's leave it at that. It's running again!\n\nOh, no it's not. I lied. There's always another problem. many possibilities about batch size and incompatible portions like CTC loss and other that may not work with TPUs. Oh well."
            },
            {
              "id": 2370080,
              "postDate": "2023-08-02T07:20:41.250Z",
              "content": "<p>Regarding batch size, you just need to add drop_reminder=True. Regarding CTC, look in the Code section for my notebook 'CTC on TPU' :)</p>",
              "rawMarkdown": "Regarding batch size, you just need to add drop_reminder=True. Regarding CTC, look in the Code section for my notebook 'CTC on TPU' :)"
            },
            {
              "id": 2370713,
              "postDate": "2023-08-02T15:37:56.413Z",
              "content": "<p>Wow. It's working! Man, people really assume too much about what I can figure out though. But I did figure it out, so far. Definitely used your notebook a lot. Thank you for sharing that. It's really cool how much faster it runs on the TPU!</p>",
              "rawMarkdown": "Wow. It's working! Man, people really assume too much about what I can figure out though. But I did figure it out, so far. Definitely used your notebook a lot. Thank you for sharing that. It's really cool how much faster it runs on the TPU!"
            },
            {
              "id": 2370895,
              "postDate": "2023-08-02T17:36:40.637Z",
              "content": "<p>one quick piggyback question to see if others deal with it. I get the \"this webpage is using significant energy and had to be reloaded\" message in Safari on my Mac sometimes, which of course ruins an interactive run when it happens. Maybe that's a problem that's really Mac and Safari-specific and you'd advise I do stuff like clear my cache, but I wonder if it's a more general problem people deal with, and of course if there's advice. I also probably need to learn to save output or state in some way so that if/when things like that happen, I don't have to start the notebook over from nothing. Like right now part of my code is writing the TFrecords, which takes time, and I should probably just export those once and save them as a new dataset I can pull from… soon.</p>",
              "rawMarkdown": "one quick piggyback question to see if others deal with it. I get the \"this webpage is using significant energy and had to be reloaded\" message in Safari on my Mac sometimes, which of course ruins an interactive run when it happens. Maybe that's a problem that's really Mac and Safari-specific and you'd advise I do stuff like clear my cache, but I wonder if it's a more general problem people deal with, and of course if there's advice. I also probably need to learn to save output or state in some way so that if/when things like that happen, I don't have to start the notebook over from nothing. Like right now part of my code is writing the TFrecords, which takes time, and I should probably just export those once and save them as a new dataset I can pull from... soon."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2391079,
      "author_name": "Zachary Grinberg",
      "author_url": "",
      "post_date": "2023-08-15T01:23:02.103000",
      "content": "<p>Hi there! To try and balance out the amount of my quota I use, I've been going back and forth between google colab and here. Whenever I'm on google colab it's mainly testing stuff and whatnot, and when I'm ready to submit I go back here and run the full thing. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2395331,
          "author_name": "Kyle Proffitt",
          "author_url": "",
          "post_date": "2023-08-17T13:08:10.267000",
          "content": "<p>well as I said, my issue with Colab was that I couldn't figure out how to re-run what I'm doing as-is, like using the entire dataset, because the dataset is too big to be usable within the free limits, and I didn't want to spend money just to try things. There's probably a way to chunk it and run test code, but you know, that would be its own entire project for me. Thanks for commenting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2369572,
      "author_name": "nymfree",
      "author_url": "",
      "post_date": "2023-08-01T20:35:23.920000",
      "content": "<p>you could use the tpu on kaggle. depending on your model size training can be as fast as 8 seconds per epoch. it is super easy if your code is in keras/tensorflow. just two extra lines. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2369587,
          "author_name": "Kyle Proffitt",
          "author_url": "",
          "post_date": "2023-08-01T21:07:51.557000",
          "content": "<p>Okay I should try this at least. Probably sorta stop gap in terms of the bigger question of computing resources, but it lets me run some experiments now hopefully. I can probably find those two lines when I start looking around, but if you want to paste them right here, that’s okay too :-)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2369737,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-08-02T02:10:06.563000",
              "content": "<p>well, I'd like to report that I think I'm getting it working. Of course when you turn on the TPU it directs you to the documentation <a href=\"https://www.kaggle.com/docs/tpu\" target=\"_blank\">here</a>. And that tells you to add:</p>\n<h1>detect and init the TPU</h1>\n<p>tpu = tf.distribute.cluster_resolver.TPUClusterResolver.connect()</p>\n<h1>instantiate a distribution strategy</h1>\n<p>tpu_strategy = tf.distribute.experimental.TPUStrategy(tpu)</p>\n<p>and then</p>\n<h1>instantiating the model in the strategy scope creates the model on the TPU</h1>\n<p>with tpu_strategy.scope():<br>\n    model = tf.keras.Sequential( … ) # define your model normally<br>\n    model.compile( … )</p>\n<p>my model = was within a def get_model function, so I think I put this line in the right place within that get_model function. </p>\n<p>I also ran into some additional issues, needed to make Jit_compile = False because TPUs don't do that, apparently. Had to throw in a </p>\n<p>!pip install pandas pyarrow</p>\n<p>for some reason where I didn't need this prior to work with the parquet files. And I'm getting a decent number more warnings-- like I'm rebuilding the TFRecord files myself, trying to use different data, and I'm getting warnings about retracing… but it's still running.</p>\n<p>And no, I didn't put the with tpu_strategy.scope() in the right place. I hit errors. But, I eventually put things in the right place. Let's leave it at that. It's running again!</p>\n<p>Oh, no it's not. I lied. There's always another problem. many possibilities about batch size and incompatible portions like CTC loss and other that may not work with TPUs. Oh well.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2370080,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2023-08-02T07:20:41.250000",
              "content": "<p>Regarding batch size, you just need to add drop_reminder=True. Regarding CTC, look in the Code section for my notebook 'CTC on TPU' :)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2370713,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-08-02T15:37:56.413000",
              "content": "<p>Wow. It's working! Man, people really assume too much about what I can figure out though. But I did figure it out, so far. Definitely used your notebook a lot. Thank you for sharing that. It's really cool how much faster it runs on the TPU!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2370895,
              "author_name": "Kyle Proffitt",
              "author_url": "",
              "post_date": "2023-08-02T17:36:40.637000",
              "content": "<p>one quick piggyback question to see if others deal with it. I get the \"this webpage is using significant energy and had to be reloaded\" message in Safari on my Mac sometimes, which of course ruins an interactive run when it happens. Maybe that's a problem that's really Mac and Safari-specific and you'd advise I do stuff like clear my cache, but I wonder if it's a more general problem people deal with, and of course if there's advice. I also probably need to learn to save output or state in some way so that if/when things like that happen, I don't have to start the notebook over from nothing. Like right now part of my code is writing the TFrecords, which takes time, and I should probably just export those once and save them as a new dataset I can pull from… soon.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2369544": "Okay, so once again, mostly new user here. I've hit my GPU quota for the first time in trying things with this competition. Well now what?! - even before I hit that, I wondered if I might not get further trying to run some of this offline. I found this post: https://www.kaggle.com/competitions/asl-fingerspelling/discussion/421924#2337500 And I attempted to follow some of that, but I think maybe it's never going to work for me? because I'm on a Mac with their M1 chip or whatever, yes GPU, but not Nvidia GPU, so from what I can tell it's not going to work. \n\nI also sleuthed a little and found recommendation to use colab. I tried that some, and it seems like it might work, except that there's a storage limit, and so I can't download this full dataset to work on it unless I spend some $. And I'm cheap. Perhaps I could break the data into chunks someway so I could at least run smaller experiments, but that will be its own effort of course, instead of just using the code as I already have it... and so far, every time I try to make a small change, I spend an hour troubleshooting to make it work again...\n\nI wondered before joining this competition what some of the factors are that separate out the top performers- no doubt it's knowledge and experience, but I think I also read/heard somewhere that one of the major limits in a competition is just how many things you can try. More people and more resources allows you to try a lot of experiments-- so is it the case that the top performers are often spending significantly on computing resources to achieve what they do? and/or do they just have pretty nice home setups because they're into this stuff anyway (not lost on me that I have my own set of privileges). \n\nIn any case, I wonder if anyone would chime in to say how they usually approach something like this and if there are any tips on how else to try and run through experiments when you meet the kaggle gpu quota. I did also try turning off gpu, and I think this slowed the model performance to the point it was never going to complete. Glad I get to participate.",
    "2391079": "Hi there! To try and balance out the amount of my quota I use, I've been going back and forth between google colab and here. Whenever I'm on google colab it's mainly testing stuff and whatnot, and when I'm ready to submit I go back here and run the full thing. ",
    "2369572": "you could use the tpu on kaggle. depending on your model size training can be as fast as 8 seconds per epoch. it is super easy if your code is in keras/tensorflow. just two extra lines. "
  }
}