{
  "id": 148040,
  "title": "Error after some epochs !",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/148040",
  "author_name": "Yaheaal",
  "post_date": "2020-05-03T00:23:41.851000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've build a model on TPU , and there wasn't any problem, but every time I run the model after epoch number 13 it gave me this error :</p>\n\n<blockquote>\n  <p>Unable to find a context_id matching the specified one (4950896286011778607). Perhaps the worker was restarted, or the context was GC'd?</p>\n</blockquote>\n\n<p>Note : In some times the error was different but still after epoch 13 .\nAny ideas ?</p>",
  "messages": [
    {
      "id": 830885,
      "postDate": "2020-05-03T00:23:41.850Z",
      "content": "<p>I've build a model on TPU , and there wasn't any problem, but every time I run the model after epoch number 13 it gave me this error :</p>\n\n<blockquote>\n  <p>Unable to find a context_id matching the specified one (4950896286011778607). Perhaps the worker was restarted, or the context was GC'd?</p>\n</blockquote>\n\n<p>Note : In some times the error was different but still after epoch 13 .\nAny ideas ?</p>",
      "rawMarkdown": "I've build a model on TPU , and there wasn't any problem, but every time I run the model after epoch number 13 it gave me this error :\n&gt; Unable to find a context_id matching the specified one (4950896286011778607). Perhaps the worker was restarted, or the context was GC'd?\n\nNote : In some times the error was different but still after epoch 13 .\nAny ideas ?",
      "votes": 3
    },
    {
      "id": 831906,
      "postDate": "2020-05-03T17:11:46.730Z",
      "content": "<p>I was also getting the same error. I think it has to do with tf.data module. I fixed it by removing .cache() method from the training and validation set declaration. This issue can be found on github <a href=\"https://github.com/huan/tensorflow-handbook-tpu/issues/1\">here</a></p>",
      "rawMarkdown": "I was also getting the same error. I think it has to do with tf.data module. I fixed it by removing .cache() method from the training and validation set declaration. This issue can be found on github [here](https://github.com/huan/tensorflow-handbook-tpu/issues/1)",
      "votes": 1,
      "replies": [
        {
          "id": 832126,
          "postDate": "2020-05-03T22:05:35.030Z",
          "content": "<p>Good</p>",
          "rawMarkdown": "Good"
        },
        {
          "id": 832986,
          "postDate": "2020-05-04T14:36:45.797Z",
          "content": "<p>It works , Thanks :)</p>",
          "rawMarkdown": "It works , Thanks :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 831017,
      "postDate": "2020-05-03T03:33:13.207Z",
      "content": "<p>I am also facing the same problem in my <a href=\"https://www.kaggle.com/prateekagnihotri/image-augmentation-on-tpu?scriptVersionId=33165179#Train\">kernel</a>. It is working fine till 12th epoch, then giving this error.</p>",
      "rawMarkdown": "I am also facing the same problem in my [kernel](https://www.kaggle.com/prateekagnihotri/image-augmentation-on-tpu?scriptVersionId=33165179#Train). It is working fine till 12th epoch, then giving this error."
    }
  ],
  "comments": [
    {
      "id": 831906,
      "author_name": "Nikhil Dange",
      "author_url": "",
      "post_date": "2020-05-03T17:11:46.730000",
      "content": "<p>I was also getting the same error. I think it has to do with tf.data module. I fixed it by removing .cache() method from the training and validation set declaration. This issue can be found on github <a href=\"https://github.com/huan/tensorflow-handbook-tpu/issues/1\">here</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 832126,
          "author_name": "Deoth GUEI",
          "author_url": "",
          "post_date": "2020-05-03T22:05:35.030000",
          "content": "<p>Good</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 832986,
          "author_name": "Yaheaal",
          "author_url": "",
          "post_date": "2020-05-04T14:36:45.797000",
          "content": "<p>It works , Thanks :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 831017,
      "author_name": "Prateek",
      "author_url": "",
      "post_date": "2020-05-03T03:33:13.207000",
      "content": "<p>I am also facing the same problem in my <a href=\"https://www.kaggle.com/prateekagnihotri/image-augmentation-on-tpu?scriptVersionId=33165179#Train\">kernel</a>. It is working fine till 12th epoch, then giving this error.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "830885": "I've build a model on TPU , and there wasn't any problem, but every time I run the model after epoch number 13 it gave me this error :\n&gt; Unable to find a context_id matching the specified one (4950896286011778607). Perhaps the worker was restarted, or the context was GC'd?\n\nNote : In some times the error was different but still after epoch 13 .\nAny ideas ?",
    "831906": "I was also getting the same error. I think it has to do with tf.data module. I fixed it by removing .cache() method from the training and validation set declaration. This issue can be found on github [here](https://github.com/huan/tensorflow-handbook-tpu/issues/1)",
    "831017": "I am also facing the same problem in my [kernel](https://www.kaggle.com/prateekagnihotri/image-augmentation-on-tpu?scriptVersionId=33165179#Train). It is working fine till 12th epoch, then giving this error."
  }
}