{
  "id": 67501,
  "title": "Trying to read the raw csv data results in error at steak.csv line 42414",
  "url": "/competitions/quickdraw-doodle-recognition/discussion/67501",
  "author_name": "OxFEE1DEAD",
  "post_date": "2018-10-03T11:04:49.497000",
  "votes": -1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I was trying to iterate over the csv files and so some post processing, however when it comes down to the steak.csv file I am getting an error, which is the following:</p>\n\n<pre><code>Error: line contains NULL byte\n</code></pre>\n\n<p>which I tried solving by:</p>\n\n<pre><code># 147 is the steak file on all csv files\nfor f in tqdm(files[147:]):\n    null_byte= False\n    if '\\0' in open(PATH_TO_DATA + f).read():\n        null_byte = True\n\n    with open(PATH_TO_DATA + f) as csvfile:\n\n\n        if null_byte:\n            reader = csv.reader(x.replace('\\0', ',') for x in csvfile)\n        else:\n            reader = csv.reader(csvfile,delimiter=',')\n</code></pre>\n\n<p>Which resulted in </p>\n\n<pre><code>Error: field larger than field limit (131072)\n</code></pre>\n\n<p>When replacing the NULL Byte with a white space instead, it will cut of some information. </p>\n\n<p>Has anyone met the same problem or a solution for this? \nIs there a kind of unique data reader for this data which handles this problem better? </p>\n\n<p>I also tried it with pandas, where the error is:</p>\n\n<pre><code>ParserError: Error tokenizing data. C error: EOF inside string starting at line 42413\n</code></pre>\n\n<p>when changing the engine in the <code>csv_read</code> call to 'python' the following (known) errror occurs:</p>\n\n<pre><code>ParserError: NULL byte detected. This byte cannot be processed in Python's native csv library at the moment, so please pass in engine='c' instead\n</code></pre>\n\n<p>when passed <code>engine = 'c'</code> the above error is coming again. </p>\n\n<p>Thank you for your help</p>\n\n<p>I am working on a Windows 10 Machine, Python 3.6 with Spyder IDE </p>",
  "messages": [
    {
      "id": 397939,
      "postDate": "2018-10-03T11:04:49.497Z",
      "content": "<p>I was trying to iterate over the csv files and so some post processing, however when it comes down to the steak.csv file I am getting an error, which is the following:</p>\n\n<pre><code>Error: line contains NULL byte\n</code></pre>\n\n<p>which I tried solving by:</p>\n\n<pre><code># 147 is the steak file on all csv files\nfor f in tqdm(files[147:]):\n    null_byte= False\n    if '\\0' in open(PATH_TO_DATA + f).read():\n        null_byte = True\n\n    with open(PATH_TO_DATA + f) as csvfile:\n\n\n        if null_byte:\n            reader = csv.reader(x.replace('\\0', ',') for x in csvfile)\n        else:\n            reader = csv.reader(csvfile,delimiter=',')\n</code></pre>\n\n<p>Which resulted in </p>\n\n<pre><code>Error: field larger than field limit (131072)\n</code></pre>\n\n<p>When replacing the NULL Byte with a white space instead, it will cut of some information. </p>\n\n<p>Has anyone met the same problem or a solution for this? \nIs there a kind of unique data reader for this data which handles this problem better? </p>\n\n<p>I also tried it with pandas, where the error is:</p>\n\n<pre><code>ParserError: Error tokenizing data. C error: EOF inside string starting at line 42413\n</code></pre>\n\n<p>when changing the engine in the <code>csv_read</code> call to 'python' the following (known) errror occurs:</p>\n\n<pre><code>ParserError: NULL byte detected. This byte cannot be processed in Python's native csv library at the moment, so please pass in engine='c' instead\n</code></pre>\n\n<p>when passed <code>engine = 'c'</code> the above error is coming again. </p>\n\n<p>Thank you for your help</p>\n\n<p>I am working on a Windows 10 Machine, Python 3.6 with Spyder IDE </p>",
      "rawMarkdown": "I was trying to iterate over the csv files and so some post processing, however when it comes down to the steak.csv file I am getting an error, which is the following:\n\n    Error: line contains NULL byte\n\nwhich I tried solving by:\n\n    # 147 is the steak file on all csv files\n    for f in tqdm(files[147:]):\n        null_byte= False\n        if '\\0' in open(PATH_TO_DATA + f).read():\n            null_byte = True\n    \n        with open(PATH_TO_DATA + f) as csvfile:\n        \n    \n            if null_byte:\n                reader = csv.reader(x.replace('\\0', ',') for x in csvfile)\n            else:\n                reader = csv.reader(csvfile,delimiter=',')\n\nWhich resulted in \n\n    Error: field larger than field limit (131072)\n\nWhen replacing the NULL Byte with a white space instead, it will cut of some information. \n\nHas anyone met the same problem or a solution for this? \nIs there a kind of unique data reader for this data which handles this problem better? \n\nI also tried it with pandas, where the error is:\n\n    ParserError: Error tokenizing data. C error: EOF inside string starting at line 42413\n\nwhen changing the engine in the `csv_read` call to 'python' the following (known) errror occurs:\n\n    ParserError: NULL byte detected. This byte cannot be processed in Python's native csv library at the moment, so please pass in engine='c' instead\n\nwhen passed `engine = 'c'` the above error is coming again. \n\nThank you for your help\n\nI am working on a Windows 10 Machine, Python 3.6 with Spyder IDE ",
      "votes": -1
    },
    {
      "id": 456555,
      "postDate": "2019-01-16T03:31:34.993Z",
      "content": "<p>I am facing a similar issue with another file. I will like to know how you managed to read the file using pandas.</p>",
      "rawMarkdown": "I am facing a similar issue with another file. I will like to know how you managed to read the file using pandas."
    }
  ],
  "comments": [
    {
      "id": 456555,
      "author_name": "Shantanu Oak",
      "author_url": "",
      "post_date": "2019-01-16T03:31:34.993000",
      "content": "<p>I am facing a similar issue with another file. I will like to know how you managed to read the file using pandas.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "397939": "I was trying to iterate over the csv files and so some post processing, however when it comes down to the steak.csv file I am getting an error, which is the following:\n\n    Error: line contains NULL byte\n\nwhich I tried solving by:\n\n    # 147 is the steak file on all csv files\n    for f in tqdm(files[147:]):\n        null_byte= False\n        if '\\0' in open(PATH_TO_DATA + f).read():\n            null_byte = True\n    \n        with open(PATH_TO_DATA + f) as csvfile:\n        \n    \n            if null_byte:\n                reader = csv.reader(x.replace('\\0', ',') for x in csvfile)\n            else:\n                reader = csv.reader(csvfile,delimiter=',')\n\nWhich resulted in \n\n    Error: field larger than field limit (131072)\n\nWhen replacing the NULL Byte with a white space instead, it will cut of some information. \n\nHas anyone met the same problem or a solution for this? \nIs there a kind of unique data reader for this data which handles this problem better? \n\nI also tried it with pandas, where the error is:\n\n    ParserError: Error tokenizing data. C error: EOF inside string starting at line 42413\n\nwhen changing the engine in the `csv_read` call to 'python' the following (known) errror occurs:\n\n    ParserError: NULL byte detected. This byte cannot be processed in Python's native csv library at the moment, so please pass in engine='c' instead\n\nwhen passed `engine = 'c'` the above error is coming again. \n\nThank you for your help\n\nI am working on a Windows 10 Machine, Python 3.6 with Spyder IDE ",
    "456555": "I am facing a similar issue with another file. I will like to know how you managed to read the file using pandas."
  }
}