{
  "topic": {
    "id": 733965,
    "title": "Use of Commercially Hosted LLMs",
    "authorName": "Po-Hao \"Howard\" Chen",
    "commentCount": 12,
    "votes": 18,
    "postDate": "2026-08-09T09:52:50.447000"
  },
  "comments": [
    {
      "id": 3521208,
      "authorName": "Navneet",
      "votes": -1,
      "postDate": "2026-09-05T07:20:15.053000",
      "content": "<p>Thank you for the Use of Commercially Hosted LLMs info <a href=\"https://www.kaggle.com/javaduke95\" target=\"_blank\">@javaduke95</a> </p>"
    },
    {
      "id": 3515499,
      "authorName": "Mayobanex Santana",
      "votes": 1,
      "postDate": "2026-08-21T21:48:19.597000",
      "content": "<p>Ok, by the sentence: <em>\"(for example, extracting labels from reports) will not, by itself, be considered prohibited PRIVATE SHARING of Competition Data outside the Team.\"</em> i can conclude that i can use LLM API (i.e. openAI) to read the reports, which are actually in diifferent languages, to generate the labels, so i don't have to struggle?\nI know that the answer may be abvious but i just want to make 100% sure i can save me so much trouble and focus on model not in getting the target labels without been kicked out from the competence.\nThanks!</p>"
    },
    {
      "id": 3517694,
      "authorName": "Po-Hao \"Howard\" Chen",
      "votes": 1,
      "postDate": "2026-08-27T21:59:45.243000",
      "content": "<p>You can use LLM API, such as those from OpenAI, to read the reports to generate the labels.</p>"
    },
    {
      "id": 3510701,
      "authorName": "Nicolai Karcher",
      "votes": 4,
      "postDate": "2026-08-09T10:23:11.330000",
      "content": "<p>\"Participants remain responsible for ensuring that any external service complies with all applicable Competition Rules and the service's terms of use, including the requirements for reasonable accessibility and minimal cost.</p>\n<p>The Competition Host reserves the right to determine whether a particular service, model, or configuration is reasonably accessible, is prohibitively costly, or otherwise creates an unfair competitive advantage.\"</p>\n<p>Could you give us your assessment regarding the use of common public MRI datasets for this challenge? Many of them only allow research (and no commercial) use, and getting a price to me could be classified as commercial use (I'm not a lawyer though so I don't know).</p>\n<p>For example, <a href=\"https://stanford.redivis.com/datasets/4a2c-4cpkzrn2c?access&amp;requirement=474\" target=\"_blank\">this dataset</a> specifically states: \"Permission is granted to view and use the Dataset without charge for personal, <strong>non-commercial research purposes only</strong>. Any commercial use, sale, or <strong>other monetization is prohibited</strong>.\"</p>\n<p>I understand these restrictions are describing external datasets, so we should probably get in contact with those databases. But I was wondering what your independent assessment would be. Thanks!</p>"
    },
    {
      "id": 3514129,
      "authorName": "PC Jimmmy",
      "votes": 0,
      "postDate": "2026-08-18T15:07:22.713000",
      "content": "<p>That's a great suggestion - contact the source and ask if they are ok with you using the data in a kaggle competition.  Not sure if they would reply, but also past experience suggest kaggle reluctant to reply to this type of question during the competition.</p>"
    },
    {
      "id": 3512977,
      "authorName": "UGUR OZER",
      "votes": 1,
      "postDate": "2026-08-14T21:30:30.767000",
      "content": "<p>To make the dataset question above concrete: the Osteoarthritis Initiative (OAI) is free of charge and open to any researcher after a click-through data use agreement, but the agreement limits use to non-commercial research. The same access pattern covers MRNet, fastMRI+ and SKM-TEA, so a ruling on the category would settle it for essentially every public knee-MRI resource. Given the prize money, could the host confirm whether external datasets under research-only agreements are permitted, and if so whether the winners' open-source obligation is affected? Teams need this before committing pretraining work ahead of the entry deadline. Thank you.</p>"
    },
    {
      "id": 3517696,
      "authorName": "Po-Hao \"Howard\" Chen",
      "votes": 6,
      "postDate": "2026-08-27T22:14:33.257000",
      "content": "<p>From the competition’s perspective, use of these datasets would be considered non-commercial; receipt of competition prize money does not, by itself, make their use a commercial sale or commercial exploitation. These datasets would not be excluded by the competition rules solely because their licenses restrict use to non-commercial purposes.</p>\n<p>A separate consideration is whether an external dataset is reasonably accessible to all competitors. Datasets that can be accessed through a straightforward registration or click-through data-use agreement would generally satisfy this requirement. By contrast, datasets requiring institution-specific approvals, negotiated written agreements, IRB approval, lengthy credentialing, or other substantial administrative steps may present a meaningful accessibility barrier, particularly given the remaining competition timeline. Such requirements could therefore affect whether a dataset is considered allowable under the competition’s external-data rules, independent of whether its use is characterized as commercial or non-commercial.</p>\n<p>Our job (Host) is tough here because RSNA and Kaggle do not own or control these external datasets and cannot interpret or waive their individual data-use agreements, expedite requests and so on, on behalf of the dataset owners. Teams remain responsible for ensuring that their intended use, including participation in a prize-bearing competition and any required release of code or models, is permitted under the applicable dataset license or agreement.</p>\n<p>Therefore, to the extent the question concerns the <em>competition rules</em>, readily accessible datasets with non-commercial restrictions are not prohibited on that basis alone. Dataset-specific licensing obligations remain the responsibility of each team, and datasets involving substantial approval or access hurdles may need to be considered separately under the competition’s equal-accessibility requirements.</p>\n<p>Sorry for the verbose response. I hope it's at least a little helpful…</p>"
    },
    {
      "id": 3511329,
      "authorName": "Ragheb haddara",
      "votes": 1,
      "postDate": "2026-08-10T21:45:23.640000",
      "content": "<p>The rules about commercially hosted LLMs are important here. Since knee MRI interpretation requires domain expertise, it's tempting to use an LLM to help with reasoning. But the competition rules likely restrict this to ensure the solutions are based on the image data, not external reasoning. Check the rules page carefully before incorporating any LLM-based features.</p>"
    },
    {
      "id": 3521660,
      "authorName": "Manish",
      "votes": 0,
      "postDate": "2026-09-06T16:59:14.213000",
      "content": "<p>This is a very welcome clarification from the organizers. Allowing the use of commercially hosted LLMs (like OpenAI, Anthropic, or Google APIs) for tasks like text parsing, label extraction, or metadata enrichment really opens up creative pipeline designs especially for handling the report text efficiently.It’s great to see the rules explicitly state that API inference doesn't count as prohibited private sharing, provided the tools remain lowcost and accessible to everyone. It gives teams the green light to leverage state-of-the-art text processing without the fear of accidental disqualification. Thanks for highlighting this section!</p>"
    },
    {
      "id": 3521200,
      "authorName": "Sebastian Garcia",
      "votes": 0,
      "postDate": "2026-09-05T06:31:11.090000",
      "content": "<p><a href=\"https://www.kaggle.com/javaduke95\" target=\"_blank\">@javaduke95</a> Thanks for all the clarifications.</p>\n<p>Would it be allowed to publish a Kaggle Dataset like the one below if it contains the original radiology report text from the competition data, or fragments/excerpts of that text?</p>\n<p><a href=\"https://www.kaggle.com/datasets/laymond/rsna-knee-abnormality-qwen3-8b-weak-labels\" target=\"_blank\">https://www.kaggle.com/datasets/laymond/rsna-knee-abnormality-qwen3-8b-weak-labels</a></p>\n<p>I’m specifically wondering whether including either the full report text or extracted portions of it in a public dataset would be considered redistribution of Competition Data under the competition rules.</p>"
    },
    {
      "id": 3514027,
      "authorName": "Oscar Yáñez Feijóo",
      "votes": 0,
      "postDate": "2026-08-18T09:26:21.707000",
      "content": "<p>Thanks for clarifying this. Allowing hosted LLMs opens some interesting possibilities for extracting structured labels or features from the reports. Would using LLM-derived labels or embeddings from the report text be fully acceptable as model inputs at both training and inference time? Thanks in advance</p>"
    },
    {
      "id": 3517695,
      "authorName": "Po-Hao \"Howard\" Chen",
      "votes": 0,
      "postDate": "2026-08-27T22:05:41.070000",
      "content": "<p>\"Would using LLM-derived labels or embeddings from the report text be fully acceptable as model inputs at both training and inference time?\" </p>\n<p>Yes. Just make sure you read some of the other discussion threads about what the model is provided as input at \"inference time.\"</p>"
    }
  ],
  "index": {
    "id": "733965",
    "title": "Use of Commercially Hosted LLMs",
    "authorName": "",
    "commentCount": "12",
    "votes": "19",
    "postDate": "2026-08-09 09:52:50.447000"
  },
  "competition": "rsna-knee-abnormality-detection"
}