{
  "topic": {
    "id": 585144,
    "title": "What if you can only use logistic regression...",
    "authorName": "broccoli beef",
    "commentCount": 23,
    "votes": 49,
    "postDate": "2025-06-18T08:41:59.611000"
  },
  "comments": [
    {
      "id": 3233520,
      "authorName": "Ali Osama",
      "votes": 1,
      "postDate": "2025-06-26T22:44:04.570000",
      "content": "<p>Amazing job!!</p>"
    },
    {
      "id": 3233338,
      "authorName": "GurSimran",
      "votes": 1,
      "postDate": "2025-06-26T17:29:00.830000",
      "content": "<p>Thank's for sharing your magic tips, i will try definitely..</p>"
    },
    {
      "id": 3233202,
      "authorName": "shashikant kunwar",
      "votes": 1,
      "postDate": "2025-06-26T15:13:14.733000",
      "content": "<p>it made me to start from the very simplest version and yet i got a good score than my previous multi model complex versions.</p>"
    },
    {
      "id": 3231622,
      "authorName": "Cody Anderson",
      "votes": 1,
      "postDate": "2025-06-24T16:49:15.617000",
      "content": "<p>Impressive! I didn't realize LogisticRegression could perform this well. I've been focusing primarily on LightGBM and XGBoost models.</p>"
    },
    {
      "id": 3231395,
      "authorName": "aci.patlican",
      "votes": 1,
      "postDate": "2025-06-24T09:33:47.160000",
      "content": "<p>Hi,<br>\nIn  a part of the code <strong>\"weight\"</strong> assigned <strong>\"4\"</strong>:</p>\n<pre><code>model = Augmented(\n    make_pipeline(\n        OneHotEncoder(handle_unknown=),\n        LogisticRegression(C=, max_iter=, random_state=)\n    ), X_o_e, y_o, weight_arg=, \n    weight=\n)\n</code></pre>\n<p>Is it assigned consciously or randomly? Then how did it specified?</p>"
    },
    {
      "id": 3231435,
      "authorName": "gowtham-dd",
      "votes": 1,
      "postDate": "2025-06-24T10:42:07.227000",
      "content": "<p>Yeah, the weight=4.0 wasn’t just a random choice — it was something he might tried out during cross-validation and found that it gave the best results with this setup.  There’s no fixed rule for picking this; it really depends on the dataset and how much signal the original data brings in. So in this case, 4 just works best empirically.</p>\n<p>Hope that helps!</p>"
    },
    {
      "id": 3228854,
      "authorName": "Ogulcan",
      "votes": 1,
      "postDate": "2025-06-20T15:32:18.403000",
      "content": "<p>Treating all features as categorical and generating cross-term features is an effective way to help logistic regression capture the underlying structure of the data. However, could including the original dataset multiple times lead to overfitting on those specific samples?</p>"
    },
    {
      "id": 3230138,
      "authorName": "broccoli beef",
      "votes": 0,
      "postDate": "2025-06-22T14:13:39.910000",
      "content": "<p>If you are tuning the number of copies using a CV score, there is always a risk of overfitting to the folds. This is true of hyperparameter tuning in general. You could use a nested CV scheme to validate the tuning methodology; this is seldom done on kaggle though.</p>"
    },
    {
      "id": 3227398,
      "authorName": "paperxd",
      "votes": 1,
      "postDate": "2025-06-19T01:08:20.013000",
      "content": "<p>I am planning to now focus my attention back to this competition, this will surely help, thanks!</p>"
    },
    {
      "id": 3233250,
      "authorName": "Shubham Veer",
      "votes": 0,
      "postDate": "2025-06-26T15:57:23.477000",
      "content": "<p>Did you notice any performance difference when including triplet interactions versus only pairwise interactions?</p>"
    },
    {
      "id": 3232336,
      "authorName": "Arko Bera",
      "votes": 0,
      "postDate": "2025-06-25T15:34:37.483000",
      "content": "<p>Does copying the original data multiple times within a fold make any difference.<br>\nI saw a public notebook and somewhere in the discussion that adding the original data within the fold turned out to be beneficial, but I wonder whether doing the same multiple times will make any difference.</p>"
    },
    {
      "id": 3232305,
      "authorName": "Mahira Banu",
      "votes": 0,
      "postDate": "2025-06-25T15:11:26.867000",
      "content": "<p>Great breakdown — love how you squeeze extra capacity out of a plain LR pipeline with clever feature-crosses and class weighting!</p>\n<p>Highlights I found useful:</p>\n<p>Using LEAD + mutual_info_score to rank 3-way crosses is a neat, lightweight heuristic.</p>\n<p>The Augmented() wrapper for re-weighting original vs. “hallucinated” data is clean and reusable.</p>\n<p>MAP@k scorer + stratified CV keeps evaluation in line with the public LB.</p>"
    },
    {
      "id": 3228709,
      "authorName": "Fantasy284",
      "votes": 0,
      "postDate": "2025-06-20T12:40:15.417000",
      "content": "<p>Thanks for launching this fun playground challenge!</p>"
    },
    {
      "id": 3228631,
      "authorName": "Sarah Arshad",
      "votes": 0,
      "postDate": "2025-06-20T11:29:14.237000",
      "content": "<p>Thanks for launching this fun playground challenge…!</p>"
    },
    {
      "id": 3228528,
      "authorName": "Harsh Gupta",
      "votes": 0,
      "postDate": "2025-06-20T08:57:27.600000",
      "content": "<p>This is a really well-explained and inspiring approach—especially for someone like me who's still building up experience. It’s great to see how logistic regression, when combined with smart feature engineering like categorical combinations and dataset augmentation, can still produce strong results. I found the use of adjusted mutual information and one-hot encoding particularly insightful. Thanks for demonstrating how much can be achieved with a simple model and thoughtful design.</p>"
    },
    {
      "id": 3228263,
      "authorName": "Shamanthak Reddy Mallu",
      "votes": 0,
      "postDate": "2025-06-20T02:15:38.933000",
      "content": "<p>what is the public score you are getting ?<br>\ni have tried logistic regression too.. was getting a public score of around 0.29</p>"
    },
    {
      "id": 3228047,
      "authorName": "gowtham-dd",
      "votes": 0,
      "postDate": "2025-06-19T15:51:20.190000",
      "content": "<p>Since logistic regression inherently assumes no feature interactions, does the inclusion of all 2-combinations and top mutual-info-based 3-combinations effectively turn it into a shallow polynomial model?</p>"
    },
    {
      "id": 3228379,
      "authorName": "broccoli beef",
      "votes": 0,
      "postDate": "2025-06-20T04:47:59.560000",
      "content": "<p>If \\(\\mathcal{C}\\) is a selection of combinations, the model is<br>\n\\[<br>\nf(x)=\\text{softmax}\\left(w_0+\\sum_ig_i(x_i)+\\sum_{c\\in\\mathcal{C}}g_c(x_{c_1},\\ldots,x_{c_{|c|}})\\right)<br>\n\\]<br>\nwhere the \\(g\\)'s are arbitrary functions. It is not the same as <code>PolynomialFeatures</code> in scikit-learn or any \"polynomial models\" in any sense.</p>"
    },
    {
      "id": 3227452,
      "authorName": "",
      "votes": 1,
      "postDate": "2025-06-19T02:55:59.013000",
      "content": ""
    },
    {
      "id": 3232431,
      "authorName": "Shubham Veer",
      "votes": 0,
      "postDate": "2025-06-25T17:41:06.763000",
      "content": "<p>feel free to check out my notebook for more insights and feature engineering ideas that helped improve performance.</p>"
    },
    {
      "id": 3233902,
      "authorName": "Vu-Linh Nguyen",
      "votes": 1,
      "postDate": "2025-06-27T09:59:24.060000",
      "content": "<p>Very helpful. Thank you 👏</p>"
    },
    {
      "id": 3233371,
      "authorName": "Aryan Sukhadia",
      "votes": 2,
      "postDate": "2025-06-26T18:01:13.057000",
      "content": "<p>Thank you for sharing</p>"
    },
    {
      "id": 3231874,
      "authorName": "Biswajit",
      "votes": 0,
      "postDate": "2025-06-25T04:48:30.467000",
      "content": "<p>Thank you  it's Very helpful </p>"
    }
  ],
  "index": {
    "id": "585144",
    "title": "What if you can only use logistic regression...",
    "authorName": "broccoli beef",
    "commentCount": "23",
    "votes": "49",
    "postDate": "2025-06-18 08:41:59.611000"
  }
}