{"cells":[{"cell_type":"markdown","id":"cell-00","metadata":{},"source":"# Deadline Week: 0.940x-1.448x by Venue\n\n## tl;dr\n\nA competition deadline is **not one portable traffic cliff**. In a public Meta Kaggle snapshot,\nthe first post-deadline week's competition-level median cumulative-view ratio ranged from **0.940x\nto 1.448x** across Kaggle's five `HostSegmentTitle` groups. The pooled 1.051x number hides that\nsign change.\n\nUse the checker below before discarding a finished venue or assuming its attention persists.\n**Copy & Edit**, change `WINDOW_DAYS`, `MIN_NOTEBOOK_AGE_DAYS`, and `MIN_SIDE_N`, then point\n`COMPETITION_SLUG` at your own competition. The notebook writes the complete reproducible table to\n`deadline_window_results.csv`."},{"cell_type":"markdown","id":"cell-01","metadata":{},"source":"## Context & Methods\n\n### Key assumptions\n\n- Unit: one competition. Notebook medians are computed *inside* each competition before ratios are\n  aggregated, so a large venue cannot dominate by supplying more notebooks.\n- Comparison: symmetric windows immediately before and after `DeadlineDate`; both sides need at\n  least `MIN_SIDE_N` currently attributed public notebooks.\n- Outcome: cumulative `TotalViews` at one snapshot, **not** a longitudinal accrual rate.\n- Population: surviving kernels whose **current** version names a competition source. Deleted or\n  formerly attributed notebooks are absent.\n- Interpretation: descriptive association. Post-deadline solution notebooks differ from\n  pre-deadline notebooks in content and author intent, so the deadline is not a causal treatment."},{"cell_type":"code","id":"cell-02","execution_count":null,"metadata":{},"outputs":[],"source":"import glob\nimport time\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nSMOKE = False\nSNAPSHOT_DATE = pd.Timestamp(\"2026-08-26 12:00:00\")\nMIN_DEADLINE_DATE = pd.Timestamp(\"2020-01-01\")\nWINDOW_DAYS = [7, 30, 60, 90]          # <-- change the window\nMIN_NOTEBOOK_AGE_DAYS = [90, 365, 730] # <-- change the maturity floor\nMIN_SIDE_N = 5                         # <-- change required notebooks per side\nPRIMARY_WINDOW = 7\nPRIMARY_MIN_AGE = 365\nCOMPETITION_SLUG = \"rsna-knee-abnormality-detection\"  # <-- point at your venue\n\ndef at_or_before_snapshot(data):\n    \"\"\"Keep the population aligned to the frozen snapshot printed above.\"\"\"\n    return data[data[\"MadePublicDate\"] <= SNAPSHOT_DATE].copy()\n\ndef find(pattern):\n    hits = sorted(glob.glob(f\"/kaggle/input/**/{pattern}\", recursive=True))\n    assert hits, f\"{pattern} not found -- mount kaggle/meta-kaggle\"\n    return hits[0]\n\nKERNELS_PATH = find(\"Kernels.csv\")\nSOURCES_PATH = find(\"KernelVersionCompetitionSources.csv\")\nCOMPETITIONS_PATH = find(\"Competitions.csv\")\nprint(\"mode:\", \"SMOKE\" if SMOKE else \"FULL\")\nprint(\"snapshot cutoff:\", SNAPSHOT_DATE)\nprint(\"paths:\", KERNELS_PATH, SOURCES_PATH, COMPETITIONS_PATH, sep=\"\\n  \")"},{"cell_type":"markdown","id":"cell-03","metadata":{},"source":"## Data\n\n### Load only the columns used by the calculation"},{"cell_type":"code","id":"cell-04","execution_count":null,"metadata":{},"outputs":[],"source":"t0 = time.time()\n\ncompetition_columns = [\"Id\", \"Slug\", \"HostSegmentTitle\", \"DeadlineDate\"]\nkernel_columns = [\"Id\", \"CurrentKernelVersionId\", \"MadePublicDate\", \"TotalViews\"]\nsource_columns = [\"KernelVersionId\", \"SourceCompetitionId\"]\n\nif SMOKE:\n    # Execution test: touch every real path/schema, then run the same analysis functions on a\n    # deterministic fixture. The full 644 MB source scan is deliberately reserved for FULL.\n    competitions_raw = pd.read_csv(COMPETITIONS_PATH, usecols=competition_columns, nrows=2_000)\n    kernels_raw = pd.read_csv(KERNELS_PATH, usecols=kernel_columns, nrows=10_000)\n    sources_raw = pd.read_csv(SOURCES_PATH, usecols=source_columns, nrows=20_000)\n    print(\"smoke source rows:\", len(kernels_raw), len(sources_raw), len(competitions_raw))\nelse:\n    competitions_raw = pd.read_csv(\n        COMPETITIONS_PATH, usecols=competition_columns,\n        parse_dates=[\"DeadlineDate\"], low_memory=False,\n    )\n    kernels_raw = pd.read_csv(\n        KERNELS_PATH, usecols=kernel_columns,\n        parse_dates=[\"MadePublicDate\"], low_memory=False,\n    )\n    valid_kernels_all = kernels_raw[\n        kernels_raw[\"MadePublicDate\"].notna()\n        & kernels_raw[\"CurrentKernelVersionId\"].notna()\n    ].copy()\n    excluded_after_snapshot = int(\n        (valid_kernels_all[\"MadePublicDate\"] > SNAPSHOT_DATE).sum()\n    )\n    valid_kernels = at_or_before_snapshot(valid_kernels_all)\n    valid_kernels[\"CurrentKernelVersionId\"] = valid_kernels[\"CurrentKernelVersionId\"].astype(\"int64\")\n    current_version_ids = valid_kernels[\"CurrentKernelVersionId\"].unique()\n\n    source_parts = []\n    for source_chunk in pd.read_csv(SOURCES_PATH, usecols=source_columns, chunksize=500_000):\n        source_parts.append(source_chunk[source_chunk[\"KernelVersionId\"].isin(current_version_ids)])\n    sources_raw = pd.concat(source_parts, ignore_index=True)\n\n    joined = valid_kernels.merge(\n        sources_raw,\n        left_on=\"CurrentKernelVersionId\", right_on=\"KernelVersionId\",\n        how=\"inner\", validate=\"one_to_many\",\n    ).merge(\n        competitions_raw,\n        left_on=\"SourceCompetitionId\", right_on=\"Id\",\n        how=\"inner\", suffixes=(\"\", \"_competition\"), validate=\"many_to_one\",\n    )\n    joined[\"days_after_close\"] = (\n        joined[\"MadePublicDate\"] - joined[\"DeadlineDate\"]\n    ).dt.total_seconds() / 86_400\n    joined[\"age_days\"] = (\n        SNAPSHOT_DATE - joined[\"MadePublicDate\"]\n    ).dt.total_seconds() / 86_400\n\nprint(f\"load + join: {time.time() - t0:.1f}s\")"},{"cell_type":"markdown","id":"cell-05","metadata":{},"source":"### Validate grain, dates, and join behavior before reading views"},{"cell_type":"code","id":"cell-06","execution_count":null,"metadata":{},"outputs":[],"source":"if SMOKE:\n    # Ten notebooks on each side for each segment; ratios are deliberately heterogeneous.\n    # One extra post-cutoff row proves the frozen-snapshot filter is exercised by the smoke.\n    fixture_rows = []\n    fixture_ratios = {\n        \"Analytics\": 0.94, \"Community\": 1.01, \"Featured\": 1.45,\n        \"Playground\": 1.03, \"Research\": 1.18,\n    }\n    for segment_index, (segment, ratio) in enumerate(fixture_ratios.items()):\n        for competition_index in range(2):\n            competition_id = 90_000 + 10 * segment_index + competition_index\n            deadline = pd.Timestamp(\"2024-01-15\")\n            for side, multiplier in ((\"pre\", 1.0), (\"post\", ratio)):\n                for notebook_index in range(5):\n                    offset = -(notebook_index + 1) if side == \"pre\" else notebook_index + 1\n                    fixture_rows.append({\n                        \"Id\": competition_id * 100 + notebook_index + (50 if side == \"post\" else 0),\n                        \"SourceCompetitionId\": competition_id,\n                        \"HostSegmentTitle\": segment,\n                        \"DeadlineDate\": deadline,\n                        \"MadePublicDate\": deadline + pd.Timedelta(days=offset),\n                        \"TotalViews\": 100.0 * multiplier,\n                        \"days_after_close\": float(offset),\n                        \"age_days\": float((SNAPSHOT_DATE - (deadline + pd.Timedelta(days=offset))).days),\n                        \"Slug\": f\"smoke-{segment.lower()}-{competition_index}\",\n                    })\n    fixture_rows.append({\n        \"Id\": 99_999_999,\n        \"SourceCompetitionId\": 99_999,\n        \"HostSegmentTitle\": \"Analytics\",\n        \"DeadlineDate\": pd.Timestamp(\"2026-08-20\"),\n        \"MadePublicDate\": SNAPSHOT_DATE + pd.Timedelta(days=1),\n        \"TotalViews\": 999_999.0,\n        \"days_after_close\": 7.5,\n        \"age_days\": -1.0,\n        \"Slug\": \"smoke-excluded-after-snapshot\",\n    })\n    joined_unfiltered = pd.DataFrame(fixture_rows)\n    excluded_after_snapshot = int(\n        (joined_unfiltered[\"MadePublicDate\"] > SNAPSHOT_DATE).sum()\n    )\n    joined = at_or_before_snapshot(joined_unfiltered)\n    quality = {\n        \"kernel_id_duplicates\": 0,\n        \"mapping_row_duplicates\": 0,\n        \"multi_competition_current_versions\": 0,\n        \"negative_views\": 0,\n        \"excluded_after_snapshot\": excluded_after_snapshot,\n        \"future_public_dates\": 0,\n        \"analyzed_rows\": len(joined),\n        \"deadline_boundary_rows\": 0,\n        \"competition_join_coverage\": 1.0,\n    }\nelse:\n    valid_kernel_count = int(\n        (kernels_raw[\"MadePublicDate\"].notna() & kernels_raw[\"CurrentKernelVersionId\"].notna()).sum()\n    )\n    quality = {\n        \"kernel_id_duplicates\": int(kernels_raw[\"Id\"].duplicated().sum()),\n        \"mapping_row_duplicates\": int(sources_raw.duplicated().sum()),\n        \"multi_competition_current_versions\": int(\n            (sources_raw.groupby(\"KernelVersionId\").size() > 1).sum()\n        ),\n        \"negative_views\": int((kernels_raw[\"TotalViews\"] < 0).sum()),\n        \"excluded_after_snapshot\": excluded_after_snapshot,\n        \"future_public_dates\": int((joined[\"age_days\"] < 0).sum()),\n        \"analyzed_rows\": len(joined),\n        \"deadline_boundary_rows\": int((joined[\"days_after_close\"] == 0).sum()),\n        \"competition_join_coverage\": joined[\"Id\"].nunique() / valid_kernel_count,\n    }\n\nprint(pd.Series(quality).to_string())\nassert quality[\"kernel_id_duplicates\"] == 0\nassert quality[\"mapping_row_duplicates\"] == 0\nassert quality[\"negative_views\"] == 0\nassert quality[\"future_public_dates\"] == 0\nif SMOKE:\n    assert quality[\"excluded_after_snapshot\"] == 1\n    assert quality[\"analyzed_rows\"] == 100\nelse:\n    assert 1 <= quality[\"excluded_after_snapshot\"] <= 1_000_000\nassert quality[\"competition_join_coverage\"] > 0.10\nassert joined[\"DeadlineDate\"].notna().all()"},{"cell_type":"markdown","id":"cell-07","metadata":{},"source":"## Results\n\n### Compute competition-level ratios for every registered arm"},{"cell_type":"code","id":"cell-08","execution_count":null,"metadata":{},"outputs":[],"source":"def competition_ratios(data, window_days, min_age_days):\n    eligible = data[\n        (data[\"DeadlineDate\"] >= MIN_DEADLINE_DATE)\n        & (data[\"age_days\"] >= min_age_days)\n        & data[\"days_after_close\"].between(-window_days, window_days)\n    ].copy()\n    eligible[\"side\"] = np.where(eligible[\"days_after_close\"] >= 0, \"post\", \"pre\")\n\n    side_stats = eligible.groupby([\"SourceCompetitionId\", \"side\"]).agg(\n        n=(\"Id\", \"size\"),\n        median_views=(\"TotalViews\", \"median\"),\n    ).reset_index()\n    paired = side_stats.pivot(\n        index=\"SourceCompetitionId\", columns=\"side\", values=[\"n\", \"median_views\"]\n    ).dropna()\n    paired.columns = [f\"{metric}_{side}\" for metric, side in paired.columns]\n\n    competition_meta = eligible.groupby(\"SourceCompetitionId\").agg(\n        segment=(\"HostSegmentTitle\", \"first\"),\n        deadline=(\"DeadlineDate\", \"first\"),\n        slug=(\"Slug\", \"first\"),\n    )\n    paired = paired.join(competition_meta)\n    paired = paired[\n        (paired[\"n_pre\"] >= MIN_SIDE_N) & (paired[\"n_post\"] >= MIN_SIDE_N)\n    ].copy()\n    paired[\"ratio\"] = paired[\"median_views_post\"] / paired[\"median_views_pre\"].replace(0, np.nan)\n    paired[\"deadline_year\"] = pd.to_datetime(paired[\"deadline\"]).dt.year\n    paired[\"window_days\"] = window_days\n    paired[\"min_age_days\"] = min_age_days\n    return paired.reset_index()\n\narm_rows = []\narm_tables = {}\nfor min_age_days in MIN_NOTEBOOK_AGE_DAYS:\n    for window_days in WINDOW_DAYS:\n        arm = competition_ratios(joined, window_days, min_age_days)\n        arm_tables[(window_days, min_age_days)] = arm\n        arm_rows.append({\n            \"window_days\": window_days,\n            \"min_age_days\": min_age_days,\n            \"eligible_competitions\": len(arm),\n            \"median_post_pre_ratio\": arm[\"ratio\"].median(),\n            \"share_post_above_pre\": (arm[\"ratio\"] > 1).mean(),\n        })\n\narm_summary = pd.DataFrame(arm_rows)\nassert len(arm_summary) == len(WINDOW_DAYS) * len(MIN_NOTEBOOK_AGE_DAYS)\nassert np.isfinite(arm_summary[\"median_post_pre_ratio\"]).all()\nprint(arm_summary.to_string(index=False, float_format=lambda x: f\"{x:.3f}\"))"},{"cell_type":"markdown","id":"cell-09","metadata":{},"source":"### The pooled number is not the result—the segment interaction is"},{"cell_type":"code","id":"cell-10","execution_count":null,"metadata":{},"outputs":[],"source":"primary = arm_tables[(PRIMARY_WINDOW, PRIMARY_MIN_AGE)]\nsegment_summary = primary.groupby(\"segment\").agg(\n    eligible_competitions=(\"SourceCompetitionId\", \"size\"),\n    median_post_pre_ratio=(\"ratio\", \"median\"),\n    share_post_above_pre=(\"ratio\", lambda values: (values > 1).mean()),\n).reset_index().sort_values(\"median_post_pre_ratio\")\n\nprint(segment_summary.to_string(index=False, float_format=lambda x: f\"{x:.3f}\"))\n\nif SMOKE:\n    expected_fixture = {\n        \"Analytics\": 0.94,\n        \"Community\": 1.01,\n        \"Featured\": 1.45,\n        \"Playground\": 1.03,\n        \"Research\": 1.18,\n    }\n    observed = segment_summary.set_index(\"segment\")\n    assert set(observed.index) == set(expected_fixture)\n    for segment, expected_ratio in expected_fixture.items():\n        assert observed.loc[segment, \"eligible_competitions\"] == 2\n        assert abs(observed.loc[segment, \"median_post_pre_ratio\"] - expected_ratio) <= 0.001\nelse:\n    registered_bands = {\n        \"Analytics\": (0.85, 1.05),\n        \"Community\": (0.95, 1.10),\n        \"Featured\": (1.25, 1.65),\n        \"Playground\": (0.95, 1.15),\n        \"Research\": (1.05, 1.35),\n    }\n    observed = segment_summary.set_index(\"segment\")\n    assert set(registered_bands).issubset(observed.index)\n    for segment, (lower, upper) in registered_bands.items():\n        ratio = observed.loc[segment, \"median_post_pre_ratio\"]\n        n_competitions = observed.loc[segment, \"eligible_competitions\"]\n        assert lower <= ratio <= upper, (segment, ratio, lower, upper)\n        assert n_competitions >= 10, (segment, n_competitions)\n\nfig, ax = plt.subplots(figsize=(8, 4.3))\ncolors = [\"#D55E00\" if value < 1 else \"#0072B2\" for value in segment_summary[\"median_post_pre_ratio\"]]\nax.barh(segment_summary[\"segment\"], segment_summary[\"median_post_pre_ratio\"], color=colors)\nax.axvline(1.0, color=\"black\", linewidth=1, linestyle=\"--\", label=\"same median views\")\nax.set_xlabel(\"median within-competition post/pre cumulative-view ratio\")\nax.set_title(\"First post-deadline week: the direction changes by venue class\")\nax.legend(loc=\"lower right\")\nplt.tight_layout()\nplt.show()"},{"cell_type":"markdown","id":"cell-11","metadata":{},"source":"The range is the useful result. `Analytics` is below 1 while the other four groups are above 1 in\nthis snapshot; `Featured` is much higher than the rest. Pooling them into 1.051x would turn a venue\nmix into a universal rule—the ecological fallacy.\n\nThis does **not** show that deadlines create or preserve attention. Final-solution writeups,\nhost promotion, author mix, competition size, and topic all change around a deadline. It shows\nthat a reader should measure their venue class rather than inherit a pooled cutoff."},{"cell_type":"markdown","id":"cell-12","metadata":{},"source":"### Time-window and year checks"},{"cell_type":"code","id":"cell-13","execution_count":null,"metadata":{},"outputs":[],"source":"year_summary = primary.groupby(\"deadline_year\").agg(\n    eligible_competitions=(\"SourceCompetitionId\", \"size\"),\n    median_post_pre_ratio=(\"ratio\", \"median\"),\n).reset_index()\nprint(\"deadline-year strata\")\nprint(year_summary.to_string(index=False, float_format=lambda x: f\"{x:.3f}\"))\n\nall_results = pd.concat(arm_tables.values(), ignore_index=True)\nall_results.to_csv(\"deadline_window_results.csv\", index=False)\nsegment_summary.to_csv(\"deadline_week_segment_summary.csv\", index=False)\nprint(f\"wrote deadline_window_results.csv: {len(all_results):,} competition-arm rows\")\nprint(\"wrote deadline_week_segment_summary.csv\")\nif SMOKE:\n    print(\"SMOKE_OK\")"},{"cell_type":"markdown","id":"cell-14","metadata":{},"source":"## Takeaways\n\n### Check your own competition"},{"cell_type":"code","id":"cell-15","execution_count":null,"metadata":{},"outputs":[],"source":"own = primary[primary[\"slug\"] == COMPETITION_SLUG]\nif len(own):\n    print(own[[\n        \"slug\", \"segment\", \"n_pre\", \"n_post\", \"median_views_pre\",\n        \"median_views_post\", \"ratio\",\n    ]].to_string(index=False, float_format=lambda x: f\"{x:.3f}\"))\nelse:\n    print(\n        f\"{COMPETITION_SLUG!r} does not yet have >= {MIN_SIDE_N} public notebooks on both \"\n        f\"sides of a {PRIMARY_WINDOW}-day deadline window in this snapshot. \"\n        \"Copy & Edit after its deadline, or change COMPETITION_SLUG.\"\n    )"},{"cell_type":"markdown","id":"cell-16","metadata":{},"source":"1. **Do not treat the deadline as a universal traffic cutoff.** The first-week direction differs\n   by public venue class in the same snapshot.\n2. **Use the narrowest window your sample supports.** The notebook prints every registered window\n   and age floor; do not choose the friendliest arm after looking.\n3. **A current snapshot is not an accrual curve.** To claim that notebooks *keep gaining* views\n   requires historical snapshots of the same kernels. This notebook deliberately does not make\n   that claim.\n\nSource: Kaggle's public `kaggle/meta-kaggle` dataset, files `Kernels.csv`,\n`KernelVersionCompetitionSources.csv`, and `Competitions.csv`, snapshot cutoff shown in the first\ncell. The code uses only current-version competition attribution and records its limitations above."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}