{"cells":[{"cell_type":"markdown","id":"010a37d6","metadata":{"papermill":{"duration":0.011701,"end_time":"2026-09-12T11:17:37.645918+00:00","exception":false,"start_time":"2026-09-12T11:17:37.634217+00:00","status":"completed"},"tags":[]},"source":"# 🩺 RSNA Knee Abnormality Detection: A Diagnostic Quest\n\nWelcome to the clinic of the future!  \nYou are a radiologist — but not an ordinary one. You are armed with neural networks.  \nYour task is to make **12 key diagnoses** from knee MRI studies, using both images and, during training, the accompanying radiology reports.\n\n---\n\n## 🧩 The Core Challenge\n\nEach study is a **puzzle** made of several MRI series — slices acquired in different planes and with different contrast weightings.  \nYou must assemble the full picture and answer 12 questions: Is there a ligament tear? Meniscal damage? Osteoarthritis? Inflammation? A cyst? A fracture?\n\n**Metric:** macro-averaged ROC AUC across all 12 labels.  \n**Key twist:** during training, you have access not only to the images but also to the **radiologists’ free-text reports**. At inference time, reports are absent — so you must learn to squeeze every bit of signal from them while training.\n\n---\n\n## 🧠 Analogy: “Training a Virtual Radiology Resident”\n\nImagine you are teaching a new resident:\n\n- **DICOM series** — different examination tools: sagittal, coronal, and axial views.\n- **Fluid-sensitive sequences** (T2, PD, STIR) — a flashlight that highlights fluid, edema, and tears.\n- **train.csv** — patient charts with partially completed diagnoses.\n- **train_series.csv** — a description of each scan: plane, fat suppression, fluid sensitivity.\n- **Radiology reports** — notes from an experienced mentor who sometimes writes “ACL tear” and sometimes “intact menisci.”\n\nYour resident (the model) must learn to make diagnoses even when the mentor is not there (at test time).\n\n---\n\n## 🏆 Difficulty Tiers of the Diagnoses\n\n| Tier | Diagnoses | Why They Are Hard | What Helps |\n|------|-----------|-------------------|------------|\n| 🟢 **Easier** | Baker's cyst, Effusion | Clearly visible on fluid-sensitive sequences | A single sagittal series |\n| 🟡 **Moderate** | MCL, Medial/Lateral Meniscus, PF OA | Need assessment across multiple planes | 2–3 series + attention |\n| 🔴 **Hard** | ACL, Synovitis, Contusion, Fracture | May be subtle; require broader context | Multi-plane aggregation, 2.5D/3D inputs |\n| 🟣 **Very Hard** | Medial OA, Lateral OA | Cartilage thickness assessment; need thin slices | Specialized models + post-processing |\n\n---\n\n## 💡 The Main Takeaway\n\n**This competition is won by whoever does three things best:**\n\n1. **Select the right series** — not all 500 GB of data is useful.\n2. **Exploit the reports** to expand weak labels — they are free supervision.\n3. **Aggregate carefully** across planes and slices.\n\nNo magic — only solid data engineering and thoughtful modeling.\n\n---\n\n**Let’s go!** 🚀  \nWe start with a simple diagnostic tool and gradually turn it into a radiology expert."},{"cell_type":"markdown","id":"d0861b90","metadata":{"papermill":{"duration":0.01004,"end_time":"2026-09-12T11:17:37.666253+00:00","exception":false,"start_time":"2026-09-12T11:17:37.656213+00:00","status":"completed"},"tags":[]},"source":"# 📊 Evolution of Public Solutions: 0.937 → 0.94 → 0.941\n\n## 🥉 Work 1 — 0.937\n**Head and Shoulders, Knees and Toes**\n\n🔗 https://www.kaggle.com/code/prvsiyan/head-and-shoulders-knees-and-toes\n\n- CoAtNet branch expanded to **4 arms**: v5, v10, v5-reverse, v8.\n- Reverse-flip adds diversity without new training.\n- Outer CoAtNet weight becomes **global 0.60** instead of per-target.\n\n---\n\n## 🥈 Work 2 — 0.94\n**RSNA Knee Hybrid Raptor — Lateral Meniscus R100**\n\n🔗 https://www.kaggle.com/code/renta0426/rsna-knee-hybrid-raptor-lateral-meniscus-r100\n\n- Starts from the exact 0.939 parent (reproduced 1-to-1).\n- Changes **only one label**: Lateral Meniscus.\n- Outer routing for Lateral Meniscus: Transformer/Rad **0.00** / strengthened Raptor **1.00**.\n- All other 11 labels remain **bit-for-bit equal** to the parent at 0.40 / 0.60.\n\n> **Key idea:** isolate one problematic target and give it the full weight of the strongest branch.\n\n---\n\n## 🥇 Work 3 — 0.941\n**Bend the Knee to the Dinosaurs (ALL PUBLIC)**\n\n🔗 https://www.kaggle.com/code/mattiaangeli/bend-the-knee-to-the-dinosaurs\n\n- Keeps the 4-arm Raptor from Work 1.\n- Adds **residual-gated CoAtNet** (e4/e6/e8) inside the Raptor branch.\n- Inner blend: public Raptor 0.60 / residual-CoAt 0.40.\n- Adds a public **CoAtNet MRI reader**: 384 px adjacent-slice triplets → CoAtNet RMLP-2 backbone.\n- Finding-specific gated spatial head + second attention stage pools windows across the study.\n- Updated CoAtNet arm weights: maxspan-v5 = **0.60**, reverse = **0.10**.\n- Per-target outer routing: ACL **0.75**, Medial Meniscus **0.80**, Lateral Meniscus **1.00**, Lateral OA **0.75**, Fracture **0.75**.\n\n> **Key idea:** strengthen the Raptor arm with a genuinely different visual inductive bias and let each finding lean on whichever branch predicts it best.\n\n---\n\n## 🧪 Work 4 — our v15\n**Extension of Bend the Knee**\n\nBuilding directly on Work 3, we added one targeted change:\n\n1. **Reverse-flip TTA on RadImageNet V48-pass2.** Each slice is evaluated both as-is and horizontally flipped; the two rank vectors are averaged before entering the calibration stack. This adds TTA diversity at the RadImageNet stage without any new training and without touching the DINOv2, DINOv3 or CoAtNet branches.\n\n> **Key idea:** even a single additional TTA pass on one branch of a multi-branch ensemble can shift the calibrated output, because the calibration transformer reads from the pre-blend rank vectors.\n\n---\n\n## 🧪 Work 5 — our v16\n\nFour changes on top of v15. Three of them change nothing on a clean run and only\ndecide what happens on a bad one; one is a correctness fix to v15's own TTA.\n\n1. **The mirror TTA now mirrors the labels too.** `build_cache()` calls\n   `normalise_laterality()`, so every study reaches the Rad cache already turned\n   to one side. v15's `pixels[..., ::-1]` therefore does not produce a second\n   view of the same knee — it produces the other knee, in which what was medial\n   is lateral. The head was fitted on one orientation, so for the four\n   medial/lateral targets the flipped pass answered the *partner* question and\n   was averaged in as if it had answered the original one. v16 swaps\n   Medial↔Lateral Meniscus and Medial↔Lateral OA back before averaging, and\n   drops MCL and Baker's from the flip entirely: the MCL mirrors to the LCL and\n   a Baker's cyst is posteromedial by definition, so neither has a scored\n   target on the other side. Simulated on the competition's own label\n   prevalence and co-occurrence, v15's flip is worth **−0.009 to −0.017** macro\n   AUC on the Rad branch at every level of TTA decorrelation; the realigned\n   version is positive at every level.\n\n2. **One unreadable study no longer kills the run.** `_micro()` returns NaN for\n   a study with no usable slot, and the DINOv3 blend asserted every value\n   finite. A single such study raised in cell 40, which meant the Rad and\n   Raptor branches never ran and the submitted file was the DINOv2-only\n   output. v16 keeps the parent's rank for studies the branch could not read,\n   and skips the blend outright if coverage falls below 98%.\n\n3. **A Rad failure no longer takes the Raptor branch with it.** `_rad_main()`\n   raises on a dozen integrity conditions. Every one of them is a reason to\n   distrust the Rad branch; none is a reason to skip Raptor, which carries\n   0.60–1.00 of every target in the final routing. v16 restores the pre-Rad\n   submission and carries on.\n\n4. **A missing Raptor arm drops that arm, not the run**, and a failed CoAt\n   substitution falls back to the public four-arm Raptor — which is what the\n   0.939 parent used.\n\n> **Key idea:** in a stack this deep, the largest expected-value item is no\n> longer a weight. It is the difference between the run that finishes and the\n> run that dies in cell 40 and submits a 0.899 file.\n\n### What was measured and rejected\n\n- **Rank-averaging the four CoAtNet arms instead of probability-averaging.** The\n  nominal 0.60/0.10/0.10/0.20 is not the weighting probability averaging\n  actually realises (it drifts to ~0.54/0.14/0.06/0.25 under plausible\n  calibration spread), but the arms are correlated enough that the macro effect\n  is within ±0.0006 either way, and not signed. Left alone.\n- **Blending the two DINOv2 aggregations the run computes and discards**\n  (`submission_native_v38.csv`, `submission_legacy_fold_blend.csv`). Free in GPU\n  time, but at the ρ≈0.98 you get from re-pooling the same 24 checkpoints it\n  only pays if the discarded aggregations are as strong as the promoted one, and\n  turns negative as soon as they are 0.005 AUC weaker. Not deployed.\n- **Retuning the per-target outer weights.** On a ~390-study public split a\n  single target's AUC has a standard deviation of 0.015–0.031, which is\n  ±0.0024–0.0051 of 95% band on the macro number those weights were chosen\n  against. Fitting a weight on that split costs 0.000–0.0013 private AUC versus\n  leaving it at 0.60. The weights are kept exactly as v15 had them.\n\n---\n\n## 🧪 Work 6 — our v17\n\nv16 closed cells 40, 42 and 44. v17 closes the one upstream of all of them,\npins the one dependency whose drift is silent, and gives back the wall clock\nthat four branches were spending on the same DICOM files. **No weight, no\nalpha and no routing constant moves.** On a clean run the submitted file is\nbit-identical to v16's; what changes is how long it takes and what happens\nwhen something goes wrong.\n\n1. **Cell 30 is guarded.** `infer_from_package()` was called bare. Inside it,\n   `combine_public_members_by_fold()` raises on anything but five complete\n   folds and `_combine()` raises when a target has no vote — and any of those\n   took cell 30 down, and with it cells 32–44, submitting the DINOv2-only\n   file. That is the same ~0.899 cliff v16 closed at cell 40, one cell\n   earlier. Now it degrades: the last banked file stays, and the DINOv3,\n   RadImageNet and Raptor branches still run.\n\n2. **A run that is not the tuned recipe says so.** The promotion at cell 30\n   depends on `len(public_frontier_members) == len(members)`, and **three**\n   independent paths break it: a window-truncated member under time pressure,\n   a member dropped after failing on both devices, and a degenerate member.\n   Two of them need nothing more than one unreadable study. When it happens,\n   `submission.csv` holds the native blend — a lineage nobody has scored —\n   and the Rad calibrator and the outer routing then apply constants fitted\n   against the other one. v16 logged that as `skipped safely`. v17 names the\n   missing members, warns again at the outer routing, and writes\n   `submission_profile_status.csv`.\n\n3. **timm is pinned to the wheel that was already mounted.** `timm` is\n   load-bearing twice — the five DINOv3 folds and all four Raptor arms — and\n   was imported bare from whatever the base image carried, while\n   `timm-1.0.22-py3-none-any.whl` sat unused inside `knee-mri-fold-weights`.\n   cv2 is pinned and version-asserted two cells down; timm now is too, with a\n   fallback to the image's copy if the wheel cannot be installed.\n\n4. **The encoder is counted, not assumed.** `assert not [k for k in missing\n   if not k.startswith('enc.')]` is blind to the entire backbone. Renames and\n   shape changes still fail loudly — a rename puts the checkpoint's orphaned\n   keys in `unexpected`, and `strict=False` never suppresses a size mismatch\n   — but *additive* drift is silent: an installed timm that builds a module\n   the checkpoint does not carry leaves that tensor at its random init, and\n   the result enters the submission at `A5_W = 0.45`. There is no fingerprint\n   check in this branch and a damaged encoder still emits finite, confident\n   probabilities. v17 logs `enc N/M` per fold and, on a shortfall, skips the\n   branch into v16's own tested fallback instead of blending it in.\n\n5. **Four DICOM passes become three, and the jitter pass is gone.** Arms 0 and\n   2 are the *same checkpoint* at the same `img`, the same slots and the same\n   span — arm 2 only flips the finished window stack — and arm 2 was re-reading\n   the volume arm 0 had just built. They now share one pass. Jitter TTA was\n   doubling the DINOv2 branch to reach `submission_native_v38.csv`, which\n   nothing reads: the promoted public frontier is pooled from the *un-jittered*\n   view. And the DINOv3 chunk loop now decodes the next chunk while the GPU\n   runs the current one, instead of the pool and the GPU taking turns.\n\n6. **Housekeeping the audit turned up.** `find_weights()` no longer *raises* on\n   a foreign incomplete `manifest.json` — two of the seventeen mounts are other\n   people's kernel outputs and can gain one without notice; `find_weight_file()`\n   prunes `test_series`/`train_series` by name like the other six finders,\n   instead of a `'competition' in path` test that misses one of the two\n   competition layouts; the three never-called head-refit functions are\n   deleted; and `_coat_substitute()` reads `torch.cuda.device_count()` itself\n   rather than letting a child process's `assert` decide silently.\n\n> **Key idea:** every item here is either insurance or wall clock. Both\n> `V16.md` and `MEMORY.md` concluded that the weight space is exhausted, and\n> a 0.94+ attempt needs a genuinely decorrelated component. There is nowhere\n> to put one while four branches only just fit inside the kernel.\n\n### Still open, deliberately\n\n- **Two DINOv2 seeds could probably go.** 20 votes, ~5 effective; at ρ ≈ 0.99\n  the loss is ~0 and the GPU saving is 50%. The premise is read off a\n  `manifest.json` nobody has opened. Download it before touching this.\n- **Spreading the branches over both T4s.** Only the DINOv2 branch uses\n  `DEVS`; DINOv3, Rad and Raptor all run on `cuda:0` while the second card\n  idles — and the CoAt child *requires* both. That is a bigger rewrite than\n  this patch set.\n- **Detaching `dinov2/base/1` and `rsna-knee-llm-labels`.** Both are dead\n  mounts, but they live in the kernel's input list, not in the notebook.\n- **The three laterality conventions, the unmasked Raptor attention, and the\n  periphery each branch discards.** Ceilings, not bugs: each branch is\n  internally consistent with how its checkpoints were fitted, so correcting\n  one at inference would break train/test parity. They bound what this model\n  family can reach; they are notes for the next trained component.\n\n---\n\n## Key Pattern\n\nEach step isolates a single high-impact change:\n\n1. **Diversity** — more CoAtNet arms + reverse-flip.\n2. **Inner strengthening** — residual-gated CoAt.\n3. **Target-specific routing** — give full weight to the best model for a problematic label.\n4. **Per-branch TTA** — reverse-flip on one branch of the ensemble stack.\n\nThis is how public solutions moved from **0.937 → 0.94 → 0.941**, and where we continue to build.\n\n---\n\n## 📚 Credits\n\nWe stand on the shoulders of giants. Our work reproduces, extends, and improves on the public solutions above. Full credit to their authors and contributors.\n\n- **Work 1 (0.937)** — [prvsiyan](https://www.kaggle.com/code/prvsiyan/head-and-shoulders-knees-and-toes)\n- **Work 2 (0.94)** — [renta0426](https://www.kaggle.com/code/renta0426/rsna-knee-hybrid-raptor-lateral-meniscus-r100)\n- **Work 3 (0.941)** — [Mattia Angeli](https://www.kaggle.com/code/mattiaangeli/bend-the-knee-to-the-dinosaurs) + credits: Pilkwang, Sofia Anjenje, Antoine G., prvsiyan, Marwan Mahmoud, dreaddevelopment/Roman Tamrazov, renta.k, Anvith Pothula"},{"cell_type":"markdown","id":"7428324f","metadata":{"papermill":{"duration":0.009692,"end_time":"2026-09-12T11:17:37.685607+00:00","exception":false,"start_time":"2026-09-12T11:17:37.675915+00:00","status":"completed"},"tags":[]},"source":"# 1. 🚀 Инициализация инференса и проверка GPU\n\n## Что делает эта ячейка\n\nГотовит окружение для предсказаний на тестовых данных:\n\n- Настраивает количество потоков CPU.\n- Проверяет доступность CUDA на каждой GPU.\n- Создаёт список устройств `DEVS`, на которых будет выполняться инференс.\n- Фиксирует seed для воспроизводимости.\n- Задаёт все параметры предобработки, слотов и аугментаций.\n\nЗдесь не происходит чтения DICOM или загрузки моделей — только подготовка."},{"cell_type":"code","execution_count":null,"id":"a9eed16a","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:37.707162Z","iopub.status.busy":"2026-09-12T11:17:37.706514Z","iopub.status.idle":"2026-09-12T11:17:50.065164Z","shell.execute_reply":"2026-09-12T11:17:50.064218Z"},"papermill":{"duration":12.371513,"end_time":"2026-09-12T11:17:50.066864+00:00","exception":false,"start_time":"2026-09-12T11:17:37.695351+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"from __future__ import annotations\nimport os\nfor _v in ('OMP_NUM_THREADS', 'OPENBLAS_NUM_THREADS', 'MKL_NUM_THREADS'):\n    os.environ.setdefault(_v, '4')\nimport gc\nimport hashlib\nimport json\nimport re\nimport time\nimport traceback\nimport threading\nfrom concurrent.futures import ThreadPoolExecutor\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\ndef _cuda_execution_probe(index):\n    dev = torch.device(f'cuda:{index}')\n    try:\n        major, minor = torch.cuda.get_device_capability(index)\n        probe = nn.Conv2d(3, 4, kernel_size=3, padding=1).eval().to(dev)\n        with torch.inference_mode():\n            out = probe(torch.zeros((1, 3, 16, 16), device=dev))\n            if tuple(out.shape) != (1, 4, 16, 16):\n                raise RuntimeError(f'unexpected CUDA probe shape {tuple(out.shape)}')\n        torch.cuda.synchronize(index)\n        print(f'cuda:{index} probe PASS (compute {major}.{minor})')\n        del probe, out\n        torch.cuda.empty_cache()\n        return True\n    except Exception as exc:\n        print(f'cuda:{index} probe FAIL ({type(exc).__name__}: {exc}); using CPU fallback')\n        try:\n            torch.cuda.empty_cache()\n        except Exception:\n            pass\n        return False\nDEVS = []\nif torch.cuda.is_available():\n    DEVS = [torch.device(f'cuda:{i}') for i in range(torch.cuda.device_count()) if _cuda_execution_probe(i)]\nif not DEVS:\n    DEVS = [torch.device('cpu')]\nprint(f'devices: {[str(d) for d in DEVS]}')\nT0 = time.time()\nSEED = 2026\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\nTARGETS = ['ACL', 'MCL', 'Medial Meniscus', 'Lateral Meniscus', 'Medial OA', 'Lateral OA', 'PF OA', 'Effusion', 'Synovitis', \"Baker's\", 'Contusion', 'Fracture']\nCROP_MM = 130.0\nCACHE_IMG = 336\nGROUP = 3\nN_GROUP_MAX = 1\nCACHE_FRACTION = 0.45\nCACHE_BUDGET_MAX_GB = 24.0\nCACHE_BUDGET_GB = 12.0\nTEST_SHARE = 0.3\nHDR_THREADS = 16\nPIX_THREADS = 12\nORDER_THREADS = 32\nORDER_BUDGET_S = 5400\nRUNS = [{'name': 'r224', 'img': 224}, {'name': 'r336', 'img': 336}]\nEPOCHS = 10\nBATCH_STUDIES = 8\nAUG_ROT_DEG = 8.0\nAUG_SCALE = 0.08\nAUG_SHIFT = 0.05\nAUG_INTENSITY = 0.1\nLAT_MIN_OFFSET_MM = 20.0\nSLICE_BAND = (0.2, 0.8)\nRULES_NATIVE = {'order': 'normal', 'lat': 'centre', 'slot_fallback': False, 'decode_fill': 'nearest'}\nRULES_LEGACY = {'order': 'dominant_axis', 'lat': 'corner_x', 'slot_fallback': True, 'decode_fill': 'zero'}\nRULES = dict(RULES_NATIVE)\nLEGACY_LAT_OFFSET_MM = 5.0\nLR_HEAD = 0.001\nLR_BACKBONE = 8e-06\nUNFREEZE_LAST = 6\nWEIGHT_DECAY = 0.02\nEVAL_BATCH = 8\nTIME_BUDGET = 8.0 * 3600\nSLOTS_RECOVERED = [('SAG_FLUID_FS', 'Sagittal', True, True), ('COR_FLUID_FS', 'Coronal', True, True), ('AX_FLUID_FS', 'Axial', True, True), ('SAG_FLUID_NOFS', 'Sagittal', True, False), ('COR_T1', 'Coronal', False, False), ('SAG_T1', 'Sagittal', False, False)]\nSLOTS_PUBLIC = [('SAG_FLUID', 'Sagittal', None, True), ('COR_FLUID', 'Coronal', None, True), ('AX_FLUID', 'Axial', None, True), ('SAG_STRUCT', 'Sagittal', None, False), ('COR_STRUCT', 'Coronal', None, False), ('AX_STRUCT', 'Axial', None, False)]\nSLOT_SCHEME = os.environ.get('SLOT_SCHEME', 'recovered')\nSLOTS = SLOTS_PUBLIC if SLOT_SCHEME == 'public' else SLOTS_RECOVERED\nN_SLOT = len(SLOTS)\nPOOL_PARTS = {'cls_mean': 2, 'cls_mean_focal': 3}\nSLOT_PRIOR_TABLE = {'ACL': (0, 3, 5), 'MCL': (1, 4), 'Medial Meniscus': (0, 1, 3, 4), 'Lateral Meniscus': (0, 1, 3, 4), 'Medial OA': (1, 4, 5), 'Lateral OA': (1, 4, 5), 'PF OA': (0, 2, 5), 'Effusion': (0, 2), 'Synovitis': (0, 2), \"Baker's\": (0,), 'Contusion': (0, 1, 2), 'Fracture': (0, 1, 2, 4, 5)}\nSLOT_PRIOR_STRENGTH = 0.55\nFATSAT_OPTS = {'FS', 'FATSAT', 'FAT_SAT', 'FSAT'}\n_SEP = re.compile('[_\\\\-.]')\n_FATSAT_RX = re.compile('\\\\bfs\\\\b|fatsat|fat sat|\\\\bstir\\\\b|\\\\bspair\\\\b|\\\\bspir\\\\b|\\\\bwe\\\\b|water excit|\\\\btirm\\\\b|\\\\bsting\\\\b|\\\\bfatsup\\\\b')\n_T1_RX = re.compile('\\\\bt1\\\\b|\\\\bt1w\\\\b')\n_T2_RX = re.compile('\\\\bt2\\\\b|\\\\bt2w\\\\b')\n_PD_RX = re.compile('\\\\bpd\\\\b|\\\\bpdw\\\\b|proton|\\\\bdp\\\\b|dens')"},{"cell_type":"markdown","id":"e341b2b8","metadata":{"papermill":{"duration":0.008109,"end_time":"2026-09-12T11:17:50.084129+00:00","exception":false,"start_time":"2026-09-12T11:17:50.07602+00:00","status":"completed"},"tags":[]},"source":"# 2. 🧭 Вспомогательные функции и планировщик кэша\n\n## Что делает эта ячейка\n\nОпределяет функции, которые используются на протяжении всего инференса:\n\n- `find_root()` — находит корень данных соревнования.\n- `find_dinov2()` — находит DINOv2 модель нужного варианта.\n- `plan_cache()` — рассчитывает, сколько срезов можно хранить в памяти.\n- `read_labels()` — загружает метки из отчётов (не используется в инференсе, но оставлена для совместимости).\n\nТакже сразу вычисляется `N_GROUP` — сколько групп срезов будет загружено для каждой серии.\n\nВ инференсе эти функции обеспечивают правильные пути и распределение памяти."},{"cell_type":"code","execution_count":null,"id":"0da7d8d2","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.102498Z","iopub.status.busy":"2026-09-12T11:17:50.102108Z","iopub.status.idle":"2026-09-12T11:17:50.312472Z","shell.execute_reply":"2026-09-12T11:17:50.311821Z"},"papermill":{"duration":0.221378,"end_time":"2026-09-12T11:17:50.314125+00:00","exception":false,"start_time":"2026-09-12T11:17:50.092747+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"def log(msg):\n    print(f'[{time.time() - T0:7.1f}s] {msg}', flush=True)\n\ndef find_root():\n    for c in [Path('/kaggle/input/competitions/rsna-knee-abnormality-detection'), Path('/kaggle/input/rsna-knee-abnormality-detection'), Path('data'), Path('.')]:\n        if (c / 'test.csv').is_file() and (c / 'test_series').is_dir():\n            return c\n    base = Path('/kaggle/input')\n    if base.is_dir():\n        for depth1 in sorted((p for p in base.iterdir() if p.is_dir())):\n            for cand in [depth1] + sorted((p for p in depth1.iterdir() if p.is_dir())):\n                if (cand / 'test.csv').is_file():\n                    return cand\n    raise FileNotFoundError(f'competition mount not found (cwd {Path.cwd()}); expected a directory holding test.csv and test_series/')\n\ndef find_dinov2(variant='small'):\n    base = Path('/kaggle/input')\n    if not base.is_dir():\n        return None\n    hits = []\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in ('train_series', 'test_series')]\n        if 'config.json' in files and 'dinov2' in root.lower():\n            hits.append(Path(root))\n    for h in hits:\n        if variant in str(h).lower():\n            return h\n    return hits[0] if hits else None\nLABEL_COLS = TARGETS + [t + '__conf' for t in TARGETS]\n\nclass LabelSourceError(RuntimeError):\n    pass\n\ndef find_label_table():\n    base = Path('/kaggle/input')\n    cands = []\n    if base.is_dir():\n        for root, dirs, files in os.walk(base):\n            dirs[:] = [d for d in dirs if d not in ('train_series', 'test_series')]\n            cands += [Path(root) / f for f in files if f.startswith('report_labels') and f.endswith('.csv')]\n    cands += [p for p in (Path('data/derived/report_labels_v2.csv'),) if p.is_file()]\n    for c in cands:\n        try:\n            head = pd.read_csv(c, nrows=1)\n        except Exception:\n            continue\n        if 'StudyInstanceUID' in head.columns and all((t in head.columns for t in TARGETS)):\n            return c\n    return None\n\ndef label_mount_attached():\n    base = Path('/kaggle/input')\n    if not base.is_dir():\n        return False\n    return any(('label' in p.name.lower() for p in base.iterdir() if p.is_dir()))\n\ndef read_labels(train_df):\n    n = len(train_df)\n    lab = pd.DataFrame([extract(r) for r in train_df['Report'].fillna('')])\n    lab['StudyInstanceUID'] = train_df['StudyInstanceUID'].values\n    lab = lab.set_index('StudyInstanceUID')\n    src = find_label_table()\n    if src is None:\n        if label_mount_attached():\n            raise LabelSourceError('LABEL SOURCE: a label dataset is mounted but no usable table was found in it. Falling back to the lexicon here would train on the weaker labels and say so only in a log line, so the run stops instead.')\n        log(f'LABEL SOURCE: lexicon, {n} studies (no table mounted)')\n        return lab\n    tab = pd.read_csv(src).set_index('StudyInstanceUID')\n    missing = [c for c in LABEL_COLS if c not in tab.columns]\n    if missing:\n        raise LabelSourceError(f'LABEL SOURCE: {src} is missing {len(missing)} expected columns (first: {missing[0]!r}). Refusing to fall back silently.')\n    hit = lab.index.intersection(tab.index)\n    if not len(hit):\n        raise LabelSourceError(f'LABEL SOURCE: {src} shares no StudyInstanceUID with train.csv.')\n    log(f'LABEL SOURCE: {src.name} covers {len(hit)} of {n} studies, lexicon for the remaining {n - len(hit)}')\n    lab.loc[hit, LABEL_COLS] = tab.loc[hit, LABEL_COLS].values\n    return lab\nROOT = find_root()\nlog(f'input root: {ROOT}')\nIMG = CACHE_IMG\n\ndef available_gb():\n    try:\n        with open('/proc/meminfo') as fh:\n            info = {k.strip(): v for k, v in (l.split(':', 1) for l in fh if ':' in l)}\n        return int(info['MemAvailable'].split()[0]) / 1024 ** 2\n    except Exception:\n        return CACHE_BUDGET_GB / CACHE_FRACTION\n\ndef plan_cache(n_study, n_test=0):\n    avail = available_gb()\n    budget = min(avail * CACHE_FRACTION, CACHE_BUDGET_MAX_GB)\n    n_total = n_study + max(n_test, int(TEST_SHARE * n_study))\n    per_slice = n_total * N_SLOT * IMG * IMG\n    afford = int(budget * 1024 ** 3 // max(per_slice, 1))\n    groups = max(1, min(N_GROUP_MAX, afford // GROUP))\n    log(f'memory: {avail:.1f} GB available, {budget:.1f} GB to the cache; sizing for {n_study} train + {n_total - n_study} test studies -> {groups} group(s) of {GROUP} = {groups * GROUP} slices per slot' + (f' (wanted {N_GROUP_MAX})' if groups < N_GROUP_MAX else ''))\n    return groups\nN_GROUP = plan_cache(len(pd.read_csv(ROOT / 'train.csv')), len(pd.read_csv(ROOT / 'test.csv')))\nCACHE_SLICES = GROUP * N_GROUP\nlog(f'cache layout: {N_GROUP} groups x {GROUP} slices = {CACHE_SLICES} per slot')"},{"cell_type":"markdown","id":"da0e5de8","metadata":{"papermill":{"duration":0.00835,"end_time":"2026-09-12T11:17:50.331614+00:00","exception":false,"start_time":"2026-09-12T11:17:50.323264+00:00","status":"completed"},"tags":[]},"source":"# 3. 🦴 Чтение DICOM-заголовков и определение латеральности\n\n## Что делает эта ячейка\n\nСчитывает метаданные DICOM для каждой серии без загрузки пикселей:\n\n- `probe()` — открывает один срез из серии и извлекает ключевые теги.\n- `walk()` — обходит все серии train/test и собирает метаданные в DataFrame.\n- `annotate()` — восстанавливает тип последовательности (T1/T2/PD/GRE) и наличие fat suppression.\n- `lat_of()` — определяет сторону колена (левое/правое) по тегу Laterality или геометрии.\n\nЭто необходимо для правильного выбора слотов и нормализации laterality перед инференсом."},{"cell_type":"code","execution_count":null,"id":"a925918b","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.349891Z","iopub.status.busy":"2026-09-12T11:17:50.349618Z","iopub.status.idle":"2026-09-12T11:17:50.37118Z","shell.execute_reply":"2026-09-12T11:17:50.370351Z"},"papermill":{"duration":0.032371,"end_time":"2026-09-12T11:17:50.372561+00:00","exception":false,"start_time":"2026-09-12T11:17:50.34019+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"HDR_TAGS = ['SeriesDescription', 'SequenceName', 'ScanOptions', 'ScanningSequence', 'RepetitionTime', 'EchoTime', 'Laterality', 'PixelSpacing', 'Rows', 'Columns', 'RescaleSlope', 'RescaleIntercept', 'ImagePositionPatient', 'ImageOrientationPatient']\n\ndef _hdr_vec(s, n):\n    if not isinstance(s, str):\n        return None\n    try:\n        v = [float(x) for x in s.split('|')]\n    except ValueError:\n        return None\n    return np.array(v) if len(v) >= n else None\n\ndef side_from_geometry(h):\n    cx = {}\n    for r in h.itertuples(index=False):\n        ipp = _hdr_vec(getattr(r, 'ImagePositionPatient', None), 3)\n        iop = _hdr_vec(getattr(r, 'ImageOrientationPatient', None), 6)\n        ps = _hdr_vec(getattr(r, 'PixelSpacing', None), 2)\n        rows, cols = (getattr(r, 'Rows', None), getattr(r, 'Columns', None))\n        if ipp is None or iop is None or ps is None or (not rows) or (not cols):\n            continue\n        try:\n            c = ipp[:3] + iop[:3] * ps[1] * float(cols) / 2 + iop[3:6] * ps[0] * float(rows) / 2\n        except (TypeError, ValueError):\n            continue\n        cx.setdefault(r.StudyInstanceUID, []).append(float(c[0]))\n    out = {}\n    for st, xs in cx.items():\n        m = float(np.median(xs))\n        out[st] = None if abs(m) < LAT_MIN_OFFSET_MM else 'R' if m < 0 else 'L'\n    return out\n\ndef side_from_corner_x(h):\n    out = {}\n    for st, g in h.groupby('StudyInstanceUID'):\n        xs = []\n        for r in g.itertuples(index=False):\n            ipp = _hdr_vec(getattr(r, 'ImagePositionPatient', None), 3)\n            if ipp is not None and np.isfinite(ipp).all():\n                xs.append(float(ipp[0]))\n        if not xs:\n            out[st] = None\n            continue\n        x = float(np.median(xs))\n        out[st] = None if abs(x) < LEGACY_LAT_OFFSET_MM else 'R' if x < 0 else 'L'\n    return out\n\ndef lat_of(h, tag=''):\n    geo = side_from_corner_x(h) if RULES['lat'] == 'corner_x' else side_from_geometry(h)\n    d, n_tag, n_geo, n_none, n_disagree = ({}, 0, 0, 0, 0)\n    for st, g in h.groupby('StudyInstanceUID'):\n        v = [str(x).strip().upper() for x in g['Laterality'].dropna()]\n        if RULES['lat'] == 'corner_x' and 'ImageLaterality' in g.columns:\n            v += [str(x).strip().upper() for x in g['ImageLaterality'].dropna()]\n        v = [x[0] for x in v if x and x[0] in ('L', 'R')]\n        side = v[0] if v else None\n        if side is not None:\n            n_tag += 1\n            if geo.get(st) is not None and geo[st] != side:\n                n_disagree += 1\n        else:\n            side = geo.get(st)\n            n_geo += side is not None\n            n_none += side is None\n        d[st] = side\n    log(f'{tag}laterality: {n_tag} from the tag, {n_geo} from geometry, {n_none} unresolved; tag and geometry disagree on {n_disagree} ({n_disagree / max(n_tag, 1):.1%} of the tagged)')\n    return d\n\ndef probe(item):\n    split, study, series, path = item\n    row = {'split': split, 'StudyInstanceUID': study, 'SeriesInstanceUID': series, 'dir': path}\n    try:\n        files = sorted((e.name for e in os.scandir(path) if e.name.endswith('.dcm')))\n        row['files'] = files\n        row['n_slices'] = len(files)\n        if not files:\n            return row\n        ds = pydicom.dcmread(os.path.join(path, files[len(files) // 2]), stop_before_pixels=True, force=True)\n        for t in HDR_TAGS:\n            v = getattr(ds, t, None)\n            if v is None:\n                row[t] = None\n            elif isinstance(v, (list, tuple)) or type(v).__name__ == 'MultiValue':\n                row[t] = '|'.join((str(x) for x in v))\n            else:\n                row[t] = str(v)\n    except Exception as exc:\n        row['err'] = str(exc)[:120]\n    return row\n\ndef walk(split):\n    base = ROOT / split\n    items = []\n    if not base.is_dir():\n        return pd.DataFrame(columns=['split', 'StudyInstanceUID', 'SeriesInstanceUID', 'dir', 'files', 'n_slices'] + HDR_TAGS)\n    for study in os.scandir(base):\n        if study.is_dir():\n            for series in os.scandir(study.path):\n                if series.is_dir():\n                    items.append((split, study.name, series.name, series.path))\n    with ThreadPoolExecutor(max_workers=HDR_THREADS) as pool:\n        rows = list(pool.map(probe, items))\n    return pd.DataFrame(rows)\n\ndef annotate(df):\n    desc = df['SeriesDescription'].fillna('') + ' ' + df['SequenceName'].fillna('')\n    desc = desc.str.lower().str.replace(_SEP, ' ', regex=True)\n    opts = df['ScanOptions'].fillna('').str.upper().str.split('|')\n    opts_fs = opts.apply(lambda ts: any((t.strip() in FATSAT_OPTS for t in ts)))\n    df['fatsat'] = desc.str.contains(_FATSAT_RX) | opts_fs\n    tr = pd.to_numeric(df['RepetitionTime'], errors='coerce')\n    te = pd.to_numeric(df['EchoTime'], errors='coerce')\n    gre = df['ScanningSequence'].fillna('').str.upper().str.contains('GR')\n    t1, t2, pdw = (desc.str.contains(_T1_RX), desc.str.contains(_T2_RX), desc.str.contains(_PD_RX))\n    df['weight'] = np.where(t1 & ~t2 & ~pdw, 'T1', np.where(t2 & ~pdw, 'T2', np.where(pdw, 'PD', np.where(gre, 'GRE', np.where(tr < 800, 'T1', np.where(te > 60, 'T2', np.where(tr >= 800, 'PD', 'UNK')))))))\n    df['fluid'] = np.isin(df['weight'], ['PD', 'T2'])\n    df['px'] = pd.to_numeric(df['PixelSpacing'].fillna('').str.split('|').str[0].replace('', np.nan), errors='coerce')\n    return df"},{"cell_type":"markdown","id":"7df9e241","metadata":{"papermill":{"duration":0.008028,"end_time":"2026-09-12T11:17:50.389462+00:00","exception":false,"start_time":"2026-09-12T11:17:50.381434+00:00","status":"completed"},"tags":[]},"source":"# 4. 🗂️ Выбор серий для слотов\n\n## Что делает эта ячейка\n\nДля каждого исследования выбирает наиболее подходящую серию под каждый из шести слотов:\n\n- Sagittal / Coronal / Axial × Fluid-Sensitive / Non-Fluid-Sensitive / T1.\n- Если точного совпадения нет, используется fallback.\n\nРезультат — словарь `{study: {slot_name: series_row}}`, который потом используется для построения кэша изображений."},{"cell_type":"code","execution_count":null,"id":"9681b454","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.406862Z","iopub.status.busy":"2026-09-12T11:17:50.40653Z","iopub.status.idle":"2026-09-12T11:17:50.411647Z","shell.execute_reply":"2026-09-12T11:17:50.411073Z"},"papermill":{"duration":0.015399,"end_time":"2026-09-12T11:17:50.412983+00:00","exception":false,"start_time":"2026-09-12T11:17:50.397584+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"def pick_slots(series_df, plane_map):\n    series_df = series_df.copy()\n    series_df['plane'] = series_df['SeriesInstanceUID'].map(plane_map)\n    out = {}\n    for study, g in series_df.groupby('StudyInstanceUID'):\n        chosen = {}\n        for name, plane, fluid, fs in SLOTS:\n            sel = (g['plane'] == plane) & (g['fatsat'] == fs)\n            if fluid is not None:\n                sel &= g['fluid'] == fluid\n            cand = g[sel]\n            if len(cand) == 0 and RULES['slot_fallback'] and (fluid is False):\n                cand = g[(g['plane'] == plane) & ~g['fatsat']]\n            if len(cand):\n                chosen[name] = cand.sort_values('n_slices', ascending=False).iloc[0]\n        out[study] = chosen\n    return out"},{"cell_type":"markdown","id":"dd0339cd","metadata":{"papermill":{"duration":0.008358,"end_time":"2026-09-12T11:17:50.429798+00:00","exception":false,"start_time":"2026-09-12T11:17:50.42144+00:00","status":"completed"},"tags":[]},"source":"# 5. 📐 Геометрическая сортировка и чтение срезов\n\n## Что делает эта ячейка\n\n- `order_slices()` — сортирует DICOM-файлы по физическому положению среза в пространстве, а не по имени файла.\n- `read_slot()` — загружает центральные срезы, кадрирует по 130 мм, нормализует интенсивность по 1–99 перцентилям и приводит к `IMG×IMG`.\n\nРезультат — тензор uint8 размером `(n_slices, IMG, IMG)`, готовый для модели."},{"cell_type":"code","execution_count":null,"id":"dd2a5282","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.447777Z","iopub.status.busy":"2026-09-12T11:17:50.447502Z","iopub.status.idle":"2026-09-12T11:17:50.468044Z","shell.execute_reply":"2026-09-12T11:17:50.467259Z"},"papermill":{"duration":0.031762,"end_time":"2026-09-12T11:17:50.469698+00:00","exception":false,"start_time":"2026-09-12T11:17:50.437936+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"ORDER_TAGS = [(32, 50), (32, 55), (32, 19)]\nDECODE_FAILED = []\n\n\ndef _natural_key(name):\n    return tuple((int(x) if x.isdigit() else x.lower() for x in re.split('(\\\\d+)', str(name))))\n\ndef _order_dominant_axis(rec):\n    files, d = (rec['files'], rec['dir'])\n    rows = []\n    for pos, f in enumerate(files):\n        ipp = inst = None\n        try:\n            ds = pydicom.dcmread(os.path.join(d, f), force=True, stop_before_pixels=True, specific_tags=['ImagePositionPatient', 'InstanceNumber'])\n            raw = getattr(ds, 'ImagePositionPatient', None)\n            if raw is not None and len(raw) >= 3:\n                c = np.asarray(raw[:3], dtype=np.float64)\n                if np.isfinite(c).all():\n                    ipp = c\n            n = getattr(ds, 'InstanceNumber', None)\n            if n is not None:\n                inst = float(n)\n        except Exception:\n            pass\n        rows.append((f, ipp, inst, pos))\n    placed = [r for r in rows if r[1] is not None]\n    need = max(2, int(0.8 * len(rows)))\n    if len(placed) >= need:\n        xyz = np.stack([r[1] for r in placed])\n        axis = int(np.argmax(np.ptp(xyz, axis=0)))\n        spare = float(np.nanmedian(xyz[:, axis]))\n        rows.sort(key=lambda r: (float(r[1][axis]) if r[1] is not None else spare, r[2] if r[2] is not None else float('inf'), r[3]))\n    elif sum((r[2] is not None for r in rows)) >= need:\n        rows.sort(key=lambda r: (r[2] if r[2] is not None else float('inf'), r[3]))\n    else:\n        rows.sort(key=lambda r: _natural_key(r[0]))\n    return ([r[0] for r in rows], True)\n\ndef order_slices(rec):\n    if RULES['order'] == 'dominant_axis':\n        return _order_dominant_axis(rec)\n    files, d = (rec['files'], rec['dir'])\n    keyed = []\n    for f in files:\n        k = None\n        try:\n            ds = pydicom.dcmread(os.path.join(d, f), force=True, stop_before_pixels=True, specific_tags=ORDER_TAGS)\n            iop = np.asarray(ds.ImageOrientationPatient, dtype=float)\n            ipp = np.asarray(ds.ImagePositionPatient, dtype=float)\n            k = float(np.dot(ipp, np.cross(iop[:3], iop[3:])))\n        except Exception:\n            try:\n                k = float(ds.InstanceNumber)\n            except Exception:\n                k = None\n        keyed.append((k, f))\n    if any((k is None for k, _ in keyed)):\n        return (files, False)\n    return ([f for _, f in sorted(keyed, key=lambda t: t[0])], True)\n\ndef read_slot(rec, n_slice=None, out_size=None):\n    n_slice = GROUP if n_slice is None else n_slice\n    out_size = IMG if out_size is None else out_size\n    files, d, px = (rec.get('ordered') or rec['files'], rec['dir'], rec['px'])\n    n = len(files)\n    if n == 0:\n        return None\n    lo, hi = (int(SLICE_BAND[0] * (n - 1)), int(SLICE_BAND[1] * (n - 1)))\n    idx = np.unique(np.linspace(lo, hi, n_slice).astype(int)) if hi > lo else np.array([n // 2])\n    while len(idx) < n_slice:\n        idx = np.append(idx, idx[-1])\n    planes = []\n    for i in idx[:n_slice]:\n        try:\n            ds = pydicom.dcmread(os.path.join(d, files[int(i)]), force=True)\n            a = ds.pixel_array.astype(np.float32)\n            sl = float(getattr(ds, 'RescaleSlope', 1) or 1)\n            ic = float(getattr(ds, 'RescaleIntercept', 0) or 0)\n            a = a * sl + ic\n        except Exception:\n            a = None\n        planes.append(a)\n    got = [k for k, p in enumerate(planes) if p is not None]\n    if RULES['decode_fill'] == 'zero':\n        if not got:\n            DECODE_FAILED.append(rec.get('SeriesInstanceUID', d))\n        planes = [np.zeros((out_size, out_size), np.float32) if p is None else p for p in planes]\n        got = list(range(len(planes)))\n    if not got:\n        DECODE_FAILED.append(rec.get('SeriesInstanceUID', d))\n        return None\n    if len(got) < len(planes):\n        DECODE_FAILED.append(rec.get('SeriesInstanceUID', d))\n        for k, p in enumerate(planes):\n            if p is None:\n                planes[k] = planes[min(got, key=lambda j: abs(j - k))]\n    shp = planes[0].shape\n    planes = [p if p.shape == shp else np.zeros(shp, np.float32) for p in planes]\n    vol = np.stack(planes)\n    if px and np.isfinite(px) and (px > 0):\n        want = int(round(CROP_MM / px))\n        h, w = shp\n        if 16 < want < min(h, w):\n            cy, cx = (h // 2, w // 2)\n            half = want // 2\n            vol = vol[:, max(0, cy - half):cy + half, max(0, cx - half):cx + half]\n    lo_v, hi_v = np.percentile(vol, [1, 99])\n    vol = np.clip((vol - lo_v) / max(hi_v - lo_v, 1e-06), 0, 1)\n    t = torch.from_numpy(np.ascontiguousarray(vol)).unsqueeze(0)\n    t = F.interpolate(t, size=(out_size, out_size), mode='bilinear', align_corners=False)\n    return (t.squeeze(0) * 255).round().clamp(0, 255).to(torch.uint8)"},{"cell_type":"markdown","id":"b89b6601","metadata":{"papermill":{"duration":0.008414,"end_time":"2026-09-12T11:17:50.487072+00:00","exception":false,"start_time":"2026-09-12T11:17:50.478658+00:00","status":"completed"},"tags":[]},"source":"# 6. 🔄 Нормализация латеральности\n\n## Что делает эта ячейка\n\nПриводит все колени к единой (левой) ориентации:\n\n- Для корональных и аксиальных срезов правого колена — горизонтальный flip.\n- Для сагиттальных — реверс порядка срезов.\n\nЭто важно для медиальных/латеральных структур, которые зависят от стороны."},{"cell_type":"code","execution_count":null,"id":"4a645bc4","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.504556Z","iopub.status.busy":"2026-09-12T11:17:50.504344Z","iopub.status.idle":"2026-09-12T11:17:50.508586Z","shell.execute_reply":"2026-09-12T11:17:50.507727Z"},"papermill":{"duration":0.014827,"end_time":"2026-09-12T11:17:50.510164+00:00","exception":false,"start_time":"2026-09-12T11:17:50.495337+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"def normalise_laterality(img, plane, lat):\n    if lat != 'R':\n        return img\n    if plane in ('Coronal', 'Axial'):\n        return torch.flip(img, dims=[-1])\n    return torch.flip(img, dims=[0])"},{"cell_type":"markdown","id":"1e7d075b","metadata":{"papermill":{"duration":0.008538,"end_time":"2026-09-12T11:17:50.527135+00:00","exception":false,"start_time":"2026-09-12T11:17:50.518597+00:00","status":"completed"},"tags":[]},"source":"# 7. 💾 Построение кэша изображений\n\n## Что делает эта ячейка\n\nЧитает все выбранные серии и сохраняет их в памяти как uint8-массив.\n\n- Сначала сортирует срезы геометрически.\n- Затем декодирует и нормализует каждый слот.\n- Возвращает `(studies, cache, mask)`.\n\nЭто самая тяжёлая по I/O часть инференса. После неё модели работают быстро, без повторного чтения DICOM."},{"cell_type":"code","execution_count":null,"id":"c3cc9422","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.546911Z","iopub.status.busy":"2026-09-12T11:17:50.546218Z","iopub.status.idle":"2026-09-12T11:17:50.560413Z","shell.execute_reply":"2026-09-12T11:17:50.559816Z"},"papermill":{"duration":0.024924,"end_time":"2026-09-12T11:17:50.56173+00:00","exception":false,"start_time":"2026-09-12T11:17:50.536806+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"ORDER_CACHE = os.environ.get('RSNA_ORDER_CACHE') or None\n\ndef build_cache(slot_map, plane_map, lat_map, tag):\n    studies = sorted(slot_map)\n    sidx = {s: i for i, s in enumerate(studies)}\n    cache = np.zeros((len(studies), N_SLOT, CACHE_SLICES, IMG, IMG), np.uint8)\n    mask = np.zeros((len(studies), N_SLOT), np.float32)\n    log(f'{tag}: cache {cache.shape} = {cache.nbytes / 1024 ** 3:.1f} GB')\n    jobs = [(st, k, plane, slot_map[st][name]) for st in studies for k, (name, plane, _, _) in enumerate(SLOTS) if name in slot_map[st]]\n    n_job = len(jobs)\n    t_ord = time.time()\n    n_slice_total = sum((len(j[3]['files']) for j in jobs))\n    log(f'{tag}: ordering {len(jobs)} slot-series ({n_slice_total} slice headers)')\n    ok = done = 0\n    CHUNK_O = 1024\n    seen = {}\n    if ORDER_CACHE and Path(ORDER_CACHE).is_file():\n        try:\n            import json as _json\n            seen = _json.loads(Path(ORDER_CACHE).read_text())\n        except (OSError, ValueError):\n            seen = {}\n        hit = 0\n        for _, _, _, rec in jobs:\n            e = seen.get(rec['SeriesInstanceUID'])\n            if e and len(e['files']) == len(rec['files']):\n                rec['ordered'] = e['files']\n                ok += int(e['good'])\n                hit += 1\n        jobs = [j for j in jobs if 'ordered' not in j[3]]\n        log(f'{tag}: {hit} slot-series ordered from {ORDER_CACHE}, {len(jobs)} to read')\n    with ThreadPoolExecutor(max_workers=ORDER_THREADS) as pool:\n        for c0 in range(0, len(jobs), CHUNK_O):\n            block = jobs[c0:c0 + CHUNK_O]\n            for (_, _, _, rec), (files, good) in zip(block, pool.map(lambda j: order_slices(j[3]), block)):\n                rec['ordered'] = files\n                ok += int(good)\n                done += 1\n                if ORDER_CACHE:\n                    seen[rec['SeriesInstanceUID']] = {'files': files, 'good': bool(good)}\n            budget = min(ORDER_BUDGET_S, max(60.0, (TIME_BUDGET - (time.time() - T0)) * 0.35))\n            if time.time() - t_ord > budget:\n                log(f'{tag}: ordering budget spent at {done}/{len(jobs)}; the rest keep file order')\n                break\n    if ORDER_CACHE and done:\n        import json as _json\n        _t = Path(ORDER_CACHE).with_suffix('.tmp')\n        _t.write_text(_json.dumps(seen))\n        _t.replace(Path(ORDER_CACHE))\n    log(f'{tag}: ordered {ok}/{n_job} by geometry ({n_job - ok} kept arbitrary) in {time.time() - t_ord:.0f}s')\n    jobs = [(st, k, plane, slot_map[st][name]) for st in studies for k, (name, plane, _, _) in enumerate(SLOTS) if name in slot_map[st]]\n    log(f'{tag}: decoding {len(jobs)} slot-series')\n    n_failed_before = len(DECODE_FAILED)\n    CHUNK = 512\n    done = 0\n    with ThreadPoolExecutor(max_workers=PIX_THREADS) as pool:\n        for c0 in range(0, len(jobs), CHUNK):\n            block = jobs[c0:c0 + CHUNK]\n            for (st, k, plane, _), img in zip(block, pool.map(lambda j: read_slot(j[3], CACHE_SLICES, IMG), block)):\n                done += 1\n                if img is None:\n                    continue\n                cache[sidx[st], k] = normalise_laterality(img, plane, lat_map.get(st)).numpy()\n                mask[sidx[st], k] = 1.0\n            if done % 4096 < CHUNK:\n                log(f'  {tag} {done}/{len(jobs)}')\n            if time.time() - T0 > TIME_BUDGET:\n                log(f'  {tag}: time budget reached during decode')\n                break\n    n_failed = len(DECODE_FAILED) - n_failed_before\n    log(f'{tag}: {int(mask.sum())}/{len(jobs)} slots filled' + (f'; {n_failed} series had a slice that would not decode' if n_failed else ''))\n    gc.collect()\n    return (studies, cache, mask)"},{"cell_type":"markdown","id":"2c330baa","metadata":{"papermill":{"duration":0.008279,"end_time":"2026-09-12T11:17:50.578432+00:00","exception":false,"start_time":"2026-09-12T11:17:50.570153+00:00","status":"completed"},"tags":[]},"source":"# 8. 🧠 SlotHead: внимание по слотам\n\n## Что делает эта ячейка\n\nОпределяет голову, которая агрегирует слоты с помощью внимания.\n\n- Каждый слот проецируется в скрытое пространство.\n- Для каждого диагноза обучается свой query.\n- Внимание маскирует отсутствующие слоты.\n- Итоговый вектор — взвешенная сумма слотов.\n\nЭто ключевая часть DINOv2 трансформера."},{"cell_type":"code","execution_count":null,"id":"ab3ae8ac","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.596956Z","iopub.status.busy":"2026-09-12T11:17:50.596188Z","iopub.status.idle":"2026-09-12T11:17:50.603599Z","shell.execute_reply":"2026-09-12T11:17:50.603069Z"},"papermill":{"duration":0.017841,"end_time":"2026-09-12T11:17:50.604896+00:00","exception":false,"start_time":"2026-09-12T11:17:50.587055+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"class SlotHead(nn.Module):\n\n    def __init__(self, dim, n_slot, n_out, hidden=256, p=0.2, prior=False):\n        super().__init__()\n        self.proj = nn.Sequential(nn.LayerNorm(dim), nn.Linear(dim, hidden), nn.GELU())\n        self.slot_emb = nn.Parameter(torch.randn(n_slot, hidden) * 0.02)\n        self.query = nn.Parameter(torch.randn(n_out, hidden) * 0.02)\n        self.drop = nn.Dropout(p)\n        self.out = nn.Linear(hidden, n_out)\n        self.hidden = hidden\n        p_ = torch.zeros(n_out, n_slot)\n        if prior and n_slot == len(SLOTS) and (n_out == len(TARGETS)):\n            for t, slots in SLOT_PRIOR_TABLE.items():\n                if t in TARGETS:\n                    p_[TARGETS.index(t), list(slots)] = SLOT_PRIOR_STRENGTH\n        self.prior = prior\n        if prior:\n            self.register_buffer('slot_prior', p_)\n\n    def forward(self, x, mask):\n        h = self.proj(x) + self.slot_emb\n        att = torch.einsum('bsh,oh->bos', h, self.query) / self.hidden ** 0.5\n        if self.prior:\n            att = att + self.slot_prior.unsqueeze(0)\n        att = att.masked_fill(mask.unsqueeze(1) < 0.5, -10000.0).softmax(-1)\n        ctx = self.drop(torch.einsum('bos,bsh->boh', att, h))\n        return (ctx * self.out.weight.unsqueeze(0)).sum(-1) + self.out.bias"},{"cell_type":"markdown","id":"be8d0eef","metadata":{"papermill":{"duration":0.008247,"end_time":"2026-09-12T11:17:50.621504+00:00","exception":false,"start_time":"2026-09-12T11:17:50.613257+00:00","status":"completed"},"tags":[]},"source":"# 9. 🏗️ Основная модель DINOv2\n\n## Что делает эта ячейка\n\nОборачивает DINOv2 backbone и SlotHead в единую модель.\n\n- Принимает батч изображений из кэша.\n- Прогоняет через DINOv2.\n- Извлекает CLS-токен и средний патч.\n- Передаёт в SlotHead для получения 12 логитов."},{"cell_type":"code","execution_count":null,"id":"35c32aa2","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.640852Z","iopub.status.busy":"2026-09-12T11:17:50.640588Z","iopub.status.idle":"2026-09-12T11:17:50.648833Z","shell.execute_reply":"2026-09-12T11:17:50.647863Z"},"papermill":{"duration":0.019974,"end_time":"2026-09-12T11:17:50.650503+00:00","exception":false,"start_time":"2026-09-12T11:17:50.630529+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"class Model(nn.Module):\n\n    def __init__(self, backbone, dim, pool='cls_mean', prior=False):\n        super().__init__()\n        self.backbone = backbone\n        self.pool = pool\n        self.head = SlotHead(dim * POOL_PARTS[pool], N_SLOT, len(TARGETS), prior=prior)\n        self.register_buffer('mean', torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))\n        self.register_buffer('std', torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))\n\n    def forward(self, imgs, mask, img_size=None):\n        B, S = imgs.shape[:2]\n        x = imgs.reshape(B * S, *imgs.shape[2:]).float().div_(255.0)\n        if img_size is not None and img_size != x.shape[-1]:\n            x = F.interpolate(x, size=(img_size, img_size), mode='bilinear', align_corners=False)\n        x = (x - self.mean) / self.std\n        out = self.backbone(pixel_values=x).last_hidden_state\n        patch = out[:, 1:]\n        parts = [out[:, 0], patch.mean(1)]\n        if self.pool == 'cls_mean_focal':\n            k = max(1, patch.shape[1] // 8)\n            parts.append(patch.topk(k, dim=1).values.mean(1))\n        feat = torch.cat(parts, dim=1).reshape(B, S, -1)\n        return self.head(feat, mask)"},{"cell_type":"markdown","id":"7406860b","metadata":{"papermill":{"duration":0.008444,"end_time":"2026-09-12T11:17:50.668046+00:00","exception":false,"start_time":"2026-09-12T11:17:50.659602+00:00","status":"completed"},"tags":[]},"source":"# 10. 📦 Загрузка DINOv2\n\n## Что делает эта ячейка\n\nЗагружает предобученный DINOv2 из подключённой модели.\n\n- Замораживает все параметры.\n- Размораживает последние `unfreeze_last` блоков.\n- Возвращает модель с головой."},{"cell_type":"code","execution_count":null,"id":"7b279ae6","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.686125Z","iopub.status.busy":"2026-09-12T11:17:50.685834Z","iopub.status.idle":"2026-09-12T11:17:50.691783Z","shell.execute_reply":"2026-09-12T11:17:50.690937Z"},"papermill":{"duration":0.01659,"end_time":"2026-09-12T11:17:50.693157+00:00","exception":false,"start_time":"2026-09-12T11:17:50.676567+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"def build_model(unfreeze_last, source=None, variant='small', pool='cls_mean', prior=False):\n    from transformers import AutoModel\n    p = source if source is not None else find_dinov2(variant)\n    if p is None:\n        raise FileNotFoundError('DINOv2 weights not attached')\n    bb = AutoModel.from_pretrained(str(p))\n    n_layer = len(bb.encoder.layer)\n    for prm in bb.parameters():\n        prm.requires_grad = False\n    for blk in bb.encoder.layer[max(0, n_layer - unfreeze_last):]:\n        for prm in blk.parameters():\n            prm.requires_grad = True\n    for prm in bb.layernorm.parameters():\n        prm.requires_grad = True\n    dim = bb.config.hidden_size\n    trainable = sum((p.numel() for p in bb.parameters() if p.requires_grad))\n    log(f'backbone: {n_layer} blocks, last {unfreeze_last} trainable ({trainable / 1000000.0:.1f}M params), feature dim {dim * POOL_PARTS[pool]}')\n    return Model(bb, dim, pool=pool, prior=prior)"},{"cell_type":"markdown","id":"adc683fa","metadata":{"papermill":{"duration":0.008025,"end_time":"2026-09-12T11:17:50.70953+00:00","exception":false,"start_time":"2026-09-12T11:17:50.701505+00:00","status":"completed"},"tags":[]},"source":"# 11. 🔍 Проверка целостности, ансамблирование и инференс пакета\n\n## Что делает эта ячейка\n\nСодержит функции для:\n\n- `fingerprint` / `check_fingerprint` — проверка, что веса модели соответствуют ожидаемым.\n- `predict_member` — инференс одной модели с TTA-пулингом.\n- `infer_from_package` — загрузка манифеста из пакета весов, построение кэша и прогон всех моделей.\n- `_combine` — взвешенное ранговое усреднение предсказаний.\n\nЭто ядро DINOv2-ветки. Создаёт `submission.csv` после обработки всех 20 моделей."},{"cell_type":"code","execution_count":null,"id":"2d27d4c3","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.727381Z","iopub.status.busy":"2026-09-12T11:17:50.72716Z","iopub.status.idle":"2026-09-12T11:17:50.782658Z","shell.execute_reply":"2026-09-12T11:17:50.781781Z"},"papermill":{"duration":0.066485,"end_time":"2026-09-12T11:17:50.784243+00:00","exception":false,"start_time":"2026-09-12T11:17:50.717758+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"FINGERPRINT_TOL = 0.002\n\ndef fingerprint(model, dev, img_size, n_slot=None, group=None, seed=None):\n    n_slot = N_SLOT if n_slot is None else n_slot\n    group = GROUP if group is None else group\n    seed = SEED if seed is None else seed\n    g = torch.Generator().manual_seed(seed)\n    imgs = torch.randint(0, 256, (2, n_slot, group, img_size, img_size), generator=g, dtype=torch.uint8).to(dev)\n    mask = torch.ones(2, n_slot, device=dev)\n    mask[1, -1] = 0.0\n    was_training = model.training\n    model.eval()\n    with torch.no_grad():\n        out = model(imgs, mask, img_size).float().cpu().numpy()\n    if was_training:\n        model.train()\n    return out\n\ndef check_fingerprint(model, dev, img_size, expected, tol=FINGERPRINT_TOL, tag=''):\n    got = fingerprint(model, dev, img_size)\n    exp = np.asarray(expected, np.float32)\n    if got.shape != exp.shape:\n        raise WeightsError(f'{tag}fingerprint shape {got.shape} != stored {exp.shape}: the architecture is not the one these weights were fitted to')\n    d = float(np.abs(got - exp).max())\n    if d > tol:\n        raise WeightsError(f'{tag}fingerprint differs by {d:.4g} (tolerance {tol:g}). The weights load but do not compute what they computed when fitted - preprocessing, resolution or architecture has moved between the two runs.')\n    log(f'{tag}fingerprint matches within {d:.2g}')\n    return d\n\nclass WeightsError(RuntimeError):\n    pass\n\ndef find_weights(name='manifest.json'):\n    import json\n    base = Path('/kaggle/input')\n    if not base.is_dir():\n        return None\n    incomplete = []\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in ('train_series', 'test_series')]\n        if name not in files:\n            continue\n        try:\n            man = json.loads((Path(root) / name).read_text())\n        except (OSError, ValueError):\n            continue\n        if isinstance(man.get('members'), list) and man['members']:\n            missing = [m['file'] for m in man['members'] if not (Path(root) / m['file']).is_file()]\n            if missing:\n                incomplete.append(f\"{root} lists {len(man['members'])} member(s), {len(missing)} absent (first {missing[0]!r})\")\n                continue\n            return Path(root)\n    if incomplete:\n        raise WeightsError('no complete weights package found under /kaggle/input; ' + '; '.join(incomplete))\n    return None\nTTA_OVERLAP = True\n# The jittered view feeds win_probs only.  submission.csv is promoted from the\n# public frontier, which is pooled from original_probs -- the un-jittered view\n# -- so the second forward pass per window costs up to 2x this branch and\n# reaches nothing but submission_native_v38.csv.\nTTA_JITTER = False\nTTA_POOL = 'prob'\nPUBLIC_FRONTIER_TARGET_POOL = {'Fracture': 'max', 'Contusion': 'max', 'Medial Meniscus': 'max', 'Lateral Meniscus': 'max', 'ACL': 'top2', 'MCL': 'top2', \"Baker's\": 'max'}\nTTA_TARGET_POOL = {**PUBLIC_FRONTIER_TARGET_POOL, 'Synovitis': 'original_mean'}\n# No-extra-pass diversity branch: smooth focal pooling is evaluated from\n# the same no-jitter public-member windows already used by the parent.\nLEGACY_FOLD_SOFTPOOL_BETA = {\n    'ACL': 6.0, 'MCL': 6.0,\n    'Medial Meniscus': 8.0, 'Lateral Meniscus': 8.0,\n    \"Baker's\": 8.0, 'Contusion': 8.0, 'Fracture': 10.0,\n}\nLEGACY_FOLD_SOFTPOOL_ALPHA = {\n    'ACL': 0.20, 'MCL': 0.20,\n    'Medial Meniscus': 0.25, 'Lateral Meniscus': 0.25,\n    \"Baker's\": 0.20, 'Contusion': 0.20, 'Fracture': 0.15,\n}\nLEGACY_MEMBER_WEIGHT_BY_TARGET = {'Lateral Meniscus': 15.0, 'Medial OA': 2.5, 'Lateral OA': 15.0, 'Contusion': 5.0}\n\ndef window_starts(n_slice, group, overlap=None):\n    overlap = TTA_OVERLAP if overlap is None else overlap\n    if overlap and n_slice >= group:\n        return list(range(n_slice - group + 1))\n    return [g * group for g in range(max(n_slice // group, 1))]\n\ndef apply_target_window_pool(values, probs, logits, original_probs, mapping, target_idx):\n    for target, mode in mapping.items():\n        j = target_idx[target]\n        if mode == 'max':\n            values[:, j] = probs[:, :, j].max(0).values\n        elif mode == 'mean':\n            values[:, j] = probs[:, :, j].mean(0)\n        elif mode == 'logit_mean':\n            values[:, j] = torch.sigmoid(logits[:, :, j].mean(0))\n        elif mode == 'original_mean':\n            values[:, j] = original_probs[:, :, j].mean(0)\n        elif mode in ('top2', 'top3'):\n            k = min(int(mode[3:]), probs.shape[0])\n            values[:, j] = probs[:, :, j].topk(k, dim=0).values.mean(0)\n        else:\n            raise ValueError(f'unknown TTA pooling mode for {target}: {mode}')\n    return values\n\ndef legacy_fold_soft_window_pool(original_probs, target_idx):\n    values = original_probs.mean(0).clone()\n    for target, beta in LEGACY_FOLD_SOFTPOOL_BETA.items():\n        j = target_idx[target]\n        x = original_probs[:, :, j]\n        weight = torch.softmax(float(beta) * x, dim=0)\n        values[:, j] = (weight * x).sum(0)\n    return values\n\n@torch.no_grad()\ndef predict_member(model, cache, mask, idx, dev, img_size, group=None, pool=None, starts=None, jitter=False, jitter_seed=SEED, return_public_frontier=False):\n    group = GROUP if group is None else group\n    pool = TTA_POOL if pool is None else pool\n    starts = window_starts(cache.shape[2], group) if starts is None else list(starts)\n    if not starts:\n        raise ValueError('predict_member was given no windows to average over')\n    target_idx = {t: j for j, t in enumerate(TARGETS)}\n    unknown = (set(TTA_TARGET_POOL) | set(PUBLIC_FRONTIER_TARGET_POOL)) - set(target_idx)\n    if unknown:\n        raise ValueError(f'unknown target(s) in TTA_TARGET_POOL: {unknown}')\n    jitter_gen = torch.Generator(device=dev)\n    jitter_gen.manual_seed(int(jitter_seed) % (2 ** 63 - 1))\n    model.eval()\n    out, public_frontier_out, public_soft_out = ([], [], [])\n    for b in range(0, len(idx), EVAL_BATCH):\n        sel = idx[b:b + EVAL_BATCH]\n        m = torch.from_numpy(mask[sel]).to(dev)\n        win_probs, win_logits, win_original_probs = ([], [], [])\n        for st in starts:\n            rows = torch.from_numpy(np.ascontiguousarray(cache[sel, :, st:st + group])).to(dev)\n            views = [rows] + ([augment(rows, generator=jitter_gen)] if jitter else [])\n            view_probs, view_logits = ([], [])\n            for view in views:\n                with torch.autocast('cuda', enabled=dev.type == 'cuda'):\n                    z = model(view, m, img_size).float()\n                view_logits.append(z)\n                view_probs.append(torch.sigmoid(z))\n            win_logits.append(torch.stack(view_logits).mean(0))\n            win_probs.append(torch.stack(view_probs).mean(0))\n            win_original_probs.append(view_probs[0])\n        probs = torch.stack(win_probs)\n        logits = torch.stack(win_logits)\n        original_probs = torch.stack(win_original_probs)\n        v = torch.sigmoid(logits.mean(0)) if pool == 'logit' else probs.mean(0)\n        v = apply_target_window_pool(v, probs, logits, original_probs, TTA_TARGET_POOL, target_idx)\n        out.append(v.cpu().numpy())\n        if return_public_frontier:\n            public_v = apply_target_window_pool(original_probs.mean(0), original_probs, logits, original_probs, PUBLIC_FRONTIER_TARGET_POOL, target_idx)\n            public_frontier_out.append(public_v.cpu().numpy())\n            public_soft = legacy_fold_soft_window_pool(original_probs, target_idx)\n            public_soft_out.append(public_soft.cpu().numpy())\n    primary = np.concatenate(out) if out else np.zeros((0, len(TARGETS)), np.float32)\n    if not return_public_frontier:\n        return primary\n    public_frontier = np.concatenate(public_frontier_out) if public_frontier_out else np.zeros((0, len(TARGETS)), np.float32)\n    public_soft = np.concatenate(public_soft_out) if public_soft_out else np.zeros((0, len(TARGETS)), np.float32)\n    return (primary, public_frontier, public_soft)\nBUILD_LOCK = threading.Lock()\nSTATE_LOCK = threading.Lock()\nLEGACY_BUNDLE_FILE = 'rsna_20260807_v1.pt'\nLEGACY_WEIGHT = 0.5\n\ndef find_legacy_bundle():\n    base = Path('/kaggle/input')\n    if not base.is_dir():\n        return None\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in ('train_series', 'test_series')]\n        if LEGACY_BUNDLE_FILE in files:\n            return Path(root) / LEGACY_BUNDLE_FILE\n    return None\n\ndef legacy_group_members():\n    return {}\n\ndef _run_member(path, m, dev, Cte, Mte, idx, starts, jitter):\n    t0 = time.time()\n    with BUILD_LOCK:\n        if 'state' in m:\n            state, fp = (m['state'], None)\n        else:\n            ck = torch.load(Path(path) / m['file'], map_location='cpu', weights_only=False)\n            state, fp = (ck['model'], ck.get('fingerprint'))\n        model = build_model(int(m['config']['unfreeze_last']), variant=m['config']['variant'], pool=m['config'].get('pool', 'cls_mean'), prior=bool(m['config'].get('prior', False))).to(dev)\n        model.load_state_dict(state)\n        if fp is not None:\n            check_fingerprint(model, dev, IMG, fp, tag=f\"{m['id']}: \")\n        else:\n            log(f\"  {m['id']}: no stored fingerprint (legacy bundle) -- accepted at reduced weight\")\n    t_ready = time.time()\n    jitter_seed = SEED + int(hashlib.sha256(str(m['id']).encode()).hexdigest()[:8], 16)\n    public_member = 'state' not in m\n    predicted = predict_member(model, Cte, Mte, idx, dev, IMG, starts=starts, jitter=jitter, jitter_seed=jitter_seed, return_public_frontier=public_member)\n    if public_member:\n        p, public_p, public_soft = predicted\n    else:\n        p, public_p, public_soft = (predicted, None, None)\n    t_done = time.time()\n    del model, state\n    gc.collect()\n    if dev.type == 'cuda':\n        with torch.cuda.device(dev):\n            torch.cuda.empty_cache()\n    passes = len(starts) * (2 if jitter else 1)\n    return (p, public_p, public_soft, (t_ready - t0, (t_done - t_ready) / max(passes, 1)))\n\ndef _combine(per_member):\n    all_ids = sorted({s for m in per_member for s in m['ids']})\n    pos = {s: i for i, s in enumerate(all_ids)}\n    acc = np.zeros((len(all_ids), len(TARGETS)), np.float64)\n    tot = np.zeros(len(TARGETS), np.float64)\n    for m in per_member:\n        target_weight = m.get('target_weight')\n        w = np.asarray(target_weight if target_weight is not None else [float(m.get('weight', 1.0))] * len(TARGETS), dtype=np.float64)\n        if w.shape != (len(TARGETS),) or np.any(w < 0):\n            raise ValueError(f\"invalid target weights for {m.get('id')}: {w}\")\n        r = pd.DataFrame(m['pred']).rank(pct=True).to_numpy()\n        acc[[pos[s] for s in m['ids']]] += r * w[None, :]\n        tot += w\n    if np.any(tot <= 0):\n        raise ValueError(f'at least one target has no ensemble vote: {tot}')\n    return (all_ids, acc / tot[None, :])\n\ndef combine_public_members_by_fold(per_member, pred_key='pred'):\n    # Raw-average the four members within each fold, rank each fold,\n    # then give all five folds equal weight.\n    all_ids = sorted({study for member in per_member for study in member['ids']})\n    position = {study: i for i, study in enumerate(all_ids)}\n    groups = {}\n    for i, member in enumerate(per_member):\n        fold = member.get('fold')\n        key = f'fold_{fold}' if fold is not None else f'member_{i}'\n        groups.setdefault(key, []).append(member)\n    fold_ranks, diagnostics = ([], [])\n    for key, members_in_fold in sorted(groups.items()):\n        matrices = []\n        for member in members_in_fold:\n            values = np.full((len(all_ids), len(TARGETS)), np.nan, np.float64)\n            values[[position[study] for study in member['ids']]] = np.asarray(member[pred_key], np.float64)\n            if np.isnan(values).any():\n                raise WeightsError(f\"{member.get('id')}: incomplete {pred_key} coverage\")\n            matrices.append(values)\n        raw_fold_mean = np.mean(matrices, axis=0)\n        fold_ranks.append(pd.DataFrame(raw_fold_mean).rank(method='average', pct=True).to_numpy(np.float64))\n        diagnostics.append({'ensemble_group': key, 'members': len(members_in_fold)})\n    if len(fold_ranks) != 5:\n        raise WeightsError(f'legacy branch requires five folds, found {len(fold_ranks)}')\n    return all_ids, np.mean(fold_ranks, axis=0), pd.DataFrame(diagnostics)\n\ndef blend_legacy_frontier_and_soft(frontier_rank, soft_rank):\n    output = np.asarray(frontier_rank, np.float64).copy()\n    for j, target in enumerate(TARGETS):\n        alpha = float(LEGACY_FOLD_SOFTPOOL_ALPHA.get(target, 0.0))\n        if alpha:\n            output[:, j] = (1.0 - alpha) * frontier_rank[:, j] + alpha * soft_rank[:, j]\n    return output\n\ndef infer_from_package(path, dev=None):\n    man = json.loads((Path(path) / 'manifest.json').read_text())\n    members = man['members']\n    log(f'weights package: {len(members)} member(s) from {path}; {len(DEVS)} device(s)')\n    test_df = pd.read_csv(ROOT / 'test.csv')\n    test_series = pd.read_csv(ROOT / 'test_series.csv')\n    plane_map = dict(zip(test_series['SeriesInstanceUID'], test_series['Anatomical_Plane']))\n    hte = annotate(walk('test_series'))\n    log(f'test header pass: {len(hte)} series')\n    groups = {}\n    for m in members:\n        groups.setdefault(m['pixel_group'], []).append(m)\n    groups.update(legacy_group_members())\n    per_member, public_frontier_members = ([], [])\n    est = {'fixed': None, 'win': None}\n\n    def bank(m, ids, pred, starts, jitter, public_pred=None, public_soft=None):\n        if float(np.std(pred)) < 1e-09:\n            log(f\"  {m['id']}: degenerate predictions; not banked\")\n            return\n        with STATE_LOCK:\n            per_member.append({'id': m['id'], 'fold': m.get('fold'), 'ids': ids, 'pred': pred, 'weight': m.get('weight', 1.0), 'target_weight': m.get('target_weight'), 'holdout': m.get('holdout')})\n            if public_pred is not None and len(starts) == len(starts_full):\n                if float(np.std(public_pred)) < 1e-09:\n                    raise WeightsError(f\"{m['id']}: degenerate public-frontier prediction\")\n                public_frontier_members.append({'id': m['id'], 'fold': m.get('fold'), 'ids': ids, 'pred': public_pred, 'soft_pred': public_soft})\n            elif public_pred is not None:\n                log(f\"  {m['id']}: public-frontier vote omitted because only {len(starts)} / {len(starts_full)} windows completed\")\n            all_ids, acc = _combine(per_member)\n            write_submission(acc, all_ids, test_df, 'submission.csv')\n            log(f\"  banked {m['id']} fold {m.get('fold', '?')} ({len(starts)} window(s){(', jitter' if jitter else '')}); submission.csv = weighted rank mean of {len(per_member)} member(s)\")\n    for gi, (key, gm) in enumerate(groups.items(), 1):\n        cfg = json.loads(key)\n        adopt_config_globals(cfg)\n        log(f\"decode group {gi}/{len(groups)}: {cfg['img']}px x {cfg['slices']} slices, crop {cfg['crop_mm']} mm -> {len(gm)} member(s)\")\n        st_te, Cte, Mte = build_cache(pick_slots(hte, plane_map), plane_map, lat_of(hte, 'test '), f'test g{gi}')\n        idx = np.arange(len(st_te))\n        starts_full = window_starts(Cte.shape[2], GROUP)\n        pending = sorted(gm, key=lambda m: -(m.get('holdout') or 0))\n        left_after = sum((len(g) for j, (_, g) in enumerate(groups.items(), 1) if j > gi))\n\n        def pop_next():\n            with STATE_LOCK:\n                if not pending:\n                    return (None, None, False)\n                left = TIME_BUDGET - (time.time() - T0)\n                remaining = len(pending) + left_after\n                slots_left = -(-remaining // len(DEVS))\n                starts, jit = (starts_full, False)\n                if est['fixed'] is not None and est['win'] is not None:\n                    afford = max(left * 0.9, 0.0)\n                    room = afford / max(slots_left, 1)\n                    if est['fixed'] + est['win'] > room:\n                        log(f'  {left / 60:.0f} min left: surrendering {len(pending)} member(s); not one more fits')\n                        pending.clear()\n                        return (None, None, False)\n                    jit = TTA_JITTER and est['fixed'] + 2 * len(starts_full) * est['win'] <= room * 0.6\n                    per_win = est['win'] * (2 if jit else 1)\n                    n_win = int((room - est['fixed']) / per_win) if per_win > 0 else len(starts_full)\n                    n_win = max(1, min(len(starts_full), n_win))\n                    if n_win < len(starts_full):\n                        mid = (len(starts_full) - n_win) // 2\n                        starts = starts_full[mid:mid + n_win]\n                return (pending.pop(0), starts, jit)\n\n        def worker(dev):\n            others = [d for d in DEVS if d is not dev]\n            while True:\n                m, starts, jit = pop_next()\n                if m is None:\n                    return\n                for attempt, d in enumerate([dev] + others[:1]):\n                    try:\n                        p, public_p, public_soft, (fs, ws) = _run_member(path, m, d, Cte, Mte, idx, starts, jit)\n                        with STATE_LOCK:\n                            est['fixed'], est['win'] = (fs, ws)\n                        bank(m, st_te, p, starts, jit, public_p, public_soft)\n                        break\n                    except Exception as exc:\n                        log(f\"  MEMBER {m['id']} failed on {d} ({type(exc).__name__}: {exc}); \" + ('retrying on peer device' if attempt == 0 and others else 'dropped -- costs one vote, not the run'))\n                        if d.type == 'cuda':\n                            with torch.cuda.device(d):\n                                torch.cuda.empty_cache()\n        threads = [threading.Thread(target=worker, args=(d,)) for d in DEVS]\n        for t in threads:\n            t.start()\n        for t in threads:\n            t.join()\n        del Cte, Mte\n        gc.collect()\n    if not per_member:\n        raise WeightsError('no member produced predictions; submission stays at 0.5')\n    all_ids, acc = _combine(per_member)\n    sub = write_submission(acc, all_ids, test_df, 'submission.csv')\n    log(f'final submission.csv = weighted rank mean of {len(per_member)} member(s); {sub.shape}; nulls {int(sub[TARGETS].isna().sum().sum())}')\n    globals()['V17_PUBLIC_FRONTIER'] = len(public_frontier_members) == len(members)\n    if len(public_frontier_members) == len(members):\n        frontier_ids, frontier_acc = _combine(public_frontier_members)\n        frontier_sub = write_submission(frontier_acc, frontier_ids, test_df, 'submission_public_0899.csv')\n        log(f'submission_public_0899.csv = exact no-jitter public-frontier rank mean of {len(public_frontier_members)} member(s); {frontier_sub.shape}; nulls {int(frontier_sub[TARGETS].isna().sum().sum())}')\n        fold_ids, fold_frontier, fold_diagnostics = combine_public_members_by_fold(public_frontier_members, 'pred')\n        soft_ids, fold_soft, _ = combine_public_members_by_fold(public_frontier_members, 'soft_pred')\n        if fold_ids != soft_ids:\n            raise WeightsError('legacy hard/soft study order mismatch')\n        legacy_prediction = blend_legacy_frontier_and_soft(fold_frontier, fold_soft)\n        legacy_sub = write_submission(legacy_prediction, fold_ids, test_df, 'submission_legacy_fold_blend.csv')\n        fold_diagnostics.to_csv('legacy_fold_diagnostics.csv', index=False)\n        log(f'legacy DINO aggregation written from five folds; {legacy_sub.shape}')\n    else:\n        banked = {m['id'] for m in public_frontier_members}\n        absent = [m['id'] for m in members if m['id'] not in banked]\n        log('WARNING: public-frontier fallback NOT emitted: '\n            f'{len(public_frontier_members)} / {len(members)} required public '\n            f'members completed; missing {absent[:5]}'\n            f\"{' ...' if len(absent) > 5 else ''}. submission.csv holds the \"\n            'native blend over whatever was banked -- some members at full '\n            'windows, some truncated -- which is not the lineage the Rad '\n            'calibrator and the outer routing were fitted against. Treat this '\n            'run as degraded, not as the tuned recipe.')\n    return sub\n\ndef adopt_config_globals(cfg):\n    global IMG, CACHE_IMG, GROUP, CACHE_SLICES, N_GROUP, CROP_MM, SLICE_BAND, RULES\n    CACHE_IMG = IMG = int(cfg['img'])\n    GROUP = int(cfg['group'])\n    CACHE_SLICES = int(cfg['slices'])\n    N_GROUP = max(CACHE_SLICES // GROUP, 1)\n    CROP_MM = float(cfg['crop_mm'])\n    SLICE_BAND = tuple((float(x) for x in cfg['band']))\n    rules = cfg.get('rules') or RULES_NATIVE\n    unknown = {k: v for k, v in rules.items() if k not in RULES_NATIVE or v not in (RULES_NATIVE[k], RULES_LEGACY[k])}\n    if unknown:\n        raise WeightsError(f'the members record pixel rules this pipeline cannot reproduce: {unknown}')\n    RULES = {**RULES_NATIVE, **rules}\n    if [s[0] for s in SLOTS] != list(cfg['slots']):\n        raise WeightsError(f\"the members were fitted on slots {cfg['slots']} and this pipeline defines {[s[0] for s in SLOTS]}; a weight would be read against the wrong slot\")"},{"cell_type":"markdown","id":"b165464c","metadata":{"papermill":{"duration":0.008336,"end_time":"2026-09-12T11:17:50.801461+00:00","exception":false,"start_time":"2026-09-12T11:17:50.793125+00:00","status":"completed"},"tags":[]},"source":"# 12. 🔄 Аугментация и предсказание\n\n## Что делает эта ячейка\n\n- `augment()` — лёгкие случайные повороты, сдвиги, масштаб и изменение яркости.\n- `predict()` — прогон модели по всем группам срезов и усреднение вероятностей.\n- `macro_auc()` — метрика для валидации (в инференсе не используется, но нужна для совместимости)."},{"cell_type":"code","execution_count":null,"id":"89edfc5c","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.819898Z","iopub.status.busy":"2026-09-12T11:17:50.818969Z","iopub.status.idle":"2026-09-12T11:17:50.830443Z","shell.execute_reply":"2026-09-12T11:17:50.829842Z"},"papermill":{"duration":0.021966,"end_time":"2026-09-12T11:17:50.8318+00:00","exception":false,"start_time":"2026-09-12T11:17:50.809834+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"def take_group(cache_rows, g):\n    return cache_rows[:, :, g * GROUP:(g + 1) * GROUP]\n\ndef augment(imgs, generator=None):\n    lead = imgs.shape[:-3]\n    x = imgs.reshape(-1, *imgs.shape[-3:]).float()\n    n, dev = (x.shape[0], x.device)\n    rot = (torch.rand(n, device=dev, generator=generator) - 0.5) * 2 * (AUG_ROT_DEG * np.pi / 180)\n    sc = 1.0 + torch.rand(n, device=dev, generator=generator) * AUG_SCALE\n    tx = (torch.rand(n, device=dev, generator=generator) - 0.5) * 2 * AUG_SHIFT\n    ty = (torch.rand(n, device=dev, generator=generator) - 0.5) * 2 * AUG_SHIFT\n    cos, sin = (torch.cos(rot) / sc, torch.sin(rot) / sc)\n    theta = torch.zeros(n, 2, 3, device=dev, dtype=torch.float32)\n    theta[:, 0, 0], theta[:, 0, 1], theta[:, 0, 2] = (cos, -sin, tx)\n    theta[:, 1, 0], theta[:, 1, 1], theta[:, 1, 2] = (sin, cos, ty)\n    grid = F.affine_grid(theta, x.shape, align_corners=False)\n    x = F.grid_sample(x, grid, mode='bilinear', padding_mode='border', align_corners=False)\n    scale = 1.0 + (torch.rand(n, 1, 1, 1, device=dev, generator=generator) - 0.5) * 2 * AUG_INTENSITY\n    x = (x * scale).clamp(0, 255)\n    return x.reshape(*lead, *x.shape[-3:]).to(imgs.dtype)\n\n@torch.no_grad()\ndef predict(model, cache, mask, idx, dev, img_size=None):\n    model.eval()\n    out = []\n    for b in range(0, len(idx), EVAL_BATCH):\n        sel = idx[b:b + EVAL_BATCH]\n        m = torch.from_numpy(mask[sel]).to(dev)\n        acc = None\n        for g in range(N_GROUP):\n            rows = torch.from_numpy(np.ascontiguousarray(cache[sel, :, g * GROUP:(g + 1) * GROUP])).to(dev)\n            with torch.autocast('cuda', enabled=dev.type == 'cuda'):\n                z = model(rows, m, img_size).float()\n            acc = z if acc is None else acc + z\n        out.append(torch.sigmoid(acc / N_GROUP).cpu().numpy())\n    return np.concatenate(out) if out else np.zeros((0, len(TARGETS)), np.float32)\n\ndef macro_auc(y, p):\n    from sklearn.metrics import roc_auc_score\n    return float(np.nanmean([roc_auc_score(y[:, j], p[:, j]) if len(set(y[:, j])) > 1 else np.nan for j in range(y.shape[1])]))"},{"cell_type":"code","execution_count":null,"id":"3e00bc05","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:17:50.849415Z","iopub.status.busy":"2026-09-12T11:17:50.849201Z","iopub.status.idle":"2026-09-12T11:18:26.687442Z","shell.execute_reply":"2026-09-12T11:18:26.686758Z"},"papermill":{"duration":35.849039,"end_time":"2026-09-12T11:18:26.689156+00:00","exception":false,"start_time":"2026-09-12T11:17:50.840117+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"import base64\nimport gc\nimport hashlib\nimport io\nimport json\nimport math\nimport os\nimport random\nimport time\nimport zlib\nfrom concurrent.futures import ThreadPoolExecutor\nfrom functools import lru_cache\nfrom pathlib import Path\nimport cv2\nimport joblib\nimport numpy as np\nimport pandas as pd\nimport pydicom\nfrom scipy.stats import rankdata\nfrom sklearn.ensemble import ExtraTreesClassifier, HistGradientBoostingClassifier\nfrom sklearn.decomposition import PCA\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.pipeline import make_pipeline\nfrom sklearn.preprocessing import StandardScaler\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom transformers import AutoModel"},{"cell_type":"markdown","id":"b74cf774","metadata":{"papermill":{"duration":0.008188,"end_time":"2026-09-12T11:18:26.706153+00:00","exception":false,"start_time":"2026-09-12T11:18:26.697965+00:00","status":"completed"},"tags":[]},"source":"# 15. 📝 Запись submission и валидация формата\n\n## Что делает эта ячейка\n\n- `write_submission()` — конвертирует предсказания в ранги и сохраняет CSV.\n- `write_benchmark_submission()` — создаёт fallback-файл с 0.5.\n- `_v37_validate_submission()` — проверяет порядок колонок, уникальность StudyInstanceUID и конечность значений.\n\n**Важно:** функция `main()` в этой ячейке не нужна — она относится к обучению, а не к инференсу."},{"cell_type":"code","execution_count":null,"id":"ad4ac36d","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:18:26.723776Z","iopub.status.busy":"2026-09-12T11:18:26.723375Z","iopub.status.idle":"2026-09-12T11:18:26.731233Z","shell.execute_reply":"2026-09-12T11:18:26.730641Z"},"papermill":{"duration":0.018552,"end_time":"2026-09-12T11:18:26.732862+00:00","exception":false,"start_time":"2026-09-12T11:18:26.71431+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"def write_submission(pred, studies, test_df, path):\n    sub = pd.DataFrame(pd.DataFrame(pred).rank(pct=True).values, columns=TARGETS)\n    sub.insert(0, 'StudyInstanceUID', studies)\n    sub = test_df[['StudyInstanceUID']].merge(sub, on='StudyInstanceUID', how='left')\n    sub[TARGETS] = sub[TARGETS].fillna(0.5)\n    sub.to_csv(path, index=False)\n    return sub\n\ndef write_benchmark_submission():\n    t = pd.read_csv(ROOT / 'test.csv')\n    for c in TARGETS:\n        t[c] = 0.5\n    t.to_csv('submission.csv', index=False)\n\ndef _v37_validate_submission(path, test_df, tag):\n    path = Path(path)\n    frame = pd.read_csv(path)\n    expected = ['StudyInstanceUID'] + TARGETS\n    if list(frame.columns) != expected:\n        raise ValueError(f'{tag}: columns differ from the competition contract')\n    if len(frame) != len(test_df) or not frame['StudyInstanceUID'].is_unique:\n        raise ValueError(f'{tag}: row count or StudyInstanceUID uniqueness failed')\n    if set(frame['StudyInstanceUID'].astype(str)) != set(test_df['StudyInstanceUID'].astype(str)):\n        raise ValueError(f'{tag}: StudyInstanceUID set differs from test.csv')\n    values = frame[TARGETS].to_numpy(np.float64)\n    if not np.isfinite(values).all():\n        raise ValueError(f'{tag}: non-finite prediction')\n    return test_df[['StudyInstanceUID']].merge(frame, on='StudyInstanceUID', how='left')\n\n"},{"cell_type":"markdown","id":"23b1aaca","metadata":{"papermill":{"duration":0.008085,"end_time":"2026-09-12T11:18:26.749125+00:00","exception":false,"start_time":"2026-09-12T11:18:26.74104+00:00","status":"completed"},"tags":[]},"source":"# 16. 🦕 DINOv2: создание первого submission.csv"},{"cell_type":"code","execution_count":null,"id":"902f2557","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:18:26.766778Z","iopub.status.busy":"2026-09-12T11:18:26.766565Z","iopub.status.idle":"2026-09-12T11:19:13.19682Z","shell.execute_reply":"2026-09-12T11:19:13.196154Z"},"papermill":{"duration":46.440956,"end_time":"2026-09-12T11:19:13.198453+00:00","exception":false,"start_time":"2026-09-12T11:18:26.757497+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"write_benchmark_submission()\npkg = find_weights()\nV17_PUBLIC_FRONTIER = True\nV17_DINO_PACKAGE_OK = True\nif pkg is not None:\n    dev = DEVS[0]\n    try:\n        infer_from_package(pkg, dev)\n    except Exception as package_error:\n        traceback.print_exc()\n        V17_DINO_PACKAGE_OK = False\n        V17_PUBLIC_FRONTIER = False\n        log(f'DINOv2 package inference ABANDONED '\n            f'({type(package_error).__name__}: {package_error}); '\n            'submission.csv keeps the last file banked before the failure and '\n            'the DINOv3, RadImageNet and Raptor branches below still run. '\n            'This run is NOT the tuned recipe.')\n    try:\n        test_df = pd.read_csv(ROOT / 'test.csv')\n        native_path = Path('submission.csv')\n        public_path = Path('submission_public_0899.csv')\n        native = _v37_validate_submission(native_path, test_df, 'native 24-member')\n        public = _v37_validate_submission(public_path, test_df, 'public DINO frontier')\n        native.to_csv('submission_native_v38.csv', index=False)\n        public.to_csv(native_path, index=False)\n        promoted = _v37_validate_submission(native_path, test_df, 'V40 primary')\n        if not promoted.equals(public):\n            raise AssertionError('V40 serialization differs from validated public frontier')\n        log('V40 primary = exact no-jitter public-frontier target pooling; native 24-member output retained')\n    except Exception as public_frontier_error:\n        V17_PUBLIC_FRONTIER = False\n        log('WARNING: the public DINO frontier was NOT promoted into '\n            f'submission.csv ({type(public_frontier_error).__name__}: '\n            f'{public_frontier_error}). submission.csv holds the native blend '\n            'instead -- a lineage that has never been scored -- while the Rad '\n            'calibrator and the outer routing below apply constants fitted '\n            'against the frontier. This run is NOT the tuned recipe.')\n        traceback.print_exc()\n    log('done')\nelse:\n    raise RuntimeError('DINOv2 weights not found')"},{"cell_type":"markdown","id":"563f745c","metadata":{"papermill":{"duration":0.012217,"end_time":"2026-09-12T11:19:13.22422+00:00","exception":false,"start_time":"2026-09-12T11:19:13.212003+00:00","status":"completed"},"tags":[]},"source":"# 17. 🦖 DINOv3: инициализация\n\n## Что делает эта ячейка\n\nГотовит параметры для DINOv3 ветки:\n\n- 6 слотов × 16 срезов.\n- Ищет датасет с фолдами `knee-mri-fold-weights`.\n- Определяет устройство (CUDA/CPU).\n\nСама загрузка моделей будет в следующей ячейке."},{"cell_type":"code","execution_count":null,"id":"da0faebf","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:13.250368Z","iopub.status.busy":"2026-09-12T11:19:13.249636Z","iopub.status.idle":"2026-09-12T11:19:17.368774Z","shell.execute_reply":"2026-09-12T11:19:17.368064Z"},"papermill":{"duration":4.134378,"end_time":"2026-09-12T11:19:17.370725+00:00","exception":false,"start_time":"2026-09-12T11:19:13.236347+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"_A5_SAVED = dict(globals())\nimport gc, os, time, warnings\nfrom concurrent.futures import ProcessPoolExecutor, as_completed\nfrom pathlib import Path\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport timm\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nwarnings.filterwarnings('ignore')\ncv2.setNumThreads(1)\nCROP_MM = 130.0\nSIZE = 336\nSLICE_BAND = (0.12, 0.88)\nN_SLICE = 16\nINTENSITY = 'slice'\nSLOTS = [('Sagittal', 1), ('Sagittal', 0), ('Coronal', 1), ('Coronal', 0), ('Axial', 1), ('Axial', 0)]\nN_SLOT = len(SLOTS)\nLABELS = ['ACL', 'MCL', 'Medial Meniscus', 'Lateral Meniscus', 'Medial OA', 'Lateral OA', 'PF OA', 'Effusion', 'Synovitis', \"Baker's\", 'Contusion', 'Fracture']\n\ndef _find_dir(*names):\n    root = Path('/kaggle/input')\n    cand = []\n    for n in names:\n        cand += [root / n, root / 'competitions' / n, root / 'datasets' / n]\n        for parent in (root / 'datasets', root / 'competitions', root):\n            if parent.is_dir():\n                try:\n                    cand += [d / n for d in parent.iterdir() if d.is_dir()]\n                except OSError:\n                    pass\n    for p in cand:\n        if p.is_dir():\n            return p\n    return None\nCOMP = _find_dir('rsna-knee-abnormality-detection')\nCKPT = _find_dir('knee-mri-fold-weights')\nassert COMP is not None, 'competition data not attached'\nassert CKPT is not None, 'fold weights not attached'\nassert (COMP / 'sample_submission.csv').exists(), f'no competition data at {COMP}'\nassert list(CKPT.glob('*_f*.pt')), f'no checkpoints at {CKPT}'\n_V17_TIMM_PIN = '1.0.22'\n\n\ndef _v17_pin_timm():\n    import importlib\n    import subprocess\n    import sys as _sys\n    global timm\n    wheels = sorted(CKPT.glob(f'timm-{_V17_TIMM_PIN}-*.whl'))\n    if not wheels:\n        print(f'[timm] no timm-{_V17_TIMM_PIN} wheel under {CKPT}; running the '\n              f'base image timm {timm.__version__} unpinned', flush=True)\n        return False\n    target = '/kaggle/working/_timm_pin'\n\n    def reimport():\n        global timm\n        for name in [n for n in list(_sys.modules)\n                     if n == 'timm' or n.startswith('timm.')]:\n            del _sys.modules[name]\n        timm = importlib.import_module('timm')\n\n    try:\n        subprocess.run([_sys.executable, '-m', 'pip', 'install', '--no-deps',\n                        '--quiet', '--target', target, str(wheels[0])],\n                       check=True)\n        _sys.path.insert(0, target)\n        reimport()\n        assert timm.__version__ == _V17_TIMM_PIN, timm.__version__\n        timm.create_model('resnet18', pretrained=False, num_classes=0)\n    except Exception as error:\n        print(f'[timm] PIN FAILED ({type(error).__name__}: {error}); falling '\n              'back to the base image timm', flush=True)\n        while target in _sys.path:\n            _sys.path.remove(target)\n        reimport()\n        print(f'[timm] running unpinned timm {timm.__version__}', flush=True)\n        return False\n    print(f'[timm] pinned to {timm.__version__} from {wheels[0].name} '\n          '(no internet needed; the wheel ships with the fold weights)',\n          flush=True)\n    return True\n\n\n_V17_TIMM_PINNED = _v17_pin_timm()\nDEV = 'cuda' if torch.cuda.is_available() else 'cpu'\nprint(f'competition : {COMP}')\nprint(f'checkpoints : {CKPT}')\nprint(f'device      : {DEV}')\nfor i in range(torch.cuda.device_count() if DEV == 'cuda' else 0):\n    cc = torch.cuda.get_device_capability(i)\n    print(f'  gpu{i}       : {torch.cuda.get_device_name(i)} sm_{cc[0]}{cc[1]}, {torch.cuda.get_device_properties(i).total_memory / 2 ** 30:.0f} GiB, native bf16={cc >= (8, 0)}')"},{"cell_type":"markdown","id":"d6f79d79","metadata":{"papermill":{"duration":0.012816,"end_time":"2026-09-12T11:19:17.397305+00:00","exception":false,"start_time":"2026-09-12T11:19:17.384489+00:00","status":"completed"},"tags":[]},"source":"# 18. 🦖 DINOv3: подготовка исследований\n\n## Что делает эта ячейка\n\nЗагружает тестовые исследования и строит volume для каждого:\n\n- 6 слотов × 16 срезов.\n- Сортировка срезов по InstanceNumber.\n- Нормализация laterality.\n- Кадрирование 130 мм, ресайз до 336×336.\n\nРезультат — массив `(N_SLOT, N_SLICE, SIZE, SIZE)` и маска слотов."},{"cell_type":"code","execution_count":null,"id":"880456b7","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:17.426713Z","iopub.status.busy":"2026-09-12T11:19:17.425877Z","iopub.status.idle":"2026-09-12T11:19:17.468411Z","shell.execute_reply":"2026-09-12T11:19:17.467459Z"},"papermill":{"duration":0.058748,"end_time":"2026-09-12T11:19:17.470066+00:00","exception":false,"start_time":"2026-09-12T11:19:17.411318+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"SERIES_ROOT = COMP / 'test_series'\nif not SERIES_ROOT.exists():\n    SERIES_ROOT = COMP / 'train_series'\nprint('series root:', SERIES_ROOT)\n\ndef ordered_files(sdir, cap=64):\n    keyed = []\n    for f in sdir.glob('*.dcm'):\n        try:\n            ds = pydicom.dcmread(str(f), stop_before_pixels=True)\n            keyed.append((int(ds.InstanceNumber), str(f)))\n        except Exception:\n            continue\n        if len(keyed) >= cap * 4:\n            break\n    return [f for _, f in sorted(keyed)]\n\ndef series_side(path):\n    try:\n        return float(pydicom.dcmread(path, stop_before_pixels=True).ImagePositionPatient[0])\n    except Exception:\n        return 0.0\n\ndef read_crop(path):\n    try:\n        ds = pydicom.dcmread(path)\n        arr = ds.pixel_array.astype(np.float32)\n    except Exception:\n        return None\n    try:\n        ps = float(ds.PixelSpacing[0])\n    except Exception:\n        ps = CROP_MM / max(arr.shape)\n    half = int(round(CROP_MM / ps / 2))\n    cy, cx = (arr.shape[0] // 2, arr.shape[1] // 2)\n    y0, y1 = (max(0, cy - half), min(arr.shape[0], cy + half))\n    x0, x1 = (max(0, cx - half), min(arr.shape[1], cx + half))\n    crop = arr[y0:y1, x0:x1]\n    return None if crop.size == 0 else crop\n\ndef window(crop, lo, hi, flip):\n    c = np.clip((crop - lo) / max(hi - lo, 1e-06), 0, 1)\n    img = cv2.resize(c, (SIZE, SIZE), interpolation=cv2.INTER_AREA)\n    return img[:, ::-1].copy() if flip else img\n\ndef render(path, flip):\n    crop = read_crop(path)\n    if crop is None:\n        return None\n    lo, hi = np.percentile(crop[::4, ::4], [1, 99])\n    return window(crop, lo, hi, flip)\n\ndef build_study(args):\n    idx, study, recs = args\n    out = np.zeros((N_SLOT, N_SLICE, SIZE, SIZE), np.uint8)\n    mask = np.zeros(N_SLOT, np.uint8)\n    rows = pd.DataFrame(recs)\n    if len(rows):\n        for s_i, (plane, fs) in enumerate(SLOTS):\n            sub = rows[(rows.Anatomical_Plane == plane) & (rows.Fat_Suppression == fs)]\n            if sub.empty:\n                continue\n            files = ordered_files(SERIES_ROOT / study / sub.iloc[0].SeriesInstanceUID)\n            if not files:\n                continue\n            flip = plane != 'Sagittal' and series_side(files[0]) < 0\n            lo, hi = SLICE_BAND\n            i0 = int(round(lo * (len(files) - 1)))\n            i1 = int(round(hi * (len(files) - 1)))\n            avail = list(range(i0, i1 + 1))\n            if len(avail) >= N_SLICE:\n                picks = [avail[int(round(t))] for t in np.linspace(0, len(avail) - 1, N_SLICE)]\n                off = 0\n            else:\n                picks, off = (avail, (N_SLICE - len(avail)) // 2)\n            if INTENSITY == 'series':\n                crops = [read_crop(files[p]) for p in picks]\n                got = [x for x in crops if x is not None]\n                if got:\n                    samp = np.concatenate([x[::4, ::4].ravel() for x in got])\n                    lo_, hi_ = np.percentile(samp, [1, 99])\n                    for c, x in enumerate(crops):\n                        if x is None:\n                            x = read_crop(files[min(len(files) - 1, picks[c] + 1)])\n                        if x is not None:\n                            out[s_i, off + c] = (window(x, lo_, hi_, flip) * 255).astype(np.uint8)\n            else:\n                for c, p in enumerate(picks):\n                    img = render(files[p], flip)\n                    if img is None:\n                        img = render(files[min(len(files) - 1, p + 1)], flip)\n                    if img is not None:\n                        out[s_i, off + c] = (img * 255).astype(np.uint8)\n            mask[s_i] = len(picks)\n    return (idx, out, mask)\nsub_df = pd.read_csv(COMP / 'sample_submission.csv')\nser_csv = pd.read_csv(COMP / 'test_series.csv')\nif not (COMP / 'test_series').exists():\n    ser_csv = pd.read_csv(COMP / 'train_series.csv')\nser_csv = ser_csv.loc[:, ~ser_csv.columns.duplicated()]\nstudies = sub_df.StudyInstanceUID.tolist()\nby = {s: g.to_dict('records') for s, g in ser_csv[ser_csv.StudyInstanceUID.isin(set(studies))].groupby('StudyInstanceUID')}\nprint(f'{len(studies):,} test studies, {len(by):,} with series metadata')"},{"cell_type":"markdown","id":"c585d78a","metadata":{"papermill":{"duration":0.016071,"end_time":"2026-09-12T11:19:17.500368+00:00","exception":false,"start_time":"2026-09-12T11:19:17.484297+00:00","status":"completed"},"tags":[]},"source":"# 19. 🦖 DINOv3: архитектура и загрузка моделей\n\n## Что делает эта ячейка\n\nОпределяет все компоненты DINOv3 ветки:\n\n- Различные пулинги (MeanMax, Attention, TokenX, Codex).\n- Механизмы смешивания слотов по глубине.\n- Основную модель `Net`.\n- Загружает 5 фолдов `m_f0..m_f4` из чекпоинтов.\n\nПосле этой ячейки `models` содержит 5 готовых DINOv3 моделей."},{"cell_type":"code","execution_count":null,"id":"75f8922c","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:17.530627Z","iopub.status.busy":"2026-09-12T11:19:17.53027Z","iopub.status.idle":"2026-09-12T11:19:24.706568Z","shell.execute_reply":"2026-09-12T11:19:24.705453Z"},"papermill":{"duration":7.193844,"end_time":"2026-09-12T11:19:24.708725+00:00","exception":false,"start_time":"2026-09-12T11:19:17.514881+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"N_SLOT_TYPES, MASK_IDX = (6, 0)\n\ndef segment_softmax(scores, sidx, B):\n    T, K = scores.shape\n    idx = sidx.unsqueeze(1).expand(-1, K)\n    m = torch.full((B, K), float('-inf'), device=scores.device, dtype=scores.dtype)\n    m = m.scatter_reduce(0, idx, scores, reduce='amax', include_self=True)\n    e = (scores - m[sidx]).exp()\n    s = torch.zeros(B, K, device=scores.device, dtype=scores.dtype).index_add_(0, sidx, e)\n    return e / s[sidx].clamp(min=1e-06)\n\nclass MeanMaxPool(nn.Module):\n\n    def forward(self, f, sidx, B, slot=None, return_attn=False):\n        D = f.shape[1]\n        cnt = torch.zeros(B, device=f.device, dtype=f.dtype).index_add_(0, sidx, torch.ones(f.shape[0], device=f.device, dtype=f.dtype))\n        mean = torch.zeros(B, D, device=f.device, dtype=f.dtype).index_add_(0, sidx, f)\n        mean = mean / cnt.clamp(min=1).unsqueeze(1)\n        mx = torch.full((B, D), -10000.0, device=f.device, dtype=f.dtype)\n        mx = mx.scatter_reduce(0, sidx.unsqueeze(1).expand(-1, D), f, reduce='amax', include_self=True)\n        return (torch.cat([mean, mx], 1), None)\n\nclass LabelAttentionPool(nn.Module):\n\n    def __init__(self, d, n_labels=12, n_heads=4, slot_bias=True):\n        super().__init__()\n        self.d, self.k, self.h = (d, n_labels, n_heads)\n        self.q = nn.Parameter(torch.randn(n_labels, d) * 0.02)\n        self.key, self.val = (nn.Linear(d, d), nn.Linear(d, d))\n        self.slot_bias = nn.Parameter(torch.zeros(n_labels, N_SLOT_TYPES + 1)) if slot_bias else None\n\n    def forward(self, f, sidx, B, slot=None, return_attn=False):\n        scores = self.key(f) @ self.q.t() / self.d ** 0.5\n        if self.slot_bias is not None and slot is not None:\n            scores = scores + self.slot_bias.t()[slot]\n        a = segment_softmax(scores, sidx, B)\n        out = torch.zeros(B, self.k, self.d, device=f.device, dtype=f.dtype)\n        out = out.index_add_(0, sidx, a.unsqueeze(-1) * self.val(f).unsqueeze(1))\n        return (out, a)\n\nclass TokenXAttnPool(nn.Module):\n\n    def __init__(self, d, n_labels=12, n_heads=6, dropout=0.2):\n        super().__init__()\n        self.d, self.k = (d, n_labels)\n        self.q = nn.Parameter(torch.randn(n_labels, d) * 0.02)\n        self.slot_emb = nn.Embedding(N_SLOT_TYPES + 1, d, padding_idx=0)\n        self.kv_norm = nn.LayerNorm(d)\n        self.attn = nn.MultiheadAttention(d, n_heads, dropout=dropout, batch_first=True)\n\n    def forward(self, tok, sidx, B, slot=None, return_attn=False):\n        T, N, D = tok.shape\n        cnt = torch.bincount(sidx, minlength=B)\n        S = int(cnt.max().item())\n        starts = torch.cumsum(cnt, 0) - cnt\n        pos = torch.arange(T, device=tok.device) - starts[sidx]\n        kv = tok + self.slot_emb(slot).unsqueeze(1)\n        pad = tok.new_zeros(B, S, N, D)\n        pad[sidx, pos] = kv\n        keep = torch.zeros(B, S, dtype=torch.bool, device=tok.device)\n        keep[sidx, pos] = True\n        kpm = ~keep.repeat_interleave(N, dim=1)\n        pad = self.kv_norm(pad.reshape(B, S * N, D))\n        q = self.q.unsqueeze(0).expand(B, -1, -1)\n        att, w = self.attn(q, pad, pad, key_padding_mask=kpm, need_weights=return_attn, average_attn_weights=True)\n        cls = tok[:, 0]\n        mean = torch.zeros(B, D, device=tok.device, dtype=tok.dtype).index_add_(0, sidx, cls) / cnt.clamp(min=1).unsqueeze(1)\n        mx = torch.full((B, D), -10000.0, device=tok.device, dtype=tok.dtype)\n        mx = mx.scatter_reduce(0, sidx.unsqueeze(1).expand(-1, D), cls, reduce='amax', include_self=True)\n        base = torch.cat([mean, mx], 1).unsqueeze(1).expand(-1, self.k, -1)\n        return (torch.cat([att, base], -1), w)\n\nclass ViTSlotToken(nn.Module):\n\n    def __init__(self, vit, n_cat, dim=None):\n        super().__init__()\n        self.vit = vit\n        d = dim or vit.embed_dim\n        self.tok = nn.Embedding(n_cat + 1, d, padding_idx=MASK_IDX)\n        self.num_features = vit.num_features\n        self._orig_prefix = getattr(vit, 'num_prefix_tokens', 1)\n        vit.num_prefix_tokens = self._orig_prefix + 1\n        for blk in vit.blocks:\n            a = getattr(blk, 'attn', None)\n            if a is not None and hasattr(a, 'num_prefix_tokens'):\n                a.num_prefix_tokens = a.num_prefix_tokens + 1\n\n    @staticmethod\n    def _maybe(mod, x):\n        return x if mod is None else mod(x)\n\n    def forward_features(self, x, cat):\n        v = self.vit\n        x = v.patch_embed(x)\n        pos = v._pos_embed(x)\n        rope = None\n        if isinstance(pos, tuple):\n            x, rope = pos\n        else:\n            x = pos\n        x = self._maybe(getattr(v, 'patch_drop', None), x)\n        x = self._maybe(getattr(v, 'norm_pre', None), x)\n        npt = self._orig_prefix\n        tok = self.tok(cat).unsqueeze(1)\n        x = torch.cat([x[:, :npt], tok, x[:, npt:]], dim=1)\n        if rope is not None:\n            if getattr(v, 'rope_mixed', False):\n                for i, blk in enumerate(v.blocks):\n                    x = blk(x, rope=rope[i])\n            else:\n                for blk in v.blocks:\n                    x = blk(x, rope=rope)\n        else:\n            x = v.blocks(x)\n        return v.norm(x)\n\n    def forward_head(self, x, pre_logits=True):\n        return self.vit.forward_head(x, pre_logits=pre_logits)\nIMAGENET_MEAN = (0.485, 0.456, 0.406)\nIMAGENET_STD = (0.229, 0.224, 0.225)\n\nclass _GatedDepthBlock(nn.Module):\n\n    def __init__(self, n_slice, dropout=0.0, ls_init=0.1):\n        super().__init__()\n        self.norm = nn.GroupNorm(1, n_slice)\n        self.v = nn.Conv2d(n_slice, n_slice, 1)\n        self.g = nn.Conv2d(n_slice, n_slice, 1)\n        self.out = nn.Conv2d(n_slice, n_slice, 1)\n        self.gamma = nn.Parameter(torch.full((n_slice, 1, 1), ls_init))\n        self.drop = nn.Dropout2d(dropout) if dropout else nn.Identity()\n\n    def forward(self, x):\n        z = self.norm(x)\n        return x + self.gamma * self.drop(self.out(self.v(z) * F.silu(self.g(z))))\n\nclass DepthCompress(nn.Module):\n\n    def __init__(self, n_slice=16, out_ch=3, depth=1, dropout=0.0, ls_init=0.1, imagenet=True, proj_noise=0.25):\n        super().__init__()\n        self.imagenet = imagenet\n        self.blocks = nn.ModuleList([_GatedDepthBlock(n_slice, dropout, ls_init) for _ in range(depth)])\n        self.proj = nn.Conv2d(n_slice, out_ch, 1, bias=True)\n        if imagenet:\n            self.register_buffer('mu', torch.tensor(IMAGENET_MEAN).view(1, -1, 1, 1))\n            self.register_buffer('sd', torch.tensor(IMAGENET_STD).view(1, -1, 1, 1))\n\n    def forward(self, x):\n        keep = (x.amax(dim=1, keepdim=True) > 0).to(x.dtype)\n        z = x\n        for b in self.blocks:\n            z = b(z)\n        z = self.proj(z)\n        if self.imagenet:\n            z = (z - self.mu.to(z.dtype)) / self.sd.to(z.dtype)\n        return z * keep\nN_PLANE, N_CONTRAST = (3, 2)\n_PLANE_OF = lambda s: torch.clamp(s - 1, 0, 5) // 2\n_CONTRAST_OF = lambda s: torch.clamp(s - 1, 0, 5) % 2\n\nclass SlotDepthMixer(nn.Module):\n\n    def __init__(self, n_slice=16, ksize=5, alpha_max=0.25):\n        super().__init__()\n        self.n_slice, self.ksize, self.r = (n_slice, ksize, ksize // 2)\n        self.alpha_max = alpha_max\n        b = torch.tensor([1.0, 4.0, 6.0, 4.0, 1.0])\n        self.register_buffer('base', b.log()[self.r:])\n        n_u = self.r + 1\n        self.shared = nn.Parameter(torch.zeros(n_u))\n        self.plane_k = nn.Parameter(torch.zeros(N_PLANE, n_u))\n        self.contrast_k = nn.Parameter(torch.zeros(N_CONTRAST, n_u))\n        self.g0 = nn.Parameter(torch.zeros(()))\n        self.gate_p = nn.Parameter(torch.zeros(N_PLANE))\n        self.gate_c = nn.Parameter(torch.zeros(N_CONTRAST))\n        idx = torch.arange(n_slice)\n        self.register_buffer('off', idx[None, :] - idx[:, None])\n\n    def kernel(self, slot):\n        p, c = (_PLANE_OF(slot), _CONTRAST_OF(slot))\n        half = self.base + self.shared + self.plane_k[p] + self.contrast_k[c]\n        full = torch.cat([half.flip(-1)[..., :self.r], half], dim=-1)\n        return F.softmax(full, dim=-1)\n\n    def alpha(self, slot):\n        p, c = (_PLANE_OF(slot), _CONTRAST_OF(slot))\n        return self.alpha_max * torch.tanh(self.g0 + self.gate_p[p] + self.gate_c[c])\n\n    def forward(self, x, slot, vmask):\n        T, S, H, W = x.shape\n        if vmask is None:\n            raise ValueError('stem=mixer requires the padding mask')\n        k = self.kernel(slot)\n        v = vmask.to(k.dtype)\n        d = self.off + self.r\n        inb = (d >= 0) & (d < self.ksize)\n        kk = k[:, d.clamp(0, self.ksize - 1)] * inb\n        M = kk * v[:, None, :]\n        den = M.sum(-1, keepdim=True)\n        eye = torch.eye(S, device=x.device, dtype=M.dtype).expand(T, S, S)\n        ok = (den > 1e-06) & v[:, :, None].bool()\n        M = torch.where(ok, M / den.clamp(min=1e-06), eye)\n        a = self.alpha(slot)[:, None, None]\n        Aop = ((1.0 - a) * eye + a * M).to(x.dtype)\n        if x.is_contiguous(memory_format=torch.channels_last) and (not x.is_contiguous()):\n            y = torch.bmm(x.permute(0, 2, 3, 1).reshape(T, H * W, S), Aop.transpose(1, 2))\n            return y.reshape(T, H, W, S).permute(0, 3, 1, 2)\n        return torch.bmm(Aop, x.reshape(T, S, H * W)).reshape(T, S, H, W)\n\ndef _seg_mean_max(v, sidx, B):\n    D = v.shape[1]\n    cnt = torch.zeros(B, device=v.device, dtype=v.dtype).index_add_(0, sidx, torch.ones(v.shape[0], device=v.device, dtype=v.dtype))\n    mean = torch.zeros(B, D, device=v.device, dtype=v.dtype).index_add_(0, sidx, v)\n    mean = mean / cnt.clamp(min=1).unsqueeze(1)\n    mx = torch.full((B, D), -10000.0, device=v.device, dtype=v.dtype)\n    mx = mx.scatter_reduce(0, sidx.unsqueeze(1).expand(-1, D), v, reduce='amax', include_self=True)\n    return torch.cat([mean, mx], 1)\n\ndef _pad_kv(x, sidx, B, norm):\n    T, P, D = x.shape\n    cnt = torch.bincount(sidx, minlength=B)\n    S = int(cnt.max().item())\n    starts = torch.cumsum(cnt, 0) - cnt\n    pos = torch.arange(T, device=x.device) - starts[sidx]\n    pad = x.new_zeros(B, S, P, D)\n    pad[sidx, pos] = x\n    keep = torch.zeros(B, S, dtype=torch.bool, device=x.device)\n    keep[sidx, pos] = True\n    return (norm(pad.reshape(B, S * P, D)), ~keep.repeat_interleave(P, dim=1))\n\nclass _GatedDelta(nn.Module):\n\n    def __init__(self, d, n_labels, n_heads, dropout):\n        super().__init__()\n        self.q = nn.Parameter(torch.randn(n_labels, d) * 0.02)\n        self.kv_norm = nn.LayerNorm(d)\n        self.attn = nn.MultiheadAttention(d, n_heads, dropout=dropout, batch_first=True)\n        self.d_norm = nn.LayerNorm(d)\n        self.dw = nn.Parameter(torch.randn(n_labels, d) * (1.0 / d ** 0.5))\n        self.db = nn.Parameter(torch.zeros(n_labels))\n        self.gate = nn.Parameter(torch.zeros(n_labels))\n\n    def delta(self, pat, sidx, B, return_attn):\n        kv, kpm = _pad_kv(pat, sidx, B, self.kv_norm)\n        q = self.q.unsqueeze(0).expand(B, -1, -1)\n        att, w = self.attn(q, kv, kv, key_padding_mask=kpm, need_weights=return_attn, average_attn_weights=True)\n        return ((self.d_norm(att) * self.dw).sum(-1) + self.db, w)\n\nclass TokenResidualPool(_GatedDelta):\n\n    def __init__(self, d, n_labels=12, n_heads=6, pe=64, dropout=0.2):\n        super().__init__(d, n_labels, n_heads, dropout)\n        self.base = nn.Sequential(nn.LayerNorm(2 * d + pe), nn.Dropout(dropout), nn.Linear(2 * d + pe, n_labels))\n\n    def forward(self, tok, slot, sidx, B, pres, return_attn=False):\n        base = self.base(torch.cat([_seg_mean_max(tok[:, 1:].mean(1), sidx, B), pres], 1))\n        d_, w = self.delta(tok[:, 1:], sidx, B, return_attn)\n        return (base + self.gate * d_, w)\n\nclass CodexResidualPool(_GatedDelta):\n\n    def __init__(self, d, n_labels=12, n_heads=6, pe=64, dropout=0.2):\n        super().__init__(d, n_labels, n_heads, dropout)\n        self.base = nn.Sequential(nn.LayerNorm(2 * d + pe), nn.Dropout(dropout), nn.Linear(2 * d + pe, n_labels))\n\n    def forward(self, tok, slot, sidx, B, pres, return_attn=False):\n        base = self.base(torch.cat([_seg_mean_max(tok[:, 0], sidx, B), pres], 1))\n        d_, w = self.delta(tok[:, 1:], sidx, B, return_attn)\n        return (base + self.gate * d_, w)\n\nclass ClsAddPool(nn.Module):\n\n    def __init__(self, d, n_labels=12, pe=64, dropout=0.2):\n        super().__init__()\n        self.net = nn.Sequential(nn.LayerNorm(4 * d + pe), nn.Dropout(dropout), nn.Linear(4 * d + pe, n_labels))\n\n    def forward(self, tok, slot, sidx, B, pres, return_attn=False):\n        return (self.net(torch.cat([_seg_mean_max(tok[:, 1:].mean(1), sidx, B), _seg_mean_max(tok[:, 0], sidx, B), pres], 1)), None)\n\nclass Readout(nn.Module):\n\n    def __init__(self, pool, d, n_labels=12, pe=64):\n        super().__init__()\n        self.pool_kind, self.k = (pool, n_labels)\n        self.pres_emb = nn.Embedding(N_SLOT_TYPES + 1, pe, padding_idx=0)\n        if pool in ('xres', 'clsadd', 'xcodex'):\n            self.pool = {'xres': TokenResidualPool, 'clsadd': ClsAddPool, 'xcodex': CodexResidualPool}[pool](d, n_labels, pe=pe)\n        elif pool in ('attn', 'xattn'):\n            if pool == 'xattn':\n                self.pool = TokenXAttnPool(d, n_labels)\n                wd = 3 * d + pe\n            else:\n                self.pool = LabelAttentionPool(d, n_labels)\n                wd = d + pe\n            self.norm = nn.LayerNorm(wd)\n            self.w = nn.Parameter(torch.randn(n_labels, wd) * (1.0 / wd ** 0.5))\n            self.b = nn.Parameter(torch.zeros(n_labels))\n        else:\n            self.pool = MeanMaxPool()\n            self.net = nn.Sequential(nn.LayerNorm(2 * d + pe), nn.Dropout(0.2), nn.Linear(2 * d + pe, n_labels))\n        self.drop = nn.Dropout(0.2)\n\n    def forward(self, f, slot, sidx, B, return_attn=False):\n        pe = self.pres_emb(slot)\n        pres = torch.zeros(B, pe.shape[1], device=f.device, dtype=f.dtype).index_add_(0, sidx, pe)\n        if self.pool_kind in ('xres', 'clsadd', 'xcodex'):\n            return self.pool(f, slot, sidx, B, pres)[0]\n        pooled, attn = self.pool(f, sidx, B, slot=slot, return_attn=return_attn)\n        if self.pool_kind in ('attn', 'xattn'):\n            x = torch.cat([pooled, pres.unsqueeze(1).expand(-1, self.k, -1)], -1)\n            x = self.drop(self.norm(x))\n            return (x * self.w).sum(-1) + self.b\n        return self.net(torch.cat([pooled, pres], 1))\n\nclass Net(nn.Module):\n\n    def __init__(self, enc, cond, n_meta=0, pool='mean_max', stem='native', n_slice=16):\n        super().__init__()\n        self.enc, self.cond = (enc, cond)\n        self.compress = DepthCompress(n_slice, 3) if stem == 'compress' else None\n        self.mixer = SlotDepthMixer(n_slice) if stem == 'mixer' else None\n        self.tokens = pool in ('xattn', 'xres', 'clsadd', 'xcodex')\n        D = enc.num_features\n        self.meta_mlp = nn.Sequential(nn.LayerNorm(n_meta), nn.Linear(n_meta, 128), nn.GELU(), nn.Linear(128, D)) if n_meta > 0 else None\n        self.readout = Readout(pool, D)\n        if cond == 'post':\n            self.slot_emb = nn.Embedding(N_SLOT_TYPES + 1, D, padding_idx=MASK_IDX)\n\n    def forward(self, im, slot, smeta, sidx, B, vm=None):\n        if self.mixer is not None:\n            im = self.mixer(im, slot, vm)\n        if self.compress is not None:\n            im = self.compress(im)\n        f = self.enc.forward_features(im, slot) if self.cond == 'token' else self.enc.forward_features(im)\n        if self.tokens:\n            inner = getattr(self.enc, 'vit', self.enc)\n            orig = getattr(self.enc, '_orig_prefix', getattr(inner, 'num_prefix_tokens', 1))\n            f = torch.cat([f[:, :1], f[:, orig:]], 1)\n        else:\n            f = self.enc.forward_head(f, pre_logits=True)\n            if f.dim() > 2:\n                f = f.flatten(1)\n        ex = (lambda v: v.unsqueeze(1)) if self.tokens else lambda v: v\n        if self.cond == 'post':\n            f = f + ex(self.slot_emb(slot))\n        if self.meta_mlp is not None and smeta.shape[1] > 0:\n            mt = self.meta_mlp(smeta)\n            f = torch.cat([f, mt.unsqueeze(1)], 1) if self.tokens else f + mt\n        return self.readout(f, slot, sidx, B)\n_A5_ENC_OK = True\nmodels = []\nfor ckpt_path in sorted(CKPT.glob('*_f*.pt')):\n    z = torch.load(ckpt_path, map_location='cpu', weights_only=False)\n    cfg = z['cfg']\n    _stem = cfg.get('stem', 'native')\n    _in = 3 if _stem == 'compress' else cfg.get('n_slice', 16)\n    enc = timm.create_model(cfg['backbone'], pretrained=False, num_classes=0, in_chans=_in, **{'img_size': cfg['img']} if 'vit_' in cfg['backbone'] else {})\n    if cfg['cond'] == 'token':\n        enc = ViTSlotToken(enc, N_SLOT_TYPES)\n    m = Net(enc, cfg['cond'], cfg.get('n_meta', 0), cfg['pool'], stem=_stem, n_slice=cfg.get('n_slice', 16))\n    missing, unexpected = m.load_state_dict(z['state_dict'], strict=False)\n    assert not [k for k in missing if not k.startswith('enc.')], f'missing {missing[:5]}'\n    assert not unexpected, f'unexpected {unexpected[:5]}'\n    _enc_keys = [k for k in m.state_dict() if k.startswith('enc.')]\n    _enc_missing = [k for k in missing if k.startswith('enc.')]\n    _enc_loaded = len(_enc_keys) - len(_enc_missing)\n    if _enc_missing:\n        _A5_ENC_OK = False\n        print(f'  ENCODER DRIFT: {ckpt_path.name} filled {_enc_loaded} of '\n              f'{len(_enc_keys)} encoder tensors; {len(_enc_missing)} stayed '\n              f'at their pretrained=False random init (first '\n              f'{_enc_missing[:3]}). The installed timm {timm.__version__} '\n              'builds modules these folds were not fitted with. The branch is '\n              'disabled below rather than blended in at A5_W = 0.45.',\n              flush=True)\n    models.append(m.eval())\n    print(f\"loaded {ckpt_path.name}  fold {z['fold']}  {cfg['backbone']} pool={cfg['pool']} meta={cfg['meta']}  enc {_enc_loaded}/{len(_enc_keys)}\")\nCFG = cfg\nassert CFG.get('n_meta', 0) == 0, f\"checkpoint expects {CFG['n_meta']} metadata features -- build slot_meta for the TEST studies and pass it to predict() before submitting\"\nprint(f\"\\n{len(models)} fold models ready | input norm: {CFG.get('norm', 'none')}\")"},{"cell_type":"markdown","id":"4603b33b","metadata":{"papermill":{"duration":0.015144,"end_time":"2026-09-12T11:19:24.742211+00:00","exception":false,"start_time":"2026-09-12T11:19:24.727067+00:00","status":"completed"},"tags":[]},"source":"# 20. 🦖 DINOv3: инференс и ранговое усреднение\n\n## Что делает эта ячейка\n\n- Выполняет инференс всех 5 фолдов DINOv3.\n- Считает ранги предсказаний.\n- Усредняет ранги фолдов.\n- Сохраняет результат в `A5_PREDS`.\n\nЭто вторая ветка пайплайна."},{"cell_type":"code","execution_count":null,"id":"9f59048e","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:24.774515Z","iopub.status.busy":"2026-09-12T11:19:24.77387Z","iopub.status.idle":"2026-09-12T11:19:29.016535Z","shell.execute_reply":"2026-09-12T11:19:29.015493Z"},"papermill":{"duration":4.260925,"end_time":"2026-09-12T11:19:29.018248+00:00","exception":false,"start_time":"2026-09-12T11:19:24.757323+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"AMP_PREF = 'bf16'\n\ndef amp_for(dev):\n    if not str(dev).startswith('cuda'):\n        return (torch.float32, False)\n    cc = torch.cuda.get_device_capability(dev)\n    if AMP_PREF == 'bf16':\n        return (torch.bfloat16, True)\n    if AMP_PREF == 'fp16':\n        return (torch.float16, True)\n    if AMP_PREF == 'fp32':\n        return (torch.float32, False)\n    return (torch.bfloat16 if cc >= (8, 0) else torch.float16, True)\nAMP_DT, AMP_ON = amp_for(DEV)\nWORKERS = max(1, min(4, os.cpu_count() or 4))\nCHUNK = 48\nMICRO = 8\nmodels = [m.to(DEV).eval() for m in models]\nprint(f\"device {DEV} | amp {str(AMP_DT).split('.')[-1]} (on={AMP_ON}) | workers {WORKERS} | chunk {CHUNK} | micro {MICRO}\")\n\ndef _norm_(im):\n    k = CFG.get('norm', 'none')\n    if k == 'zscore':\n        m = (im > 0).float()\n        n = m.sum(dim=(1, 2, 3), keepdim=True).clamp(min=1.0)\n        mu = (im * m).sum(dim=(1, 2, 3), keepdim=True) / n\n        var = (((im - mu) * m) ** 2).sum(dim=(1, 2, 3), keepdim=True) / n\n        return (im - mu) / (var.sqrt() + 1e-06) * m\n    if k == 'imagenet':\n        m = (im > 0).float()\n        return (im - 0.485) / 0.229 * m\n    return im\n\n@torch.no_grad()\ndef _micro(images, masks):\n    dev = DEV\n    ims, slots, sidx, vms = ([], [], [], [])\n    for b in range(len(masks)):\n        present = np.nonzero(masks[b] > 0)[0]\n        if len(present) == 0:\n            continue\n        blk = images[b][present]\n        ims.append(torch.from_numpy(blk))\n        vms.append(torch.from_numpy(blk.reshape(blk.shape[0], blk.shape[1], -1).max(2) > 0))\n        slots.append(torch.from_numpy(present + 1).long())\n        sidx.append(torch.full((len(present),), b, dtype=torch.long))\n    out = np.full((len(models), len(masks), len(LABELS)), np.nan, np.float32)\n    if not ims:\n        return out\n    im = _norm_(torch.cat(ims).to(dev, non_blocking=True).float().div_(255.0))\n    sl = torch.cat(slots).to(dev)\n    si = torch.cat(sidx).to(dev)\n    vm = torch.cat(vms).to(dev)\n    sm = torch.zeros(len(sl), CFG.get('n_meta', 0), device=dev)\n    per = torch.zeros(len(models), len(masks), len(LABELS), device=dev, dtype=torch.float32)\n    with torch.autocast('cuda' if str(dev).startswith('cuda') else 'cpu', dtype=AMP_DT, enabled=AMP_ON):\n        for fold_index, model in enumerate(models):\n            per[fold_index] = torch.sigmoid(\n                model(im, sl, sm, si, len(masks), vm=vm).float()\n            )\n    got = per.cpu().numpy()\n    keep = np.array([(masks[b] > 0).any() for b in range(len(masks))])\n    out[:, keep] = got[:, keep]\n    return out\n\ndef predict(images, masks):\n    out = np.full((len(models), len(masks), len(LABELS)), np.nan, np.float32)\n    for a in range(0, len(masks), MICRO):\n        b = min(a + MICRO, len(masks))\n        out[:, a:b] = _micro(images[a:b], masks[a:b])\n    return out\n\n# Macro ROC-AUC depends on ordering, so combine fold orderings rather\n# than allowing a fold's probability scale to dominate the mean.\npreds = np.full((len(models), len(studies), len(LABELS)), np.nan, np.float32)\nt0, done = (time.time(), 0)\nif not globals().get('_A5_ENC_OK', True):\n    _A5_TODO = []\n    print('DINOv3 inference SKIPPED: the fold checkpoints did not fill every '\n          'encoder tensor, so this branch would contribute a partly random '\n          'encoder at A5_W = 0.45. preds stay NaN, coverage comes out at 0% '\n          'and cell 40 keeps the DINOv2 parent.', flush=True)\nelse:\n    _A5_TODO = list(studies)\n_A5_CHUNKS = list(range(0, len(_A5_TODO), CHUNK))\n\n\ndef _a5_submit(ex, c0):\n    block = _A5_TODO[c0:c0 + CHUNK]\n    return (c0, block, [ex.submit(build_study, (i, s, by.get(s, [])))\n                        for i, s in enumerate(block)])\n\n\ndef _a5_collect(job):\n    c0, block, futs = job\n    imgs = np.zeros((len(block), N_SLOT, N_SLICE, SIZE, SIZE), np.uint8)\n    msks = np.zeros((len(block), N_SLOT), np.uint8)\n    for f in as_completed(futs):\n        try:\n            i, a, k = f.result()\n            imgs[i], msks[i] = (a, k)\n        except Exception as e:\n            print(f'  study failed: {type(e).__name__}: {e}')\n    return (c0, block, imgs, msks)\n\n\nwith ProcessPoolExecutor(max_workers=WORKERS) as ex:\n    _a5_inflight = _a5_submit(ex, _A5_CHUNKS[0]) if _A5_CHUNKS else None\n    for _a5_pos in range(len(_A5_CHUNKS)):\n        c0, block, imgs, msks = _a5_collect(_a5_inflight)\n        # The next chunk decodes on the pool while this one runs on the GPU.\n        _a5_inflight = (_a5_submit(ex, _A5_CHUNKS[_a5_pos + 1])\n                        if _a5_pos + 1 < len(_A5_CHUNKS) else None)\n        preds[:, c0:c0 + len(block)] = predict(imgs, msks)\n        done += len(block)\n        el = time.time() - t0\n        print(f'  {done:,}/{len(_A5_TODO):,}  {el / 60:.1f}m  eta {el / done * (len(_A5_TODO) - done) / 60:.1f}m', flush=True)\n        del imgs, msks\n        gc.collect()\nprint(f'\\ninference done in {(time.time() - t0) / 60:.1f} min')\nA5_W = 0.45\nA5_LABELS = list(LABELS)\n_a5_ok = np.isfinite(preds).all(axis=(0, 2))\n_a5_rank_mean = np.zeros((len(studies), len(LABELS)), np.float64)\nfor fold_index in range(preds.shape[0]):\n    fold = preds[fold_index][_a5_ok]\n    ordinal = fold.argsort(0).argsort(0).astype(np.float64)\n    _a5_rank_mean[_a5_ok] += ordinal / max(len(fold) - 1, 1)\n_a5_rank_mean /= preds.shape[0]\n_a5_rank_mean[~_a5_ok] = np.nan\nA5_PREDS = dict(zip(\n    sub_df['StudyInstanceUID'].astype(str), _a5_rank_mean.astype(np.float32)\n))\nfor _a5k, _a5v in _A5_SAVED.items():\n    globals()[_a5k] = _a5v\ndel _A5_SAVED, _a5k, _a5v"},{"cell_type":"markdown","id":"253c80af","metadata":{"papermill":{"duration":0.013514,"end_time":"2026-09-12T11:19:29.046122+00:00","exception":false,"start_time":"2026-09-12T11:19:29.032608+00:00","status":"completed"},"tags":[]},"source":"# 21. 🦖 DINOv3: смешивание с DINOv2\n\n## Что делает эта ячейка\n\nСмешивает предсказания DINOv3 с текущим `submission.csv` (DINOv2):\n\n- DINOv2: вес 0.55\n- DINOv3: вес 0.45\n\nРезультат — обновлённый `submission.csv`."},{"cell_type":"code","execution_count":null,"id":"caad6a1a","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:29.073776Z","iopub.status.busy":"2026-09-12T11:19:29.073237Z","iopub.status.idle":"2026-09-12T11:19:29.087044Z","shell.execute_reply":"2026-09-12T11:19:29.086172Z"},"papermill":{"duration":0.029538,"end_time":"2026-09-12T11:19:29.088794+00:00","exception":false,"start_time":"2026-09-12T11:19:29.059256+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"_a5_sub = pd.read_csv('/kaggle/working/submission.csv',\n                      dtype={'StudyInstanceUID': str})\nassert _a5_sub.columns.tolist()[1:] == A5_LABELS, 'submission schema drift'\n_A5_MIN_COVERAGE = 0.98\n\nif A5_W > 0:\n    _a5_nan = np.full(len(A5_LABELS), np.nan, np.float32)\n    _a5_ours = np.stack([A5_PREDS.get(_u, _a5_nan)\n                         for _u in _a5_sub['StudyInstanceUID'].astype(str)])\n    _a5_have = np.isfinite(_a5_ours).all(axis=1)\n    _a5_coverage = float(_a5_have.mean())\n    _a5_base_rank = _a5_sub[A5_LABELS].rank(method='average', pct=True)\n    if _a5_coverage < _A5_MIN_COVERAGE:\n        # Too little of the branch survived to rank against the parent: a rank\n        # column built from a small covered subset is not comparable with one\n        # built from every study.\n        print(f'[a5] DINOv3 blend skipped: only {_a5_coverage:.1%} of studies '\n              f'produced a prediction (need {_A5_MIN_COVERAGE:.0%}); '\n              f'submission.csv stays at the DINOv2 parent', flush=True)\n    else:\n        _a5_ours_rank = pd.DataFrame(_a5_ours, columns=A5_LABELS,\n                                     index=_a5_sub.index).rank(method='average',\n                                                               pct=True)\n        _a5_blend = ((1.0 - A5_W) * _a5_base_rank.to_numpy(np.float64)\n                     + A5_W * _a5_ours_rank.to_numpy(np.float64))\n        # Studies this branch could not read keep the parent's rank rather than\n        # poisoning the whole column with NaN.\n        _a5_blend[~_a5_have] = _a5_base_rank.to_numpy(np.float64)[~_a5_have]\n        if not np.isfinite(_a5_blend).all():\n            raise AssertionError('DINOv3 blend still non-finite after fallback')\n        _a5_sub[A5_LABELS] = _a5_blend\n        if int((~_a5_have).sum()):\n            print(f'[a5] {int((~_a5_have).sum())} study(ies) kept the DINOv2 '\n                  f'rank because DINOv3 produced no prediction for them',\n                  flush=True)\n        _a5_sub.to_csv('/kaggle/working/submission.csv', index=False)"},{"cell_type":"markdown","id":"d4348458","metadata":{"papermill":{"duration":0.013249,"end_time":"2026-09-12T11:19:29.115696+00:00","exception":false,"start_time":"2026-09-12T11:19:29.102447+00:00","status":"completed"},"tags":[]},"source":"# 22. 🩻 RadImageNet: инференс и калибровка\n\n## Что делает эта ячейка\n\nВыполняет третью ветку пайплайна:\n\n- Загружает ResNet50 RadImageNet и две группы голов (public и reference).\n- Строит кэш для трёх раскладок слотов (E10, E13, V48-pass2).\n- Смешивает предсказания с текущим submission.\n- Применяет калибровку для части диагнозов.\n\nОбновляет `submission.csv`."},{"cell_type":"code","execution_count":null,"id":"a3b644d4","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:29.143806Z","iopub.status.busy":"2026-09-12T11:19:29.143536Z","iopub.status.idle":"2026-09-12T11:19:43.704327Z","shell.execute_reply":"2026-09-12T11:19:43.703539Z"},"papermill":{"duration":14.577079,"end_time":"2026-09-12T11:19:43.705923+00:00","exception":false,"start_time":"2026-09-12T11:19:29.128844+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"from __future__ import annotations\nimport base64 as _rad_b64\nimport zlib as _rad_zlib\n_RAD_CAL_PAYLOAD = 'eNrtmk1vI8cRhv9KsJdcKKE/q6tzc4z4ZCMBcjQWhrCRDSG2ZEjaIEGQ/57n7RlRQ3KG4jqLJAcDS4o709NdXR9vvVU9/3z30+3N/bvffRuuawgxlm7eq2ePeffrpV8v/V9euvLraD3k6ilW6zn126vYd+U6eKmx9RiaF7OSx+X1weE67K7SdSgpVSauuaecUxr3rtp1LK0HyyV7y9Gmy/E6pBhTL61Zt2gWx+WtSew6JmOoM5AHutXp+ro8+ZrlsnvJOfDdfRL+Kl4zd2bqllNmga7rvrva2OzGNFuynGxpmnxrS/069hCtpt6T1dLLOQVsTbJ1fWunG2qXAX8NiU+7VK8rdqvNU89uwQwjpWyeLdTCnVy84ua9pthSjznHXLjDpZxbLe4WPRAQCT+LrSVcGGdLzNX0XKhesC2eV5uZ67lQPDlmSzUhS2+61ELtoTOyWGjdPudc73fvnj7c/Hg7ElpyzYXjdwIZ32m7X3INZ0crtvtc833ua6fyoUYUGf4H13rJpX+GK5aa5VbeWO9ye/03bLi97i8SpaSCl49nc4gdxI2p1dZyVmQnfL56jBH/j70GqSq6dyMROLHQGvCpcSnkYsyeohNErnVjbamVjs47wOpBa7BAsa61SXulDfSIwAbAkQqDmSK38WwImKK1kFJ3d11CpFqR1mPHlj1pWVAHeDfyU2hk0vGogz0GqCBbKmFMx9xIWEAbQ8I4PUvmAplQCeuij7GNXGMk3yhBsFIfEnvVE6HFYMHDtFuLzKwpya9gisaZduAkcyAvmw1ROjmd7OnBYytpOBqJjTyK4KzT0Pi0tQJYslc0nGqcNZV7dEaRvfcSewg9h4p41iwO+4CWMkNG1326VBmHFUoB7VoakzE+JAtYN/NknS9lLJ+9S5qR56Q2kLA4qgTuplGNmUJAkbXiGbrU0HCIsgtYaGOUVXbayNeykLfpGv8lulBfgiFEmx61xm3LTlrOPoa51Ng7IyJiT4tWxrE0ZmnJRnhawZsyLgjXIC9MRu0tRAgImkJ7bUyHKk1eUgIewtRjXGDJqO1mNJrHfNUty1HRW2SSoV3FAVIA//hrUXKANxhbRbEkgqAFKiCC43qOEXsPTaKtKCdkqx3/koMYsuP+mh4HrHoQtqSIw/h4DM7EpRL5zRirAcHiUDj2SUlr9fFAsElHBX/zWgJCQlw+83Qksw8Pt9+Ty0hmOBQBh9XQAr7fxjRQ2IkGTT/w8cAiBKwRc1jwcMh+WKtMgFShNaDGJWRAechNSOMUldXTynNIUPAaEKKkCv1cLFzxaowBOmX8vU8Pj1voWa5JTPFcGN4QXm+PpZPjA8QQYYp9Qab9rZiIAuLX8AA8f/jX4kHgI+BXCtBEeOTFvNPapIyC92YAEOmHypJcCSiFOpQi712M73IL8KiBX8Hz62pDNsAH8MK/rMR+tNSRIRJwAkIIY8Rn/fD2kB2V9I7BMX+KSJ/XJtNAbwFglmQZbu/90CZcMhAeZKudKAvlSBLwNwP1aA88BYpPdJRkbADGeBzWN2IOvySVELxyU0Lx0JHQHsFpEQRsCB9iO1qT3IOHGUjMDqeczX4BiIbvCjkiLD6t6G239p+gAML1KQuAdaLLdidDV6Y6uPR+9+3LT8KrlYgzAsjg7ok+VJakimvgAjn5qUFmtTGJKXGxSadg2UuLkWok4EvGJ0Hdsr+DmpVlGsUMZkxtuTI8vDAVaIsnSIDbK0DhVSpAFlu4iRtgrLYWRaArADOcFDAfErEdwBS8dWpNqHupW3o7nIukTmxlmASq4L/51ESrznqJTSkgCzslVSMncekb8659ljeyYttxK0oX9k50n3h2kDqoapToWt8S9/yO3tTXyteLu+19rkPViKkIblLnzJhJUhYENUJCtdfWKkEFe9WTsWin6azpqJFIfEwNyLNes+PZXFV3h+tIOXjb5mzIRnCLaWKuhnPtjsCVPAOwdkhQhBSMfWLVruTgwEM/WXt1FYgv5AN4IsyBiHrOSscCBjIomksgZCPFn9grqDrE9UXDQCSPfs5RXyeuQl3gwcEIdn5JzADhpG34s4uF5NQ24eh1GWISkEmYA6iFiNezYAbkKNGJhIk+l9XNXK3r7AgLT4YnRUuHicJS0NRLLG0KDrF1IbkaMzj2MoeKzAH/kTo9KsnWJe8gDZUuxs1iPi0CayABMdLIT4qQCbfENfCiAsujIsj9KMUQ1kXwXUBFZspHZm9Zj/dB/SrVy1qOJg9ApKAlmSQNpr6SjibnNgVoqZQCr8gllgsTiFAOCFSeDcYWuKOChZQXi5dPxouF5Oo6ogVTaxDwnVE8KfQFI6Zomwgb/oI4VG2Ae1MdsTQCwQuZp5ZROZPcdudWHrZRfwe4cPE1cv9L/FDuwRwGNYOttGllMRItS1K3sFDQDAwQMvgpfIzyIeQj2yTFOhomGLsBVPtAASGLqiXKnChASW/4ILUTGlftQlUVl8FDjifbKYIzLKXmco4anNPK6ZVl8MhVBs+D5YgTo8HXuIEDQDJF4wHEYL5PvEExn4PIt7x0JumDeUgjxHeFIZDt+6xN5mALCc6pMjTZ7oy8EomCUA09MTRvvuCLxyFC+sEWCGiqjE+HIAJ4g39Bi5sadSN9MJZaFWLbpubBNBvlmwrPAONJWS1DpaIgbyKuZSbKg7qZzgNKsqnwIX8RAOvINdgfRg0jPNJBWE+5xAQ9qDaAPh4nKInagjAL9otj26U+sAnbUERTF0QYjHV8j8OEA+mohtH9SLLuoUIh6uAWHI6yAw54uEsqL6zufMAAt3yW+6xSFLmKwEmYgIPlXHYbPAo3AvKovCk/0KR5m3fmKkpFz2W3ng58h7KzqVVHqVYUpfGlAJHvQ2dN7QKmXoQITggHUz9HGZhSty5smYHeSL4pFdvbXlXZcSJV7GCHKPpsePEmVAOm41WxbuW3Y+qkrhYhisbrSBqbGiGdAkBwd5gzNehxGUVhzu5JCKa2VAyL2pfYgYgUtcTwu3RKZzO2JublhfCDbO1Qqypu8SdTWO0501Z6iAAH68g21mECvjX8bZZ7yvpViqq9BlHCnhj7DFuiXm8qEwgrAkK9hpfbuU/nN/i4l7I/qQEnpR4qVHUxgDjblWuM2UZFhOpzp+apb8i4ua8gnyHrqPMHs839gK7gaaYDDtILqNTWOF8kxVYdMxlVDwE68/F8DR2CgpCSkkzkEvKY3x9GoRqjpn8ZLj62TyEoRVV1dxrV2GtNOMosPAjqSC6pm6yRKRJFM5wnixbhh6uLl9FdjElsUvX7WxR61T1gGohOkQaej27MxiR+DQjAndQ5SgL4Yb+VMKT2UncGK4se5hfeR7Jr6nygbTJcPGK7QSpXOoBJkuMXlFTYQX0aRsObYEQn22BNrDQdKwmZ1DDaZNcANxumdkGs3pZIJcaCcxBuPSlGziTf+UNZIgqmLMTWS1/XYNBxFowH5FchZb7uT5EQqervUK+Rc15pP55mBpKra6Wju5MCbZwiKOSS/HFuZyVRBFeeIoVrc35oNnVt1ARpKDaHF2BO1z3DLILOrOVZdhIG52ojyFxsVa1k6Bq+sOS7AA0upIqpiCt9Om8+1SsIoI4/ZYFKZq9tdwnhCzrVRpoCxpqSo+0ua1DNLCOMc3X19vGSZZ9znfEAMTUpymSGqYW/kpXU6FWV6+pah9OGLveaTkopOV1d3U9oWfAhMiyK46fRgPeL9LTSDAtUjj5cFN4D8S5zmaCjmaTjJ/Kvxb3MaFrHVI27QS2vRfcMCmeqcUXY+NX6652okyj19OEC3SaFrdWyeyJMxo7gPrANn17GjQ5RgtqvregNjzr1hYOOxATeos4b/IItGRxEFZC4aNruVL1ZbGx8ojpE5moLiGNavazB+baxPoXr/gK56zjI1VGbjtG8nA3Xg3SpFl2gtqzqsM/oGgb5Axj4w+XD1g48MBFrMAnR+TRxMci1+8h+uCC1e1nmC7XEWhf2zEdIw5SsCe3Tcc18XibsU/PKhJhd7a380s67RLEvQQWUw9KKjmps7qefVeub/IZQoLbK0qxnHRFP7qlOy8BURC7Fy2LDPk7N1F1VEm1lrZbSmSWVSoNk63DQXgvIohAjYcjXU7b/zI8u7RVvPiRChSqVjkiU6YQjrT1ImVAAN/WoKEp62V2aQGAdAuQKhSW5qouyIdQ4jNe5NS6I8v10hLIeqIIMqplDn71o0dYl4fEFkRZ4LmhFJNUSQqZCNlBT2fIWnKzJ9XWwG48bOzGZWEIffcVx/q23xAziBV83gf1cE4/+mvI8/EFV3QVmPW0a6rB7ZD314ploTllOSlFYqX3X4u7TJg6jmxLRWxit1FaPtArKoHfHeb3jwPmNGDqfvS+Iv9Ns14Vv6p9rfyft+HGiFqmJIQSU++rlbegPBqDumli9zoOGSptec8O6GAWqtaAgkOgqZ3P1WhQQ08aj2KGONuEIqdZetzuLFB6qQniid8HkZTnlyGsPGnA4lY5Qu1456WnbokG9IupwU0NoPhHTwRRUqynZU0HObV8QUyZuo/lb6tFhsdxb7bCoU9qWe8vLxBwj2laVzgS+bIZGMfee1Qaldl87gBZ5jjr3Y4rq2Xaf3LatohHwxzbqBKtv4d8bHnh1eUqfVRx07jc4zfzySrj8IGV9NMFX4mBKrq6FXc4Jz154/3737u7++fbxw+3Pz9N7eg0X0mHCeHfHbHrNhrCIpdXRA/LRRC4WR9sU/ZqJB46XJsJouwU1rPT+il6uKHqNAdbJ2InUJp0wVWBV7w8NuldU7uMyajZqh2N6UfeoM6ym1zLGizF6T6AnNQZiHO/KZMijZMOV2/yKQFKv0TSXq0Hb5teE9F4HLp50QKJ3OX64edZ7ie+++PLrd7t339z+5e7mx9/88Qt+f82dx5f//Omr6e8fvv/+49Pdwz0/f3/z19vH3z7x68uH++fpKhP+/Pjw/PDh4cfv+Hz86f5Jk99/93T7eHersfff/fnmh/H3y4fH8feLv9/x96ub5+l7vq9f0wj9msf8+HH6fhnDr3kMvzRGG3p8+PizVt3vaXy/ysjox5sPzx8fbxn+7cuWv7m9v3v68PFpsfH9pcWwLc1oyEI3f/7H/cPf7p7vnhZ6ev/+X/8GYIe3xg=='\n\nimport contextlib as _rad_contextlib\nimport gc as _rad_gc\nimport hashlib as _rad_hashlib\nimport json as _rad_json\nimport os as _rad_os\nimport re as _rad_re\nimport time as _rad_time\nfrom concurrent.futures import ThreadPoolExecutor as _RadThreadPool\nfrom pathlib import Path as _RadPath\n\nimport numpy as _rad_np\nimport pandas as _rad_pd\nimport pydicom as _rad_pydicom\nimport torch as _rad_torch\nimport torch.nn as _rad_nn\nimport torch.nn.functional as _rad_F\nfrom torchvision.models import resnet50 as _rad_resnet50\n\n_RAD_LABELS = [\n    'ACL', 'MCL', 'Medial Meniscus', 'Lateral Meniscus', 'Medial OA',\n    'Lateral OA', 'PF OA', 'Effusion', 'Synovitis', \"Baker's\",\n    'Contusion', 'Fracture',\n]\n_RAD_ALPHA = 0.50\n_RAD_EXCLUDE = (\"Baker's\", 'Fracture')\n_RAD_HEADS_SHA256 = '54f657826b3458a7ba3d462e198ba380732f2b136246182312704929874a9a2c'\n_RAD_REFERENCE_HEADS_SHA256 = '0f465649799ecfbccaac1767844639e7ced44e1bc9babde6e4bac7c5d9b89eaa'\n_RAD_ENCODER_SHA256 = '08629f7e7bd3e29b8ee9522ca3f65ce4d010a7ddf74f0ea3c7e3f3d0bbab0734'\n_RAD_E13_HEADS_SHA256 = 'ad9f19af73bfdf4e49263c0e45060dc3cb239e1195039b26dc8c0a3a6bcd1a8a'\n_RAD_E13_MEMBER_WEIGHT = 0.50\n_RAD_V48_SECOND_ALPHA = 0.15\n_RAD_TWIN_ALT_WEIGHT = 0.500001\n_RAD_TOKEN_DIM, _RAD_HEAD_DIM = 2048, 512\n\n_RAD_E11_SLOTS = [\n    ('SAG_NOFS', 'Sagittal', None, False),\n    ('COR_NOFS', 'Coronal', None, False),\n    ('AX_NOFS', 'Axial', None, False),\n    ('SAG_FS', 'Sagittal', None, True),\n]\n_RAD_E11_CROP_MM = 130.0\n_RAD_E11_CACHE_SLICES = 8\n_RAD_E11_IMG = 224\n\n_RAD_E13_SLOTS = [\n    ('SAG_FS', 'Sagittal', None, True),\n    ('COR_FS', 'Coronal', None, True),\n    ('AX_FS', 'Axial', None, True),\n    ('SAG_NOFS', 'Sagittal', None, False),\n]\n_RAD_E13_CROP_MM = 130.0\n_RAD_E13_CACHE_SLICES = 8\n_RAD_E13_IMG = 224\n\n# Our independently trained five-fold family.  Its preprocessing and estimator\n# are preserved from V35: native DICOM geometry/fat-sat handling and a mean of\n# per-fold percentile ranks (rather than v15's rank of the probability mean).\n_OUR_N_SLOT, _OUR_N_SLICE, _OUR_IMG = 3, 8, 224\n\n# Exact V40/E10 test representation: three fat-suppressed planes, eight\n# acquired slices per plane, full frame, legacy ordering/laterality/fill.\nSLOTS = [\n    ('SAG_FS', 'Sagittal', None, True),\n    ('COR_FS', 'Coronal', None, True),\n    ('AX_FS', 'Axial', None, True),\n]\nN_SLOT = len(SLOTS)\nCACHE_SLICES = 8\nIMG = CACHE_IMG = 224\nCROP_MM = 10_000.0\nSLICE_BAND = (0.2, 0.8)\nRULES = dict(RULES_LEGACY)\nTIME_BUDGET = 8.0 * 3600\n\n\ndef _rad_log(message):\n    print(f'[Rad-dual5] {message}', flush=True)\n\n\ndef _rad_sha256(path, chunk=8 << 20):\n    digest = _rad_hashlib.sha256()\n    with open(path, 'rb') as handle:\n        for block in iter(lambda: handle.read(chunk), b''):\n            digest.update(block)\n    return digest.hexdigest()\n\n\ndef _rad_find_file(name, expected_sha=None, explicit_env=None):\n    if explicit_env and _rad_os.environ.get(explicit_env):\n        candidates = [_RadPath(_rad_os.environ[explicit_env])]\n    else:\n        candidates = []\n        base = _RadPath('/kaggle/input')\n        if base.is_dir():\n            for root, dirs, files in _rad_os.walk(base):\n                dirs[:] = [d for d in dirs if d not in ('train_series', 'test_series')]\n                if name in files:\n                    candidates.append(_RadPath(root) / name)\n    if not candidates:\n        raise FileNotFoundError(f'V36 missing input artifact {name}')\n    for path in candidates:\n        if expected_sha is None or _rad_sha256(path) == expected_sha:\n            return path\n    raise RuntimeError(f'V36 found {name}, but no copy has the required SHA-256')\n\n\nclass _RadEncoder(_rad_nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.backbone = _rad_nn.Sequential(\n            *list(_rad_resnet50(weights=None).children())[:-2]\n        )\n\n    def forward(self, image):\n        return self.backbone(image).mean(dim=(2, 3))\n\n\nclass _RadHead(_rad_nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.project = _rad_nn.Sequential(\n            _rad_nn.LayerNorm(_RAD_TOKEN_DIM),\n            _rad_nn.Linear(_RAD_TOKEN_DIM, _RAD_HEAD_DIM),\n            _rad_nn.GELU(),\n        )\n        self.plane = _rad_nn.Parameter(_rad_torch.randn(N_SLOT, _RAD_HEAD_DIM) * .01)\n        self.position = _rad_nn.Parameter(_rad_torch.randn(CACHE_SLICES, _RAD_HEAD_DIM) * .01)\n        self.query = _rad_nn.Parameter(_rad_torch.randn(len(_RAD_LABELS), _RAD_HEAD_DIM) * .02)\n        self.attn = _rad_nn.MultiheadAttention(\n            _RAD_HEAD_DIM, 8, dropout=.10, batch_first=True\n        )\n        self.fuse = _rad_nn.Sequential(\n            _rad_nn.LayerNorm(_RAD_HEAD_DIM * 4),\n            _rad_nn.Linear(_RAD_HEAD_DIM * 4, _RAD_HEAD_DIM),\n            _rad_nn.GELU(),\n            _rad_nn.Dropout(.15),\n        )\n        self.weight = _rad_nn.Parameter(\n            _rad_torch.randn(len(_RAD_LABELS), _RAD_HEAD_DIM) * .02\n        )\n        self.bias = _rad_nn.Parameter(_rad_torch.zeros(len(_RAD_LABELS)))\n\n    def forward(self, feature, mask):\n        token = self.project(feature.float())\n        token = token.view(len(token), N_SLOT, CACHE_SLICES, _RAD_HEAD_DIM)\n        token = token + self.plane[None, :, None] + self.position[None, None]\n        token = token.flatten(1, 2)\n        key_padding = mask <= 0\n        all_empty = key_padding.all(1)\n        if all_empty.any():\n            key_padding = key_padding.clone()\n            key_padding[all_empty, 0] = False\n        query = self.query.unsqueeze(0).expand(len(token), -1, -1)\n        attended = query + self.attn(\n            query, token, token, key_padding_mask=key_padding, need_weights=False\n        )[0]\n        denominator = mask.sum(1, keepdim=True).clamp_min(1).unsqueeze(-1)\n        mean = (token * mask.unsqueeze(-1)).sum(1, keepdim=True) / denominator\n        mean = mean.expand(-1, len(_RAD_LABELS), -1)\n        fused = self.fuse(_rad_torch.cat(\n            [attended, mean, _rad_torch.abs(attended - mean), attended * mean], dim=-1\n        ))\n        return (fused * self.weight.unsqueeze(0)).sum(-1) + self.bias\n\n\ndef _rad_load_public_heads(device, expected_sha):\n    heads_path = _rad_find_file('v52_radimagenet_heads.pt', expected_sha)\n    payload = _rad_torch.load(heads_path, map_location='cpu', weights_only=True)\n    expected = {\n        'version': 'v52-radimagenet-resnet50-official-1',\n        'targets': _RAD_LABELS,\n        'encoder_sha256': _RAD_ENCODER_SHA256,\n        'encoder_source_commit': '0ce16f7375db4236e646829d1eca61cdb4282133',\n        'img': 224,\n        'slices_per_plane': 8,\n        'feature': 'global_average_pool',\n    }\n    for key, value in expected.items():\n        if payload.get(key) != value:\n            raise RuntimeError(f'public-v15 head contract drift for {key}')\n    folds = payload.get('folds')\n    if not isinstance(folds, list) or len(folds) != 5:\n        raise RuntimeError('public-v15 bundle requires exactly five heads')\n    if sorted(int(record.get('fold', -1)) for record in folds) != list(range(5)):\n        raise RuntimeError('public-v15 fold identity drift')\n    heads = []\n    for record in folds:\n        head = _RadHead().to(device).eval()\n        head.load_state_dict(record['state_dict'], strict=True)\n        heads.append(head)\n    return heads, str(heads_path)\n\n\ndef _rad_load_e13_heads(device):\n    # V48 used an unqualified filename shared by E11 and E13. Resolve the\n    # intended E13 bundle by content and validate its complete pixel contract.\n    heads_path = _rad_find_file('v52_e11_heads.pt', _RAD_E13_HEADS_SHA256)\n    payload = _rad_torch.load(heads_path, map_location='cpu', weights_only=False)\n    expected = {\n        'version': 'e11-radimagenet-resnet50-diverse-1',\n        'targets': _RAD_LABELS,\n        'encoder_sha256': _RAD_ENCODER_SHA256,\n        'slots': [list(slot) for slot in _RAD_E13_SLOTS],\n        'crop_mm': _RAD_E13_CROP_MM,\n        'img': _RAD_E13_IMG,\n        'slices_per_plane': _RAD_E13_CACHE_SLICES,\n        'feature': 'global_average_pool',\n    }\n    for key, value in expected.items():\n        if payload.get(key) != value:\n            raise RuntimeError(f'E13 head contract drift for {key}')\n    folds = payload.get('folds')\n    if not isinstance(folds, list) or len(folds) != 5:\n        raise RuntimeError('E13 bundle requires exactly five heads')\n    if sorted(int(record.get('fold', -1)) for record in folds) != list(range(5)):\n        raise RuntimeError('E13 fold identity drift')\n    heads = []\n    for record in folds:\n        head = _RadHead().to(device).eval()\n        head.load_state_dict(record['state_dict'], strict=True)\n        heads.append(head)\n    return heads, str(heads_path)\n\n\ndef _rad_load_models(device):\n    encoder_path = _rad_find_file(\n        'ResNet50.pt', _RAD_ENCODER_SHA256, explicit_env='RSNA_RAD_WEIGHT_PATH'\n    )\n    encoder = _RadEncoder()\n    encoder.load_state_dict(\n        _rad_torch.load(encoder_path, map_location='cpu', weights_only=True), strict=True\n    )\n    if sum(parameter.numel() for parameter in encoder.parameters()) != 23_508_032:\n        raise RuntimeError('V36 RadImageNet encoder parameter-count drift')\n    encoder.eval().to(device)\n    for parameter in encoder.parameters():\n        parameter.requires_grad_(False)\n    if device.type == 'cuda' and _rad_torch.cuda.device_count() > 1:\n        encoder = _rad_nn.DataParallel(\n            encoder, device_ids=list(range(_rad_torch.cuda.device_count()))\n        )\n\n    alt_heads, alt_path = _rad_load_public_heads(device, _RAD_HEADS_SHA256)\n    reference_heads, reference_path = _rad_load_public_heads(\n        device, _RAD_REFERENCE_HEADS_SHA256\n    )\n    if alt_path == reference_path:\n        raise RuntimeError('twin E10 branches resolved to the same artifact')\n    return encoder, alt_heads, reference_heads, str(encoder_path), alt_path, reference_path\n\n\n@_rad_torch.inference_mode()\ndef _rad_encode(encoder, pixels, slot_mask, device):\n    n, slots, slices, height, width = pixels.shape\n    features = _rad_np.zeros(\n        (n, slots * slices, _RAD_TOKEN_DIM), _rad_np.float16\n    )\n    token_mask = _rad_np.repeat(slot_mask[:, :, None], slices, axis=2).reshape(n, -1)\n    valid = _rad_np.flatnonzero(token_mask.reshape(-1) > 0)\n    flat = pixels.reshape(-1, height, width)\n    batch = 192 if device.type == 'cuda' and _rad_torch.cuda.device_count() > 1 else (\n        96 if device.type == 'cuda' else 8\n    )\n    for start in range(0, len(valid), batch):\n        indices = valid[start:start + batch]\n        image = _rad_torch.from_numpy(flat[indices]).to(device).float().div_(127.5).sub_(1.0)\n        image = image.unsqueeze(1).expand(-1, 3, -1, -1).contiguous()\n        amp = (_rad_torch.autocast('cuda')\n               if device.type == 'cuda' else _rad_contextlib.nullcontext())\n        with amp:\n            feature = encoder(image)\n        values = feature.float().cpu().numpy()\n        if not _rad_np.isfinite(values).all():\n            raise RuntimeError('V36 non-finite RadImageNet feature')\n        features.reshape(-1, _RAD_TOKEN_DIM)[indices] = values.astype(_rad_np.float16)\n    return features, token_mask.astype(_rad_np.float32)\n\n\n@_rad_torch.inference_mode()\ndef _rad_predict_head(head, features, masks, device, batch=64):\n    predictions = []\n    for start in range(0, len(features), batch):\n        image = _rad_torch.from_numpy(features[start:start + batch]).to(device)\n        mask = _rad_torch.from_numpy(masks[start:start + batch]).to(device)\n        amp = (_rad_torch.autocast('cuda')\n               if device.type == 'cuda' else _rad_contextlib.nullcontext())\n        with amp:\n            predictions.append(_rad_torch.sigmoid(head(image, mask)).float().cpu())\n    return _rad_torch.cat(predictions).numpy()\n\n\ndef _rad_rank_columns(values):\n    return _rad_pd.DataFrame(\n        _rad_np.asarray(values, dtype=_rad_np.float64)\n    ).rank(method='average', pct=True).to_numpy(_rad_np.float64)\n\n\ndef _rad_validate(frame, expected_ids):\n    if frame.columns.tolist() != ['StudyInstanceUID', *_RAD_LABELS]:\n        raise RuntimeError('V36 submission schema drift')\n    ids = frame['StudyInstanceUID'].astype(str).tolist()\n    if ids != list(map(str, expected_ids)) or len(ids) != len(set(ids)):\n        raise RuntimeError('V36 submission study identity/order drift')\n    values = frame[_RAD_LABELS].to_numpy(_rad_np.float64)\n    if not _rad_np.isfinite(values).all() or values.min() < 0 or values.max() > 1:\n        raise RuntimeError('V36 invalid submission values')\n\n\ndef _rad_main():\n    started = _rad_time.time()\n    work = _RadPath(_rad_os.environ.get('RSNA_RAD_OUTPUT_DIR', '/kaggle/working'))\n    primary = work / 'submission.csv'\n    if not primary.is_file():\n        raise FileNotFoundError('V37 requires the completed DINO parent submission.csv')\n    test = _rad_pd.read_csv(ROOT / 'test.csv', dtype={'StudyInstanceUID': str})\n    expected_ids = test['StudyInstanceUID'].astype(str).tolist()\n    baseline = _rad_pd.read_csv(primary, dtype={'StudyInstanceUID': str})\n    _rad_validate(baseline, expected_ids)\n\n    device = _rad_torch.device('cuda:0' if _rad_torch.cuda.is_available() else 'cpu')\n    if device.type != 'cuda':\n        raise RuntimeError('V37 RadImageNet inference requires CUDA')\n    (encoder, public_heads, reference_heads, encoder_path,\n     public_heads_path, reference_heads_path) = _rad_load_models(device)\n\n    # Family 1: public v15/E10 legacy pixels.  Keep this path bit-for-bit as in\n    # V36, including rank(mean(fold probability)).\n    test_series = _rad_pd.read_csv(\n        ROOT / 'test_series.csv',\n        dtype={'StudyInstanceUID': str, 'SeriesInstanceUID': str},\n    )\n    plane = dict(zip(test_series.SeriesInstanceUID, test_series.Anatomical_Plane))\n    headers = annotate(walk('test_series'))\n    studies, pixels, slot_mask = build_cache(\n        pick_slots(headers, plane), plane, lat_of(headers, 'test-e10 '), 'test-e10'\n    )\n    by_uid = {str(uid): index for index, uid in enumerate(studies)}\n    missing = [uid for uid in expected_ids if uid not in by_uid]\n    if missing:\n        raise RuntimeError(f'{len(missing)} test studies absent from public-v15 cache')\n    order = _rad_np.asarray([by_uid[uid] for uid in expected_ids], dtype=_rad_np.int64)\n    pixels, slot_mask = pixels[order], slot_mask[order]\n    token_count = int(\n        _rad_np.repeat(slot_mask[:, :, None], CACHE_SLICES, axis=2).sum()\n    )\n    if token_count < int(0.85 * len(test) * N_SLOT * CACHE_SLICES):\n        raise RuntimeError(f'insufficient acquired public-v15 test slices: {token_count}')\n\n    features, token_mask = _rad_encode(encoder, pixels, slot_mask, device)\n    del pixels, slot_mask, headers\n    _rad_gc.collect()\n    public_fold_predictions = [\n        _rad_predict_head(head, features, token_mask, device)\n        for head in public_heads\n    ]\n    reference_fold_predictions = [\n        _rad_predict_head(head, features, token_mask, device)\n        for head in reference_heads\n    ]\n    if len(public_fold_predictions) != 5 or len(reference_fold_predictions) != 5:\n        raise RuntimeError('twin E10 inference did not use all ten heads')\n\n    # Preserve each public recipe's rank(mean(fold probability)) estimator.\n    public_probability = _rad_np.mean(_rad_np.stack(public_fold_predictions), axis=0)\n    reference_probability = _rad_np.mean(\n        _rad_np.stack(reference_fold_predictions), axis=0\n    )\n    public_rank = _rad_rank_columns(public_probability)\n    reference_rank = _rad_rank_columns(reference_probability)\n    del (public_heads, reference_heads, public_fold_predictions,\n         reference_fold_predictions, public_probability, reference_probability,\n         features, token_mask)\n    _rad_gc.collect()\n    _rad_torch.cuda.empty_cache()\n    _rad_log(\n        f'twin public-v15 families complete ({public_heads_path}; {reference_heads_path})'\n    )\n\n    # One new member from V48: three fat-sensitive planes plus a sagittal\n    # structural anchor, all at a 130 mm crop. Average ranks inside the Rad block\n    # and re-rank the result exactly as V48 does before the unchanged E10 vote.\n    globals().update(\n        SLOTS=list(_RAD_E13_SLOTS),\n        N_SLOT=len(_RAD_E13_SLOTS),\n        CACHE_SLICES=int(_RAD_E13_CACHE_SLICES),\n        IMG=int(_RAD_E13_IMG),\n        CACHE_IMG=int(_RAD_E13_IMG),\n        CROP_MM=float(_RAD_E13_CROP_MM),\n        RULES=dict(RULES_LEGACY),\n    )\n    e13_heads, e13_path = _rad_load_e13_heads(device)\n    headers = annotate(walk('test_series'))\n    studies, pixels, slot_mask = build_cache(\n        pick_slots(headers, plane), plane, lat_of(headers, 'test-e13 '), 'test-e13'\n    )\n    by_uid = {str(uid): index for index, uid in enumerate(studies)}\n    missing = [uid for uid in expected_ids if uid not in by_uid]\n    if missing:\n        raise RuntimeError(f'{len(missing)} test studies absent from E13 cache')\n    order = _rad_np.asarray([by_uid[uid] for uid in expected_ids], dtype=_rad_np.int64)\n    pixels, slot_mask = pixels[order], slot_mask[order]\n    e13_token_count = int(\n        _rad_np.repeat(slot_mask[:, :, None], CACHE_SLICES, axis=2).sum()\n    )\n    if e13_token_count < int(0.85 * len(test) * N_SLOT * CACHE_SLICES):\n        raise RuntimeError(f'insufficient acquired E13 test slices: {e13_token_count}')\n    e13_features, e13_token_mask = _rad_encode(\n        encoder, pixels, slot_mask, device\n    )\n    del pixels, slot_mask, headers\n    _rad_gc.collect()\n    e13_predictions = [\n        _rad_predict_head(head, e13_features, e13_token_mask, device)\n        for head in e13_heads\n    ]\n    if len(e13_predictions) != 5:\n        raise RuntimeError('E13 inference did not use all five heads')\n    e13_probability = _rad_np.mean(_rad_np.stack(e13_predictions), axis=0)\n    if (\n        e13_probability.shape != (len(test), len(_RAD_LABELS))\n        or not _rad_np.isfinite(e13_probability).all()\n    ):\n        raise RuntimeError(f'invalid E13 prediction shape/value: {e13_probability.shape}')\n    e13_rank = _rad_rank_columns(e13_probability)\n    public_rank = _rad_rank_columns(\n        (1.0 - _RAD_E13_MEMBER_WEIGHT) * public_rank\n        + _RAD_E13_MEMBER_WEIGHT * e13_rank\n    )\n    reference_rank = _rad_rank_columns(\n        (1.0 - _RAD_E13_MEMBER_WEIGHT) * reference_rank\n        + _RAD_E13_MEMBER_WEIGHT * e13_rank\n    )\n    # V48 resolves this same bundle again after switching pixel layouts.\n    del (e13_predictions, e13_probability, e13_rank,\n         e13_features, e13_token_mask)\n    _rad_gc.collect()\n    _rad_torch.cuda.empty_cache()\n    _rad_log(\n        f'E13 FS-crop member complete at Rad-block weight '\n        f'{_RAD_E13_MEMBER_WEIGHT:.2f} ({e13_path})'\n    )\n\n    # E10 keeps its audited 0.50 parent/Rad vote. The two excluded findings\n    # remain the raw parent values, matching the audited deployment.\n    baseline_rank = _rad_rank_columns(baseline[_RAD_LABELS].to_numpy())\n\n    def _rad_e10_branch(head_rank):\n        branch = baseline.copy()\n        for index, target in enumerate(_RAD_LABELS):\n            if target not in _RAD_EXCLUDE:\n                branch[target] = (\n                    (1.0 - _RAD_ALPHA) * baseline_rank[:, index]\n                    + _RAD_ALPHA * head_rank[:, index]\n                )\n        return branch\n\n    candidate_alt = _rad_e10_branch(public_rank)\n    candidate_reference = _rad_e10_branch(reference_rank)\n    for branch in (candidate_alt, candidate_reference):\n        for target in _RAD_EXCLUDE:\n            if not _rad_np.array_equal(\n                branch[target].to_numpy(), baseline[target].to_numpy()\n            ):\n                raise RuntimeError(f'E10 failed to preserve raw parent values for {target}')\n        _rad_validate(branch, expected_ids)\n    # Diagnostic E10 output only; the final equal rank mean is formed after E11\n    # and the legacy-DINO tie-break have completed independently in each branch.\n    candidate = baseline.copy()\n    alt_e10_rank = _rad_rank_columns(candidate_alt[_RAD_LABELS].to_numpy())\n    reference_e10_rank = _rad_rank_columns(\n        candidate_reference[_RAD_LABELS].to_numpy()\n    )\n    candidate[_RAD_LABELS] = (\n        _RAD_TWIN_ALT_WEIGHT * alt_e10_rank\n        + (1.0 - _RAD_TWIN_ALT_WEIGHT) * reference_e10_rank\n    )\n    _rad_validate(candidate, expected_ids)\n    e10_path = work / 'submission_e10_v2.csv'\n    candidate.to_csv(e10_path, index=False)\n    _rad_log(\n        f'twin E10 branches complete at alpha={_RAD_ALPHA:.2f}; '\n        f'preserved raw={list(_RAD_EXCLUDE)}'\n    )\n\n    # V48's successful run selected the E13 bundle a second time after\n    # installing the older E11 slot order. Express that observed behavior\n    # directly, without relying on duplicate-filename directory order.\n    globals().update(\n        SLOTS=list(_RAD_E11_SLOTS),\n        N_SLOT=len(_RAD_E11_SLOTS),\n        CACHE_SLICES=int(_RAD_E11_CACHE_SLICES),\n        IMG=int(_RAD_E11_IMG),\n        CACHE_IMG=int(_RAD_E11_IMG),\n        CROP_MM=float(_RAD_E11_CROP_MM),\n        RULES=dict(RULES_LEGACY),\n    )\n    headers = annotate(walk('test_series'))\n    studies, pixels, slot_mask = build_cache(\n        pick_slots(headers, plane), plane,\n        lat_of(headers, 'test-v48-pass2 '), 'test-v48-pass2'\n    )\n    by_uid = {str(uid): index for index, uid in enumerate(studies)}\n    missing = [uid for uid in expected_ids if uid not in by_uid]\n    if missing:\n        raise RuntimeError(f'{len(missing)} test studies absent from V48 pass-2 cache')\n    order = _rad_np.asarray([by_uid[uid] for uid in expected_ids], dtype=_rad_np.int64)\n    pixels, slot_mask = pixels[order], slot_mask[order]\n    v48_pass2_token_count = int(\n        _rad_np.repeat(slot_mask[:, :, None], CACHE_SLICES, axis=2).sum()\n    )\n    if v48_pass2_token_count < int(0.55 * len(test) * N_SLOT * CACHE_SLICES):\n        raise RuntimeError(\n            f'insufficient acquired V48 pass-2 slices: {v48_pass2_token_count}'\n        )\n    print(f'[FLIP-DEBUG] BEFORE original encode: pixels.shape={pixels.shape}', flush=True)\n    _t0 = _rad_time.time()\n    v48_features, v48_token_mask = _rad_encode(\n        encoder, pixels, slot_mask, device\n    )\n    print(f'[FLIP-DEBUG] AFTER original encode: {_rad_time.time()-_t0:.1f}s, v48_features.shape={v48_features.shape}', flush=True)\n    _t1 = _rad_time.time()\n    pixels_flip = pixels[:, :, :, :, ::-1].copy()\n    print(f'[FLIP-DEBUG] pixels_flip.shape={pixels_flip.shape}', flush=True)\n    v48_features_flip, v48_token_mask_flip = _rad_encode(\n        encoder, pixels_flip, slot_mask, device\n    )\n    print(f'[FLIP-DEBUG] AFTER flip encode: {_rad_time.time()-_t1:.1f}s, v48_features_flip.shape={v48_features_flip.shape}', flush=True)\n    del pixels_flip\n    del pixels, slot_mask, headers\n    _rad_gc.collect()\n    v48_pass2_predictions = [\n        _rad_predict_head(head, v48_features, v48_token_mask, device)\n        for head in e13_heads\n    ]\n    if len(v48_pass2_predictions) != 5:\n        raise RuntimeError('V48 second pass did not use all five E13 heads')\n    v48_pass2_probability = _rad_np.mean(\n        _rad_np.stack(v48_pass2_predictions), axis=0\n    )\n    if (\n        v48_pass2_probability.shape != (len(test), len(_RAD_LABELS))\n        or not _rad_np.isfinite(v48_pass2_probability).all()\n    ):\n        raise RuntimeError(\n            f'invalid V48 pass-2 prediction: {v48_pass2_probability.shape}'\n        )\n    v48_pass2_rank = _rad_rank_columns(v48_pass2_probability)\n    v48_pass2_predictions_flip = [\n        _rad_predict_head(head, v48_features_flip, v48_token_mask_flip, device)\n        for head in e13_heads\n    ]\n    v48_pass2_probability_flip = _rad_np.mean(\n        _rad_np.stack(v48_pass2_predictions_flip), axis=0\n    )\n    v48_pass2_rank_flip = _rad_rank_columns(v48_pass2_probability_flip)\n\n    _RAD_MIRROR_PAIRS = (\n        ('Medial Meniscus', 'Lateral Meniscus'),\n        ('Medial OA', 'Lateral OA'),\n    )\n    _RAD_MIRROR_UNPAIRED = ('MCL', \"Baker's\")\n    _rad_flip_aligned = v48_pass2_rank_flip.copy()\n    for _rad_a, _rad_b in _RAD_MIRROR_PAIRS:\n        _rad_i = _RAD_LABELS.index(_rad_a)\n        _rad_j = _RAD_LABELS.index(_rad_b)\n        _rad_flip_aligned[:, [_rad_i, _rad_j]] = (\n            v48_pass2_rank_flip[:, [_rad_j, _rad_i]]\n        )\n    _rad_merged = (v48_pass2_rank + _rad_flip_aligned) / 2.0\n    for _rad_t in _RAD_MIRROR_UNPAIRED:\n        _rad_k = _RAD_LABELS.index(_rad_t)\n        _rad_merged[:, _rad_k] = v48_pass2_rank[:, _rad_k]\n    v48_pass2_rank = _rad_rank_columns(_rad_merged)\n    _rad_log(\n        'mirror TTA realigned: medial/lateral columns swapped back, '\n        f'{list(_RAD_MIRROR_UNPAIRED)} kept from the unflipped pass'\n    )\n    del v48_features_flip, v48_token_mask_flip, v48_pass2_predictions_flip\n    del v48_pass2_probability_flip, v48_pass2_rank_flip\n    _rad_gc.collect()\n    _rad_torch.cuda.empty_cache()\n\n    reference_branch = candidate_reference.copy()\n    reference_branch[_RAD_LABELS] = _rad_rank_columns(\n        (1.0 - _RAD_V48_SECOND_ALPHA)\n        * _rad_rank_columns(candidate_reference[_RAD_LABELS].to_numpy())\n        + _RAD_V48_SECOND_ALPHA * v48_pass2_rank\n    )\n    _rad_validate(reference_branch, expected_ids)\n    _rad_log(\n        f'V48 second E13 pass complete at alpha '\n        f'{_RAD_V48_SECOND_ALPHA:.2f} on the E11 slot layout'\n    )\n\n    # V48 deploys the pinned reference branch directly after the second pass.\n    # The alternative-head twin and legacy-DINO tie-break are not part of .917.\n    _RAD_CAL = _rad_json.loads(_rad_zlib.decompress(\n        _rad_b64.b64decode(_RAD_CAL_PAYLOAD)).decode())\n    _RAD_CAL_GATE = set(_RAD_CAL['gate'])\n    _RAD_CAL_W = 0.40\n\n    def _rad_cal_protocol(uids):\n        frame = _rad_pd.read_csv(\n            ROOT / 'test_series.csv',\n            dtype={'StudyInstanceUID': str, 'SeriesInstanceUID': str},\n        )\n        frame['StudyInstanceUID'] = frame['StudyInstanceUID'].astype(str)\n        index = _rad_pd.Index([str(u) for u in uids], name='StudyInstanceUID')\n        table = _rad_pd.DataFrame(index=index)\n        table['n_series'] = frame.groupby(\n            'StudyInstanceUID').size().reindex(index).fillna(0)\n        for plane in ('Sagittal', 'Coronal', 'Axial'):\n            part = frame[frame['Anatomical_Plane'].astype(str) == plane]\n            table[f'n_{plane[:3]}'] = part.groupby(\n                'StudyInstanceUID').size().reindex(index).fillna(0)\n        for flag in ('Fat_Suppression', 'Fluid_Sensitive'):\n            marked = frame[_rad_pd.to_numeric(\n                frame[flag], errors='coerce').fillna(0) > 0]\n            table[flag[:3]] = marked.groupby(\n                'StudyInstanceUID').size().reindex(index).fillna(0)\n            for plane in ('Sagittal', 'Coronal', 'Axial'):\n                part = marked[marked['Anatomical_Plane'].astype(str) == plane]\n                table[f'{flag[:3]}_{plane[:3]}'] = part.groupby(\n                    'StudyInstanceUID').size().reindex(index).fillna(0)\n        if list(table.columns) != list(_RAD_CAL['protocol_columns']):\n            raise RuntimeError('calibration protocol layout mismatch')\n        return table.to_numpy(_rad_np.float64)\n\n    def _rad_calibrate(branch):\n        base = baseline_rank\n        public = reference_rank\n        pass2 = v48_pass2_rank\n        mean = (base + public + pass2) / 3.0\n        blocks = [base, public, pass2, public - base, pass2 - base, mean]\n        for _grp in _RAD_CAL['groups']:\n            cols = [_RAD_LABELS.index(t) for t in _grp]\n            blocks.append(mean[:, cols].mean(axis=1, keepdims=True))\n        blocks.append(_rad_cal_protocol(expected_ids))\n        x = _rad_np.concatenate(blocks, axis=1)\n        centre = _rad_np.asarray(_RAD_CAL['mean'], _rad_np.float64)\n        spread = _rad_np.asarray(_RAD_CAL['scale'], _rad_np.float64)\n        coef = _rad_np.asarray(_RAD_CAL['coef'], _rad_np.float64)\n        bias = _rad_np.asarray(_RAD_CAL['intercept'], _rad_np.float64)\n        if x.shape[1] != coef.shape[1]:\n            raise RuntimeError(\n                f'calibration expects {coef.shape[1]} columns, built {x.shape[1]}')\n        adjusted = _rad_rank_columns(((x - centre) / spread) @ coef.T + bias)\n        out = branch.copy()\n        values = out[_RAD_LABELS].to_numpy(_rad_np.float64).copy()\n        for index, target in enumerate(_RAD_LABELS):\n            if target in _RAD_CAL_GATE:\n                values[:, index] = (\n                    (1.0 - _RAD_CAL_W) * values[:, index]\n                    + _RAD_CAL_W * adjusted[:, index]\n                )\n        out[_RAD_LABELS] = _rad_rank_columns(values)\n        _rad_validate(out, expected_ids)\n        return out\n\n    final = _rad_calibrate(reference_branch)\n    globals()['V18_CALIBRATOR_APPLIED'] = True\n    globals()['V18_CAL_GATE'] = tuple(sorted(_RAD_CAL_GATE))\n    _rad_validate(final, expected_ids)\n    temporary = primary.with_suffix('.csv.tmp')\n    final.to_csv(temporary, index=False)\n    _rad_os.replace(temporary, primary)\n\n    receipt = {\n        'recipe': 'V48 deployed reference branch: correct E13@0.50-inside-Rad -> E10@0.50 -> same E13 on E11 layout@0.15',\n        'e13_member_weight_inside_rad': _RAD_E13_MEMBER_WEIGHT,\n        'e10_alpha': _RAD_ALPHA,\n        'e10_preserved_targets': list(_RAD_EXCLUDE),\n        'v48_second_alpha': _RAD_V48_SECOND_ALPHA,\n        'reference_heads_sha256': _RAD_REFERENCE_HEADS_SHA256,\n        'e13_heads_sha256': _RAD_E13_HEADS_SHA256,\n        'v48_second_heads_sha256': _RAD_E13_HEADS_SHA256,\n        'v48_second_slots': [list(slot) for slot in _RAD_E11_SLOTS],\n        'encoder_sha256': _RAD_ENCODER_SHA256,\n        'test_studies': len(expected_ids),\n        'e10_tokens': token_count,\n        'v48_second_tokens': v48_pass2_token_count,\n        'e13_tokens': e13_token_count,\n        'submission_sha256': _rad_sha256(primary),\n    }\n    (work / 'v50_v2_repro_receipt.json').write_text(\n        _rad_json.dumps(receipt, indent=2, sort_keys=True) + '\\n'\n    )\n    del (encoder, e13_heads, v48_pass2_predictions,\n         v48_pass2_probability, v48_pass2_rank, v48_features, v48_token_mask)\n    _rad_gc.collect()\n    _rad_torch.cuda.empty_cache()\n    _rad_log(\n        f'V48 reference branch complete; reference-v15={reference_heads_path}; '\n        f'e13-two-pass={e13_path}; second_alpha='\n        f'{_RAD_V48_SECOND_ALPHA:.2f}; encoder={encoder_path}; '\n        f'elapsed={(_rad_time.time()-started)/60:.1f}m'\n    )\n\n\ndef _rad_main_guarded():\n    import shutil as _rad_shutil\n    import traceback as _rad_traceback\n    primary = _RadPath('/kaggle/working/submission.csv')\n    snapshot = _RadPath('/kaggle/working/submission_pre_rad.csv')\n    if primary.is_file():\n        _rad_shutil.copyfile(primary, snapshot)\n    try:\n        _rad_main()\n    except Exception as error:\n        _rad_traceback.print_exc()\n        if snapshot.is_file():\n            _rad_shutil.copyfile(snapshot, primary)\n        _rad_log(\n            f'branch ABANDONED ({type(error).__name__}: {error}); '\n            'submission.csv restored to the pre-Rad DINO parent so the '\n            'Raptor branch can still run'\n        )\n        globals()['V18_CALIBRATOR_APPLIED'] = False\n\n\n_rad_main_guarded()"},{"cell_type":"markdown","id":"05f103b6","metadata":{"papermill":{"duration":0.014551,"end_time":"2026-09-12T11:19:43.735334+00:00","exception":false,"start_time":"2026-09-12T11:19:43.720783+00:00","status":"completed"},"tags":[]},"source":"# 23. 🦖 CoAtNet 4 руки + residual-gated + outer-routing\n\n## Что делает эта ячейка\n\nФинальная и самая важная часть пайплайна:\n\n1. Воспроизводит публичный Raptor (4 руки CoAtNet).\n2. Добавляет residual-gated CoAt (e4/e6/e8) с alpha 0.40.\n3. Смешивает с Transformer/Rad.\n4. Для Lateral Meniscus меняет outer-routing: Raptor = 100%, Transformer/Rad = 0%.\n5. Создаёт итоговый `submission.csv`.\n\nИменно эта ячейка отличает решение 0.94 от 0.939."},{"cell_type":"code","execution_count":null,"id":"e614cebf","metadata":{"execution":{"iopub.execute_input":"2026-09-12T11:19:43.766811Z","iopub.status.busy":"2026-09-12T11:19:43.766561Z","iopub.status.idle":"2026-09-12T11:21:37.846253Z","shell.execute_reply":"2026-09-12T11:21:37.84541Z"},"papermill":{"duration":114.098626,"end_time":"2026-09-12T11:21:37.84798+00:00","exception":false,"start_time":"2026-09-12T11:19:43.749354+00:00","status":"completed"},"tags":[]},"outputs":[],"source":"import numpy as _ke_np\nimport pandas as _ke_pd\nfrom pathlib import Path as _KePath\n\n_ke_primary = _KePath('/kaggle/working/submission.csv')\n_ke_ours = _ke_pd.read_csv(\n    _ke_primary,\n    dtype={'StudyInstanceUID': str},\n)\n_KE_LAB = [\n    c\n    for c in _ke_ours.columns\n    if c != 'StudyInstanceUID'\n]\n\n\n# =====================================================================\n# Public Raptor reproduction\n# =====================================================================\n\n_KE_SRC = 'import os, glob, time, gc, hashlib\\nos.environ.setdefault(\\'HF_HUB_OFFLINE\\', \\'1\\')\\nos.environ.setdefault(\\'TRANSFORMERS_OFFLINE\\', \\'1\\')\\nos.environ.setdefault(\\'HF_HUB_DISABLE_TELEMETRY\\', \\'1\\')\\nimport numpy as np\\nimport torch, torch.nn as nn, torch.nn.functional as F\\nimport timm\\ntorch.backends.cudnn.benchmark = True\\ntorch.backends.cuda.matmul.allow_tf32 = True\\nIMG = 336\\nCROP_MM = 140.0\\nSPAN_LO, SPAN_HI = 0.02, 0.98\\nSLOTS = [(\"Sagittal\", 1, 18), (\"Sagittal\", 0, 14),\\n         (\"Coronal\", 1, 12), (\"Coronal\", 0, 8), (\"Axial\", -1, 12)]\\nMAXS = sum(slot[2] for slot in SLOTS)\\nK_EVAL = 62\\nNORM = \"imagenet\"\\nLAB = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\\n       \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker\\'s\",\\n       \"Contusion\", \"Fracture\"]\\n_MEAN = torch.tensor([0.485, 0.456, 0.406]).view(3, 1, 1)\\n_STD = torch.tensor([0.229, 0.224, 0.225]).view(3, 1, 1)\\n_SLOTS64 = [(\"Sagittal\", 1, 18), (\"Sagittal\", 0, 14),\\n            (\"Coronal\", 1, 12), (\"Coronal\", 0, 8), (\"Axial\", -1, 12)]\\n_SLOTS44 = [(\"Sagittal\", 1, 12), (\"Sagittal\", 0, 10),\\n            (\"Coronal\", 1, 8), (\"Coronal\", 0, 6), (\"Axial\", -1, 8)]\\nARMS = [\\n    {\"name\": \"maxspan-v5\", \"file\": \"raptor_ft_coatnet_v5_full_swa.pt\",\\n     \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\", \"res\": 384,\\n     \"img\": 336, \"slots\": _SLOTS64, \"span\": (0.02, 0.98), \"k_eval\": 62,\\n     \"reverse\": False, \"w\": 0.60},\\n    {\"name\": \"native384dense-v10\", \"file\": \"raptor_ft_coatnet_v10_full.pt\",\\n     \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\", \"res\": 384,\\n     \"img\": 384, \"slots\": _SLOTS64, \"span\": (0.02, 0.98), \"k_eval\": 62,\\n     \"reverse\": False, \"w\": 0.10},\\n    {\"name\": \"maxspan-v5-reverse\", \"file\": \"raptor_ft_coatnet_v5_full_swa.pt\",\\n     \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\", \"res\": 384,\\n     \"img\": 336, \"slots\": _SLOTS64, \"span\": (0.02, 0.98), \"k_eval\": 62,\\n     \"reverse\": True, \"w\": 0.10},\\n    {\"name\": \"native384-v8\", \"file\": \"raptor_ft_coatnet_v8_full_swa.pt\",\\n     \"arch\": \"coatnet_rmlp_2_rw_384.sw_in12k_ft_in1k\", \"res\": 384,\\n     \"img\": 384, \"slots\": _SLOTS44, \"span\": (0.06, 0.94), \"k_eval\": 42,\\n     \"reverse\": False, \"w\": 0.20},\\n]\\n\\ndef build_backbone(arch, pretrained=False):\\n    hybrid = arch.startswith((\\'maxvit\\', \\'maxxvit\\', \\'coatnet\\', \\'coat_\\', \\'convnext\\'))\\n    is_vit = not hybrid and any((k in arch for k in (\\'vit\\', \\'deit\\', \\'dinov2\\', \\'eva\\', \\'beit\\')))\\n    kw = dict(pretrained=pretrained, num_classes=0, in_chans=3)\\n    if is_vit:\\n        kw.update(global_pool=\\'token\\', dynamic_img_size=True)\\n    else:\\n        kw.update(global_pool=\\'avg\\')\\n    return timm.create_model(arch, **kw)\\n\\nclass RaptorClassifier(nn.Module):\\n\\n    def __init__(self, backbone, F_dim=768, n=12, drop=0.2):\\n        super().__init__()\\n        self.backbone = backbone\\n        self.norm = nn.LayerNorm(F_dim)\\n        self.att = nn.Sequential(\\n            nn.Linear(F_dim, 256),\\n            nn.Tanh(),\\n            nn.Dropout(drop),\\n            nn.Linear(256, n),\\n        )\\n        self.clsW = nn.Parameter(torch.zeros(n, F_dim))\\n        self.clsb = nn.Parameter(torch.zeros(n))\\n        nn.init.trunc_normal_(self.clsW, std=0.02)\\n        self.n = n\\n\\n    def encode(self, x):\\n        B, K = x.shape[:2]\\n        f = self.backbone(x.flatten(0, 1))\\n        return f.view(B, K, -1)\\n\\n    def head(self, feats):\\n        h = self.norm(feats)\\n        a = self.att(h)\\n        a = torch.softmax(a, dim=1)\\n        pooled = torch.einsum(\\'bkn,bkf->bnf\\', a, h)\\n        logits = (pooled * self.clsW).sum(-1) + self.clsb\\n        return logits\\n\\n    def forward(self, x):\\n        return self.head(self.encode(x))\\n\\ndef load_model(pt_path, arch_default, res_default, device, ngpu=1):\\n    ck = torch.load(\\n        pt_path,\\n        map_location=\\'cpu\\',\\n        weights_only=False,\\n    )\\n    arch = ck.get(\\'arch\\', arch_default)\\n    ck_res = int(ck.get(\\'res\\', res_default))\\n    bb = build_backbone(arch, pretrained=False)\\n    model = RaptorClassifier(\\n        bb,\\n        F_dim=bb.num_features,\\n    )\\n    model.load_state_dict(\\n        ck[\\'model\\'],\\n        strict=True,\\n    )\\n    model.eval().to(device)\\n    del ck\\n    gc.collect()\\n    return (model, ck_res)\\n\\ndef _eval_centers(mask, D, k):\\n    valid = np.where(mask > 0)[0]\\n    if len(valid) < 3:\\n        valid = np.arange(min(3, D))\\n    lo, hi = (int(valid.min()), int(valid.max()))\\n    cs = [\\n        c\\n        for c in range(lo + 1, hi)\\n        if c - 1 >= lo and c + 1 <= hi\\n    ]\\n    if not cs:\\n        cs = [\\n            max(\\n                1,\\n                min((lo + hi) // 2, D - 2),\\n            )\\n        ]\\n    idx = (\\n        np.linspace(0, len(cs) - 1, k)\\n        .round()\\n        .astype(int)\\n    )\\n    return [cs[i] for i in idx]\\n\\ndef eval_windows(vol, mask, k, res, norm=NORM):\\n    D = vol.shape[0]\\n    cs = _eval_centers(mask, D, k)\\n    wins = np.empty(\\n        (len(cs), 3, res, res),\\n        np.float32,\\n    )\\n    for j, c in enumerate(cs):\\n        c = max(1, min(c, D - 2))\\n        tri = np.stack(\\n            [\\n                vol[c - 1],\\n                vol[c],\\n                vol[c + 1],\\n            ],\\n            0,\\n        ).astype(np.float32) / 255.0\\n        t = torch.from_numpy(tri)\\n        if t.shape[-1] != res:\\n            t = F.interpolate(\\n                t[None],\\n                size=(res, res),\\n                mode=\\'bilinear\\',\\n                align_corners=False,\\n            )[0]\\n        wins[j] = t.numpy()\\n    x = torch.from_numpy(wins)\\n    if norm == \\'imagenet\\':\\n        x = (x - _MEAN) / _STD\\n    return x\\n\\n@torch.no_grad()\\ndef infer_probs(model, xwins, device):\\n    x = xwins.unsqueeze(0).to(device)\\n    use_cuda = (\\n        device != \\'cpu\\'\\n        and str(device).startswith(\\'cuda\\')\\n    )\\n    if use_cuda:\\n        try:\\n            with torch.autocast(\\n                \\'cuda\\',\\n                dtype=torch.float16,\\n            ):\\n                o = torch.sigmoid(\\n                    model(x).float()\\n                )\\n            return o[0].cpu().numpy()\\n        except RuntimeError:\\n            torch.cuda.empty_cache()\\n            o = torch.sigmoid(\\n                model(x).float()\\n            )\\n            return o[0].cpu().numpy()\\n    o = torch.sigmoid(model(x).float())\\n    return o[0].cpu().numpy()\\n\\ndef rankpct(x):\\n    order = (\\n        x.argsort(0)\\n        .argsort(0)\\n        .astype(np.float64)\\n    )\\n    return order / max(\\n        1,\\n        x.shape[0] - 1,\\n    )\\n\\ndef _make_reader():\\n    import pydicom, cv2\\n    from pydicom.pixel_data_handlers.util import apply_modality_lut\\n\\n    def order_and_meta(sdir):\\n        fs = glob.glob(sdir + \\'/*.dcm\\')\\n        recs = []\\n        ps_list = []\\n        for f in fs:\\n            try:\\n                h = pydicom.dcmread(\\n                    f,\\n                    stop_before_pixels=True,\\n                )\\n                iop = getattr(\\n                    h,\\n                    \\'ImageOrientationPatient\\',\\n                    None,\\n                )\\n                ipp = getattr(\\n                    h,\\n                    \\'ImagePositionPatient\\',\\n                    None,\\n                )\\n                if (\\n                    iop is not None\\n                    and ipp is not None\\n                    and len(iop) == 6\\n                ):\\n                    r = np.array(iop[:3], float)\\n                    c = np.array(iop[3:], float)\\n                    n = np.cross(r, c)\\n                    pos = float(\\n                        np.dot(\\n                            np.array(ipp, float),\\n                            n,\\n                        )\\n                    )\\n                else:\\n                    pos = float(\\n                        getattr(\\n                            h,\\n                            \\'InstanceNumber\\',\\n                            0,\\n                        ) or 0\\n                    )\\n                ps = getattr(\\n                    h,\\n                    \\'PixelSpacing\\',\\n                    None,\\n                )\\n                ps = (\\n                    float(ps[0])\\n                    if ps is not None\\n                    else 0.5\\n                )\\n                ps_list.append(ps)\\n                recs.append((pos, f, ps))\\n            except Exception:\\n                recs.append((0.0, f, 0.5))\\n        recs.sort(key=lambda x: x[0])\\n        med_ps = (\\n            float(np.median(ps_list))\\n            if ps_list\\n            else 0.5\\n        )\\n        return (\\n            [(f, ps) for _, f, ps in recs],\\n            med_ps,\\n        )\\n\\n    def read_px(f):\\n        d = pydicom.dcmread(f)\\n        a = apply_modality_lut(\\n            d.pixel_array,\\n            d,\\n        ).astype(np.float32)\\n        if str(\\n            getattr(\\n                d,\\n                \\'PhotometricInterpretation\\',\\n                \\'\\',\\n            )\\n        ) == \\'MONOCHROME1\\':\\n            a = a.max() - a\\n        return a\\n\\n    def mm_crop_resize(a, ps):\\n        h, w = a.shape\\n        cpx = int(\\n            round(\\n                CROP_MM\\n                / max(ps, 0.001)\\n            )\\n        )\\n        cpx = min(\\n            cpx,\\n            min(h, w),\\n        )\\n        y0 = (h - cpx) // 2\\n        x0 = (w - cpx) // 2\\n        a = a[\\n            y0:y0 + cpx,\\n            x0:x0 + cpx,\\n        ]\\n        return cv2.resize(\\n            a,\\n            (IMG, IMG),\\n            interpolation=cv2.INTER_AREA,\\n        )\\n\\n    return (\\n        order_and_meta,\\n        read_px,\\n        mm_crop_resize,\\n    )\\n\\ndef _pick_series_for_slot(rows, plane, fluid, used):\\n    cands = [\\n        r\\n        for r in rows\\n        if (\\n            r[\\'Anatomical_Plane\\'] == plane\\n            and r[\\'SeriesInstanceUID\\'] not in used\\n        )\\n    ]\\n    if fluid in (0, 1):\\n        pref = [\\n            r\\n            for r in cands\\n            if int(\\n                r.get(\\n                    \\'Fluid_Sensitive\\',\\n                    0,\\n                ) or 0\\n            ) == fluid\\n        ]\\n        if pref:\\n            return pref[0]\\n    return cands[0] if cands else None\\n\\ndef build_study(sid, ser_records, tsdir, reader):\\n    order_and_meta, read_px, mm_crop_resize = reader\\n    rows = ser_records.get(sid, [])\\n    vol = np.zeros(\\n        (MAXS, IMG, IMG),\\n        np.uint8,\\n    )\\n    idx = 0\\n    used = set()\\n\\n    for plane, fluid, k in SLOTS:\\n        r = _pick_series_for_slot(\\n            rows,\\n            plane,\\n            fluid,\\n            used,\\n        )\\n        if r is None:\\n            idx += k\\n            continue\\n\\n        used.add(r[\\'SeriesInstanceUID\\'])\\n\\n        files, med_ps = order_and_meta(\\n            f\"{tsdir}/{sid}/\"\\n            f\"{r[\\'SeriesInstanceUID\\']}\"\\n        )\\n\\n        if not files:\\n            idx += k\\n            continue\\n\\n        n = len(files)\\n        lo, hi = (\\n            int(n * SPAN_LO),\\n            int(n * SPAN_HI) - 1,\\n        )\\n        hi = max(hi, lo)\\n\\n        picks = (\\n            np.linspace(\\n                lo,\\n                hi,\\n                k,\\n            ).round().astype(int)\\n            if n > 1\\n            else [0] * k\\n        )\\n\\n        arrs = []\\n        pss = []\\n\\n        for p in picks:\\n            fp, ps = files[min(p, n - 1)]\\n            try:\\n                arrs.append(read_px(fp))\\n                pss.append(ps)\\n            except Exception:\\n                arrs.append(None)\\n                pss.append(med_ps)\\n\\n        valid = [\\n            a\\n            for a in arrs\\n            if a is not None\\n        ]\\n\\n        if valid:\\n            allpx = np.concatenate(\\n                [a.ravel() for a in valid]\\n            )\\n            loq, hiq = np.percentile(\\n                allpx,\\n                [2.0, 98.0],\\n            )\\n        else:\\n            loq, hiq = (0.0, 1.0)\\n\\n        for a, ps in zip(arrs, pss):\\n            if idx >= MAXS:\\n                break\\n            if a is None:\\n                idx += 1\\n                continue\\n\\n            aw = np.clip(\\n                (a - loq)\\n                / (hiq - loq + 1e-06),\\n                0,\\n                1,\\n            )\\n\\n            aw = mm_crop_resize(\\n                aw,\\n                ps if ps > 0 else med_ps,\\n            )\\n\\n            vol[idx] = (\\n                aw * 255\\n            ).astype(np.uint8)\\n\\n            idx += 1\\n\\n        if idx >= MAXS:\\n            break\\n\\n    mask = (\\n        vol.reshape(MAXS, -1)\\n        .sum(1)\\n        > 0\\n    ).astype(np.uint8)\\n\\n    return (vol, mask)\\n\\ndef find_test_root():\\n    cands = [\\n        \\'/kaggle/input/competitions/rsna-knee-abnormality-detection\\',\\n        \\'/kaggle/input/rsna-knee-abnormality-detection\\',\\n    ]\\n\\n    for b in cands:\\n        if os.path.exists(b + \\'/test.csv\\'):\\n            return b\\n\\n    for d, _, f in os.walk(\\'/kaggle/input\\'):\\n        if (\\n            \\'test.csv\\' in f\\n            and (\\n                os.path.isdir(d + \\'/test_series\\')\\n                or os.path.isdir(d + \\'/test_images\\')\\n            )\\n        ):\\n            return d\\n\\n    for d, _, f in os.walk(\\'/kaggle/input\\'):\\n        if \\'test.csv\\' in f:\\n            return d\\n\\n    raise RuntimeError(\\n        \\'no test root under /kaggle/input\\'\\n    )\\n\\ndef find_weight_file(fname):\\n    direct = [\\n        f\\'/kaggle/input/raptor-knee-maxspan/{fname}\\',\\n        f\\'/kaggle/input/raptor-knee-native384dense/{fname}\\',\\n        f\\'/kaggle/input/raptor-knee-native384/{fname}\\',\\n        f\\'/kaggle/input/raptor-knee-arms/{fname}\\',\\n        f\\'/kaggle/input/raptor-knee-arms/1/{fname}\\',\\n        f\\'/kaggle/input/raptor-cnn336/{fname}\\',\\n    ]\\n\\n    for p in direct:\\n        if os.path.exists(p):\\n            return p\\n\\n    for d in sorted(\\n        glob.glob(\\'/kaggle/input/*/\\')\\n    ):\\n        if \\'competition\\' in d.lower():\\n            continue\\n        hits = []\\n        for root, dirs, files in os.walk(d):\\n            dirs[:] = [\\n                x\\n                for x in dirs\\n                if x not in (\\n                    \\'train_series\\',\\n                    \\'test_series\\',\\n                )\\n            ]\\n            if fname in files:\\n                hits.append(\\n                    os.path.join(root, fname)\\n                )\\n        if hits:\\n            return sorted(hits)[0]\\n\\n    raise RuntimeError(\\n        f\\'{fname} not found under /kaggle/input\\'\\n    )\\n\\ndef main():\\n    import pandas as pd\\n\\n    t0 = time.time()\\n    dev = (\\n        \"cuda\"\\n        if torch.cuda.is_available()\\n        else \"cpu\"\\n    )\\n\\n    print(\\n        f\"device {dev} | \"\\n        f\"gpus {torch.cuda.device_count()} | \"\\n        f\"torch {torch.__version__}\",\\n        flush=True,\\n    )\\n\\n    root = find_test_root()\\n    tsdir = root + \"/test_series\"\\n\\n    if not os.path.isdir(tsdir):\\n        tsdir = root + \"/test_images\"\\n\\n    print(\\n        \"test root:\",\\n        root,\\n        \"| series dir:\",\\n        tsdir,\\n        flush=True,\\n    )\\n\\n    test = pd.read_csv(root + \"/test.csv\")\\n    test[\"StudyInstanceUID\"] = (\\n        test[\"StudyInstanceUID\"].astype(str)\\n    )\\n    test_ids = test[\"StudyInstanceUID\"].tolist()\\n\\n    tser = pd.read_csv(\\n        root + \"/test_series.csv\"\\n    )\\n    tser[\"StudyInstanceUID\"] = (\\n        tser[\"StudyInstanceUID\"].astype(str)\\n    )\\n    tser[\"SeriesInstanceUID\"] = (\\n        tser[\"SeriesInstanceUID\"].astype(str)\\n    )\\n\\n    series = {\\n        key: frame.to_dict(\"records\")\\n        for key, frame\\n        in tser.groupby(\"StudyInstanceUID\")\\n    }\\n\\n    print(\\n        f\"test studies {len(test_ids)} | \"\\n        f\"test series {len(tser)}\",\\n        flush=True,\\n    )\\n\\n    sub_cols = [\\n        \"StudyInstanceUID\",\\n        *LAB,\\n    ]\\n\\n    sample = os.path.join(\\n        root,\\n        \"sample_submission.csv\",\\n    )\\n\\n    if os.path.exists(sample):\\n        sub_cols = list(\\n            pd.read_csv(\\n                sample,\\n                nrows=1,\\n            ).columns\\n        )\\n\\n    reader = _make_reader()\\n    n_study = len(test_ids)\\n    n_arm = len(ARMS)\\n\\n    arm_probs = [\\n        np.full(\\n            (n_study, len(LAB)),\\n            0.5,\\n            np.float32,\\n        )\\n        for _ in range(n_arm)\\n    ]\\n\\n    passes = []\\n\\n    for arm_index, arm in enumerate(ARMS):\\n        key = (\\n            arm[\"file\"],\\n            arm[\"arch\"],\\n            int(arm[\"res\"]),\\n            int(arm[\"img\"]),\\n            tuple(tuple(slot) for slot in arm[\"slots\"]),\\n            tuple(float(v) for v in arm[\"span\"]),\\n            int(arm[\"k_eval\"]),\\n        )\\n\\n        for group in passes:\\n            if group[\"key\"] == key:\\n                group[\"arms\"].append((arm_index, arm))\\n                break\\n        else:\\n            passes.append({\"key\": key, \"arms\": [(arm_index, arm)]})\\n\\n    print(\\n        f\"[plan] {n_arm} arm(s) in {len(passes)} decode pass(es): \"\\n        + \" | \".join(\\n            \"+\".join(a[\"name\"] for _, a in group[\"arms\"])\\n            for group in passes\\n        ),\\n        flush=True,\\n    )\\n\\n    for pass_index, group in enumerate(passes):\\n        head_arm = group[\"arms\"][0][1]\\n        globals()[\"IMG\"] = int(head_arm[\"img\"])\\n        globals()[\"SLOTS\"] = list(head_arm[\"slots\"])\\n        globals()[\"MAXS\"] = sum(\\n            slot[2]\\n            for slot in SLOTS\\n        )\\n        globals()[\"SPAN_LO\"], globals()[\"SPAN_HI\"] = map(\\n            float,\\n            head_arm[\"span\"],\\n        )\\n        globals()[\"K_EVAL\"] = int(\\n            head_arm[\"k_eval\"]\\n        )\\n\\n        weight_path = find_weight_file(\\n            head_arm[\"file\"]\\n        )\\n\\n        model, resolution = load_model(\\n            weight_path,\\n            head_arm[\"arch\"],\\n            head_arm[\"res\"],\\n            dev,\\n        )\\n\\n        names = \"+\".join(\\n            a[\"name\"] for _, a in group[\"arms\"]\\n        )\\n\\n        print(\\n            f\"[pass {pass_index}] \"\\n            f\"arms {[i for i, _ in group[\\'arms\\']]} {names} | \"\\n            f\"img {IMG} | \"\\n            f\"slices {MAXS} | \"\\n            f\"span {SPAN_LO:.2f}-{SPAN_HI:.2f} | \"\\n            f\"windows {K_EVAL} | \"\\n            f\"res {resolution} | \"\\n            f\"{time.time() - t0:.0f}s\",\\n            flush=True,\\n        )\\n\\n        for study_index, study_uid in enumerate(test_ids):\\n            try:\\n                volume, mask = build_study(\\n                    study_uid,\\n                    series,\\n                    tsdir,\\n                    reader,\\n                )\\n\\n                windows = eval_windows(\\n                    volume,\\n                    mask,\\n                    k=K_EVAL,\\n                    res=resolution,\\n                    norm=NORM,\\n                )\\n\\n                for arm_index, arm in group[\"arms\"]:\\n                    if bool(\\n                        arm.get(\"reverse\", False)\\n                    ):\\n                        xwins = (\\n                            windows.flip(1)\\n                            .contiguous()\\n                        )\\n                    else:\\n                        xwins = windows\\n\\n                    arm_probs[arm_index][study_index] = infer_probs(\\n                        model,\\n                        xwins,\\n                        dev,\\n                    )\\n\\n                    del xwins\\n\\n                del volume, mask, windows\\n\\n            except Exception as error:\\n                print(\\n                    f\"  [pass {pass_index}] \"\\n                    f\"study {study_index} \"\\n                    f\"{study_uid[:16]} FALLBACK \"\\n                    f\"({type(error).__name__}: \"\\n                    f\"{error})\",\\n                    flush=True,\\n                )\\n\\n            if (\\n                (study_index + 1) % 100 == 0\\n                or study_index + 1 == n_study\\n            ):\\n                print(\\n                    f\"  [pass {pass_index}] \"\\n                    f\"{study_index + 1}/{n_study} | \"\\n                    f\"{time.time() - t0:.0f}s\",\\n                    flush=True,\\n                )\\n\\n        del model\\n        gc.collect()\\n\\n        if str(dev).startswith(\"cuda\"):\\n            torch.cuda.empty_cache()\\n\\n        print(\\n            f\"[pass {pass_index}] done + freed | \"\\n            f\"{time.time() - t0:.0f}s\",\\n            flush=True,\\n        )\\n\\n    weights = np.array(\\n        [\\n            float(arm.get(\"w\", 1.0))\\n            for arm in ARMS\\n        ],\\n        dtype=np.float64,\\n    )\\n    weights /= weights.sum()\\n\\n    print(\\n        \"[blend] global probability mean w=\"\\n        f\"{dict(zip([arm[\\'name\\'] for arm in ARMS], weights.round(4)))}\",\\n        flush=True,\\n    )\\n\\n    probability_blend = np.tensordot(\\n        weights,\\n        np.stack(\\n            [\\n                np.clip(values, 0, 1)\\n                for values in arm_probs\\n            ]\\n        ),\\n        axes=(0, 0),\\n    )\\n\\n    ranks = rankpct(probability_blend)\\n\\n    if not np.isfinite(ranks).all():\\n        ranks[\\n            ~np.isfinite(ranks)\\n        ] = 0.5\\n\\n    submission = pd.DataFrame(\\n        ranks.astype(np.float32),\\n        columns=LAB,\\n    )\\n    submission.insert(\\n        0,\\n        \"StudyInstanceUID\",\\n        test_ids,\\n    )\\n    submission = submission[sub_cols]\\n\\n    assert (\\n        submission[\"StudyInstanceUID\"].tolist()\\n        == test_ids\\n    )\\n    assert np.isfinite(\\n        submission[LAB].values\\n    ).all()\\n\\n    out = \"/kaggle/working/_raptor.csv\"\\n    submission.to_csv(\\n        out,\\n        index=False,\\n    )\\n\\n    print(\\n        \"wrote\",\\n        out,\\n        \"|\",\\n        len(submission),\\n        \"rows x\",\\n        len(submission.columns),\\n        \"cols\",\\n        flush=True,\\n    )\\n\\n    print(\\n        submission.head().to_string(index=False),\\n        flush=True,\\n    )\\n\\n    print(\\n        f\"DONE {time.time() - t0:.0f}s\",\\n        flush=True,\\n    )\\n'\n\n_KE_NS = {\n    '__name__': '_ke_raptor'\n}\n\nexec(\n    compile(\n        _KE_SRC,\n        '<raptor>',\n        'exec',\n    ),\n    _KE_NS,\n)\n\n_ke_arms_found = []\nfor _ke_arm in _KE_NS['ARMS']:\n    try:\n        _KE_NS['find_weight_file'](_ke_arm['file'])\n    except Exception as _ke_error:\n        print(f\"[raptor] arm {_ke_arm['name']} dropped \"\n              f\"({type(_ke_error).__name__}: {_ke_error})\", flush=True)\n        continue\n    _ke_arms_found.append(_ke_arm)\n\nif not _ke_arms_found:\n    raise RuntimeError('no Raptor arm checkpoint is mounted')\n\nif len(_ke_arms_found) != len(_KE_NS['ARMS']):\n    _ke_kept = sum(float(a['w']) for a in _ke_arms_found)\n    print(f\"[raptor] running {len(_ke_arms_found)} of \"\n          f\"{len(_KE_NS['ARMS'])} arms; weights renormalised over \"\n          f\"{_ke_kept:.2f} of the original 1.00\", flush=True)\n    _KE_NS['ARMS'] = _ke_arms_found\n\n_KE_NS['main']()\n\n\ndef _coat_substitute():\n    import hashlib as _h\n    import os as _o\n    import subprocess as _sp\n    import sys as _sy\n    from pathlib import Path as _P\n\n    import pandas as _pd\n    import torch as _torch\n\n    _n_gpu = (\n        _torch.cuda.device_count()\n        if _torch.cuda.is_available()\n        else 0\n    )\n\n    if _n_gpu != 2:\n        raise RuntimeError(\n            'the residual-gated CoAt arm requires T4 x2; this session '\n            f'has {_n_gpu} CUDA device(s), so the child process would '\n            'fail on its own assert with no name attached'\n        )\n\n    MAN_SHA = (\n        '98511a8fdeb9da0e6e70c78d013ff636e'\n        '1476f31c80b5dc134d294b18c3f284e'\n    )\n\n    WHL_SHA = (\n        '236c8df54a90f4d02076e6f9c1cc763d'\n        '794542e886c576a6fee46ec8ff75a7a9'\n    )\n\n    raptor = _P(\n        '/kaggle/working/_raptor.csv'\n    )\n\n    def sha(p):\n        d = _h.sha256()\n        with _P(p).open('rb') as f:\n            for b in iter(\n                lambda: f.read(8 << 20),\n                b'',\n            ):\n                d.update(b)\n        return d.hexdigest()\n\n    def find(name, want):\n        root = _P('/kaggle/input')\n\n        if not root.is_dir():\n            return None\n\n        for base in sorted(root.iterdir()):\n            if base.name in (\n                'competitions',\n                'train_series',\n                'test_series',\n            ):\n                continue\n\n            for p in sorted(\n                base.rglob(name)\n            ):\n                if (\n                    p.is_file()\n                    and sha(p) == want\n                ):\n                    return p\n\n        return None\n\n    man = find(\n        'coat_resgated_ep10_top3_manifest.json',\n        MAN_SHA,\n    )\n\n    if man is None:\n        raise RuntimeError(\n            'coat manifest absent or hash mismatch'\n        )\n\n    art = man.parent\n\n    whl = find(\n        'opencv_python_headless-4.12.0.88-*.whl',\n        WHL_SHA,\n    )\n\n    if whl is None:\n        for c in sorted(\n            _P('/kaggle/input').rglob(\n                'opencv_python_headless-4.12.0.88-*.whl'\n            )\n        ):\n            if sha(c) == WHL_SHA:\n                whl = c\n                break\n\n    if whl is None:\n        raise RuntimeError(\n            'pinned opencv wheel absent or hash mismatch'\n        )\n\n    envd = _P(\n        '/kaggle/working/_coat_env'\n    )\n\n    _sp.run(\n        [\n            _sy.executable,\n            '-m',\n            'pip',\n            'install',\n            '--no-deps',\n            '--quiet',\n            '--target',\n            str(envd),\n            str(whl),\n        ],\n        check=True,\n    )\n\n    out = _P(\n        '/kaggle/working/_coat_arm.csv'\n    )\n\n    child = (\n        \"import sys, json\\n\"\n        f\"sys.path.insert(0, {str(envd)!r})\\n\"\n        f\"sys.path.insert(0, {str(art)!r})\\n\"\n        \"import cv2; \"\n        \"assert cv2.__version__ == '4.12.0', \"\n        \"cv2.__version__\\n\"\n        \"import torch; \"\n        \"assert torch.cuda.device_count() == 2\\n\"\n        \"import coatnet_resgated_ep10_top3_inference as rt\\n\"\n        \"assert rt.base.cv2.__version__ == '4.12.0'\\n\"\n        \"from pathlib import Path\\n\"\n        \"r = rt.run_submission(\"\n        \"competition_root=rt.base.find_competition_root(),\\n\"\n        f\"    artifact_root=Path({str(art)!r}), \"\n        f\"output_path=Path({str(out)!r}),\\n\"\n        \"    gpu_batch_studies=2, \"\n        \"backbone_micro_images=8)\\n\"\n        \"assert r['status'] == rt.SUBMISSION_STATUS\\n\"\n        \"assert r['models'] == 3\\n\"\n        \"assert [i['epoch'] for i in r['checkpoints']] \"\n        \"== [4, 6, 8]\\n\"\n        \"assert r['fallback_studies'] == 0, r['failures']\\n\"\n        \"Path('/kaggle/working/_coat_arm_receipt.json')\"\n        \".write_text(json.dumps(r, indent=2))\\n\"\n    )\n\n    env = dict(_o.environ)\n\n    env['PYTHONPATH'] = (\n        f\"{envd}:{art}:\"\n        + env.get(\n            'PYTHONPATH',\n            '',\n        )\n    )\n\n    proc = _sp.run(\n        [\n            _sy.executable,\n            '-c',\n            child,\n        ],\n        env=env,\n        capture_output=True,\n        text=True,\n    )\n\n    if proc.returncode != 0:\n        raise RuntimeError(\n            'coat child failed: '\n            f'{proc.stderr[-700:]}'\n        )\n\n    pub = _pd.read_csv(\n        raptor,\n        dtype={\n            'StudyInstanceUID': str\n        },\n    )\n\n    ours = _pd.read_csv(\n        out,\n        dtype={\n            'StudyInstanceUID': str\n        },\n    )\n\n    if (\n        list(ours.columns)\n        != list(pub.columns)\n    ):\n        raise RuntimeError(\n            'coat arm column drift'\n        )\n\n    ours = (\n        ours\n        .set_index('StudyInstanceUID')\n        .reindex(\n            pub[\n                'StudyInstanceUID'\n            ].astype(str).tolist()\n        )\n        .reset_index()\n    )\n\n    lab = [\n        c\n        for c in pub.columns\n        if c != 'StudyInstanceUID'\n    ]\n\n    if (\n        ours[lab]\n        .isna()\n        .any()\n        .any()\n    ):\n        raise RuntimeError(\n            'coat arm does not cover every study'\n        )\n\n    import json as _j\n    import numpy as _np\n\n    private_alpha = (\n        0.40000000000000002\n    )\n\n    public_rank = pub[\n        lab\n    ].rank(\n        method='average',\n        pct=True,\n    )\n\n    private_rank = ours[\n        lab\n    ].rank(\n        method='average',\n        pct=True,\n    )\n\n    hybrid = pub.copy()\n\n    hybrid[lab] = (\n        (\n            (1.0 - private_alpha)\n            * public_rank\n        )\n        + (\n            private_alpha\n            * private_rank\n        )\n    ).rank(\n        method='average',\n        pct=True,\n    )\n\n    if not _np.isfinite(\n        hybrid[\n            lab\n        ].to_numpy(\n            _np.float64\n        )\n    ).all():\n        raise RuntimeError(\n            'CoAt/Raptor hybrid contains '\n            'non-finite values'\n        )\n\n    tmp = raptor.with_name(\n        '.raptor_coat_hybrid.csv'\n    )\n\n    hybrid.to_csv(\n        tmp,\n        index=False,\n    )\n\n    _o.replace(\n        tmp,\n        raptor,\n    )\n\n    raptor.with_name(\n        '_coat_raptor_blend_receipt.json'\n    ).write_text(\n        _j.dumps(\n            {\n                'contract':\n                    'public_raptor_private_residual_'\n                    'coat_global_rank_blend_v1',\n\n                'private_alpha':\n                    private_alpha,\n\n                'public_raptor_alpha':\n                    1.0 - private_alpha,\n\n                'study_count':\n                    len(ours),\n\n                'finding_specific_weights':\n                    False,\n            },\n            indent=2,\n            sort_keys=True,\n        )\n        + '\\n'\n    )\n\n    return len(ours)\n\n\nimport shutil as _ke_shutil\nimport traceback as _ke_traceback\n\n\ndef _ke_profile_status(stage, status, detail):\n    path = _KePath('/kaggle/working/submission_profile_status.csv')\n    header = '' if path.is_file() else 'stage,status,detail\\n'\n    with path.open('a', encoding='utf-8') as handle:\n        handle.write(header + f'{stage},{status},\"{detail}\"\\n')\n\n\n_ke_raptor_path = _KePath('/kaggle/working/_raptor.csv')\n_ke_raptor_public = _KePath('/kaggle/working/_raptor_public_only.csv')\n_ke_shutil.copyfile(_ke_raptor_path, _ke_raptor_public)\n\ntry:\n    _coat_n = _coat_substitute()\n    print(\n        '[coat-arm] blended residual-gated '\n        'e4/e6/e8 into the public Raptor arm '\n        f'(private alpha 0.400; '\n        f'{_coat_n} studies)',\n        flush=True,\n    )\n    _KE_COAT_APPLIED = True\n    _ke_profile_status(\n        'coat_substitute',\n        'applied',\n        f'{_coat_n} studies at private alpha 0.400',\n    )\nexcept Exception as _ke_coat_error:\n    _ke_traceback.print_exc()\n    _ke_shutil.copyfile(_ke_raptor_public, _ke_raptor_path)\n    print(\n        '[coat-arm] SKIPPED '\n        f'({type(_ke_coat_error).__name__}: {_ke_coat_error}); '\n        'the Raptor arm stays at the public four-arm blend, which is what '\n        'the 0.939 parent used',\n        flush=True,\n    )\n    _KE_COAT_APPLIED = False\n    _ke_profile_status(\n        'coat_substitute',\n        'skipped',\n        f'{type(_ke_coat_error).__name__}: {_ke_coat_error}',\n    )\n\n\n# =====================================================================\n# Align Transformer/Rad and strengthened Raptor.\n# =====================================================================\n\n_ke_theirs = _ke_pd.read_csv(\n    '/kaggle/working/_raptor.csv',\n    dtype={\n        'StudyInstanceUID': str\n    },\n)\n\nassert (\n    list(_ke_theirs.columns)\n    == list(_ke_ours.columns)\n), 'column drift'\n\n_ke_theirs = (\n    _ke_theirs\n    .set_index('StudyInstanceUID')\n    .reindex(\n        _ke_ours[\n            'StudyInstanceUID'\n        ].astype(str).tolist()\n    )\n    .reset_index()\n)\n\nassert (\n    _ke_theirs[\n        _KE_LAB\n    ]\n    .notna()\n    .all()\n    .all()\n), 'study identity drift'\n\n\n_ke_tr = _ke_ours[\n    _KE_LAB\n].rank(\n    method='average',\n    pct=True,\n)\n\n_ke_cr = _ke_theirs[\n    _KE_LAB\n].copy()\n\n_blend_transformer = (\n    _ke_ours.copy()\n)\n\n_blend_coatnet = (\n    _ke_theirs.copy()\n)\n\n_blend_labels = list(\n    _KE_LAB\n)\n\n_blend_tr = _ke_tr.copy()\n_blend_cr = _ke_cr.copy()\n\n\n_parent = (\n    _blend_transformer.copy()\n)\n\nfor _label in _blend_labels:\n    _parent[\n        _label\n    ] = (\n        0.40\n        * _blend_tr[\n            _label\n        ]\n        + 0.60\n        * _blend_cr[\n            _label\n        ]\n    )\n\n_parent[\n    _blend_labels\n] = (\n    _parent[\n        _blend_labels\n    ]\n    .rank(\n        method='average',\n        pct=True,\n    )\n)\n\nassert _ke_np.isfinite(\n    _parent[\n        _blend_labels\n    ].to_numpy(\n        _ke_np.float64\n    )\n).all()\n\n_parent_path = _KePath(\n    '/kaggle/working/'\n    'submission_parent_exact.csv'\n)\n\n_parent.to_csv(\n    _parent_path,\n    index=False,\n)\n\n\n\nif not globals().get('V17_PUBLIC_FRONTIER', True):\n    print(\n        '[outer-routing] WARNING: the public DINO frontier was never '\n        'promoted into submission.csv, so the Transformer branch below is '\n        'the native blend -- a lineage that has never been scored. The '\n        'per-target uplifts are applied unchanged; treat this submission as '\n        'a degraded run, not as the tuned recipe.',\n        flush=True,\n    )\n\nif not globals().get('V18_CALIBRATOR_APPLIED', False):\n    print(\n        '[outer-routing] WARNING: the Rad calibrator did not run, so the '\n        'Transformer branch below is not the one these per-target uplifts '\n        'were chosen against. The uplifts are applied unchanged; treat this '\n        'submission as a degraded run, not as the tuned recipe.',\n        flush=True,\n    )\n\n_outer_raptor_weight = {\n    label: 0.60\n    for label\n    in _blend_labels\n}\n\n_outer_raptor_weight.update({\n    'ACL': 0.75,\n    'Medial Meniscus': 0.80,\n    'Lateral Meniscus': 1.00,\n    'Lateral OA': 0.75,\n    'Fracture': 0.75,\n})\n\n\n_expected_outer_weight = {\n    'ACL':\n        0.75,\n\n    'MCL':\n        0.60,\n\n    'Medial Meniscus':\n        0.80,\n\n    'Lateral Meniscus':\n        1.00,\n\n    'Medial OA':\n        0.60,\n\n    'Lateral OA':\n        0.75,\n\n    'PF OA':\n        0.60,\n\n    'Effusion':\n        0.60,\n\n    'Synovitis':\n        0.60,\n\n    \"Baker's\":\n        0.60,\n\n    'Contusion':\n        0.60,\n\n    'Fracture':\n        0.75,\n}\n\nassert (\n    set(_outer_raptor_weight)\n    == set(_expected_outer_weight)\n)\n\nfor _label, _expected in (\n    _expected_outer_weight.items()\n):\n    assert abs(\n        float(\n            _outer_raptor_weight[\n                _label\n            ]\n        )\n        - _expected\n    ) < 1e-12\n\n\nprint(\n    '[outer-routing] '\n    'single-target candidate',\n    flush=True,\n)\n\nprint(\n    '[outer-routing] '\n    'inner Raptor blend remains '\n    'public=0.60 / residual-CoAt=0.40',\n    flush=True,\n)\n\nfor _label in _blend_labels:\n    print(\n        f'  {_label:18s} '\n        f'T='\n        f'{1.0 - _outer_raptor_weight[_label]:.2f} '\n        f'R='\n        f'{_outer_raptor_weight[_label]:.2f}',\n        flush=True,\n    )\n\n\n# =====================================================================\n# Apply the candidate using the exact same final-rank operation as\n# the 0.939 parent.\n# =====================================================================\n\n_candidate = (\n    _blend_transformer.copy()\n)\n\nfor _label in _blend_labels:\n\n    _weight = float(\n        _outer_raptor_weight[\n            _label\n        ]\n    )\n\n    _candidate[\n        _label\n    ] = (\n        (\n            1.0\n            - _weight\n        )\n        * _blend_tr[\n            _label\n        ]\n        + _weight\n        * _blend_cr[\n            _label\n        ]\n    )\n\n\n_candidate[\n    _blend_labels\n] = (\n    _candidate[\n        _blend_labels\n    ]\n    .rank(\n        method='average',\n        pct=True,\n    )\n)\n\n\nassert _ke_np.isfinite(\n    _candidate[\n        _blend_labels\n    ].to_numpy(\n        _ke_np.float64\n    )\n).all()\n\n\n# =====================================================================\n# Hard isolation gate:\n# all 11 untouched targets must be EXACTLY equal to the 0.939 parent.\n# =====================================================================\n\n_changed_targets = [\n    'ACL',\n    'Medial Meniscus',\n    'Lateral Meniscus',\n    'Lateral OA',\n    'Fracture',\n]\n\n_unchanged_targets = [\n    label\n    for label in _blend_labels\n    if label not in _changed_targets\n]\n\n\nfor _label in _unchanged_targets:\n\n    _parent_values = (\n        _parent[\n            _label\n        ].to_numpy(\n            _ke_np.float64\n        )\n    )\n\n    _candidate_values = (\n        _candidate[\n            _label\n        ].to_numpy(\n            _ke_np.float64\n        )\n    )\n\n    if not _ke_np.array_equal(\n        _parent_values,\n        _candidate_values,\n    ):\n        raise RuntimeError(\n            'untouched target drifted: '\n            f'{_label}'\n        )\n\n\n# =====================================================================\n# No fullfit0033 / bag overlay is permitted in this diagnostic.\n# Otherwise Medial Meniscus would change and the result would no longer\n# isolate Lateral Meniscus.\n# =====================================================================\n\n_candidate.to_csv(\n    _ke_primary,\n    index=False,\n)\n\nprint(\n    '[outer-routing] '\n    'no additional model overlay applied',\n    flush=True,\n)\n\n\n# =====================================================================\n# Receipt\n# =====================================================================\n\nimport hashlib as _ke_hashlib\nimport json as _ke_json\n\n\ndef _ke_sha256(_path):\n\n    digest = (\n        _ke_hashlib.sha256()\n    )\n\n    with _KePath(\n        _path\n    ).open(\n        'rb'\n    ) as handle:\n\n        for block in iter(\n            lambda:\n                handle.read(\n                    8 << 20\n                ),\n            b'',\n        ):\n            digest.update(\n                block\n            )\n\n    return digest.hexdigest()\n\n\n_parent_sha = (\n    _ke_sha256(\n        _parent_path\n    )\n)\n\n_output_sha = (\n    _ke_sha256(\n        _ke_primary\n    )\n)\n\n\n_changed_numeric_count = {}\n\nfor _label in _changed_targets:\n\n    _parent_values = (\n        _parent[\n            _label\n        ].to_numpy(\n            _ke_np.float64\n        )\n    )\n\n    _candidate_values = (\n        _candidate[\n            _label\n        ].to_numpy(\n            _ke_np.float64\n        )\n    )\n\n    _changed_numeric_count[\n        _label\n    ] = int(\n        _ke_np.count_nonzero(\n            _parent_values\n            != _candidate_values\n        )\n    )\n\n\n_receipt = {\n    'contract':\n        'public0939_single_target_'\n        'outer_routing_v1',\n\n    'status':\n        'passed',\n\n    'parent':\n        'reproduced_public_0.939',\n\n    'parent_sha256':\n        _parent_sha,\n\n    'output_sha256':\n        _output_sha,\n\n    'study_count':\n        int(\n            len(_candidate)\n        ),\n\n    'inner_public_raptor_alpha':\n        0.60,\n\n    'inner_private_coat_alpha':\n        0.40,\n\n    'outer_raptor_weight_by_target':\n        {\n            label:\n                float(\n                    _outer_raptor_weight[\n                        label\n                    ]\n                )\n            for label\n            in _blend_labels\n        },\n\n    'changed_targets':\n        list(\n            _changed_targets\n        ),\n\n    'unchanged_targets':\n        list(\n            _unchanged_targets\n        ),\n\n    'changed_numeric_count_by_target':\n        _changed_numeric_count,\n\n    'unchanged_targets_exact_parent_match':\n        True,\n\n    'additional_model_overlay':\n        False,\n\n    'final_rank_after_outer_blend':\n        True,\n\n    'raptor_arms':\n        [\n            {\n                'name':\n                    arm[\n                        'name'\n                    ],\n\n                'weight':\n                    float(\n                        arm[\n                            'w'\n                        ]\n                    ),\n\n                'windows':\n                    int(\n                        arm[\n                            'k_eval'\n                        ]\n                    ),\n            }\n            for arm\n            in _KE_NS[\n                'ARMS'\n            ]\n        ],\n}\n\n\n_receipt_path = _KePath(\n    '/kaggle/working/'\n    'outer_routing_receipt.json'\n)\n\n_receipt_path.write_text(\n    _ke_json.dumps(\n        _receipt,\n        indent=2,\n        sort_keys=True,\n    )\n    + '\\n',\n    encoding='utf-8',\n)\n\n\n# Keep the existing historical receipt filename available in case a later\n# notebook cell expects it.  The payload itself reflects the current run.\n_KePath(\n    '/kaggle/working/'\n    'repro937_fallback_receipt.json'\n).write_text(\n    _ke_json.dumps(\n        _receipt,\n        indent=2,\n        sort_keys=True,\n    )\n    + '\\n',\n    encoding='utf-8',\n)\n\n\n_ke_ours = _ke_pd.read_csv(\n    _ke_primary,\n    dtype={\n        'StudyInstanceUID':\n            str\n    },\n)\n\n\nassert (\n    _ke_ours[\n        'StudyInstanceUID'\n    ].astype(str).tolist()\n    ==\n    _candidate[\n        'StudyInstanceUID'\n    ].astype(str).tolist()\n)\n\nassert _ke_np.isfinite(\n    _ke_ours[\n        _blend_labels\n    ].to_numpy(\n        _ke_np.float64\n    )\n).all()\n\n\nprint(\n    '[outer-routing] PASS',\n    flush=True,\n)\n\nprint(\n    '[outer-routing] parent sha256:',\n    _parent_sha,\n    flush=True,\n)\n\nprint(\n    '[outer-routing] candidate sha256:',\n    _output_sha,\n    flush=True,\n)\n\nprint(\n    '[outer-routing] changed row counts:',\n    _changed_numeric_count,\n    flush=True,\n)\n\nprint(\n    '[outer-routing] submission:',\n    _ke_primary,\n    flush=True,\n)\n\nprint(\n    '[outer-routing] receipt:',\n    _receipt_path,\n    flush=True,\n)"},{"cell_type":"code","execution_count":null,"id":"4c414beb","metadata":{"papermill":{"duration":0.015281,"end_time":"2026-09-12T11:21:37.880181+00:00","exception":false,"start_time":"2026-09-12T11:21:37.8649+00:00","status":"completed"},"tags":[]},"outputs":[],"source":""}],"metadata":{"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":154281},{"sourceType":"datasetVersion","sourceId":18229736},{"sourceType":"datasetVersion","sourceId":18673450},{"sourceType":"datasetVersion","sourceId":18716507},{"sourceType":"datasetVersion","sourceId":18839182},{"sourceType":"datasetVersion","sourceId":18842180},{"sourceType":"datasetVersion","sourceId":18879001},{"sourceType":"datasetVersion","sourceId":18956429},{"sourceType":"datasetVersion","sourceId":19120128},{"sourceType":"datasetVersion","sourceId":19122845},{"sourceType":"datasetVersion","sourceId":19134209},{"sourceType":"datasetVersion","sourceId":19395055},{"sourceType":"datasetVersion","sourceId":19395706},{"sourceType":"kernelVersion","sourceId":342671664},{"sourceType":"kernelVersion","sourceId":342849430},{"sourceType":"modelInstanceVersion","sourceId":4533},{"sourceType":"modelInstanceVersion","sourceId":4534}],"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.12.13"},"papermill":{"default_parameters":{},"duration":246.93066,"end_time":"2026-09-12T11:21:40.961632+00:00","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2026-09-12T11:17:34.030972+00:00","version":"2.7.0"},"widgets":{"application/vnd.jupyter.widget-state+json":{"state":{"001fff04ef0b4aeca1e0881dc9f5d158":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_6eca7796d91b4cec880eb8b2bc922c0c","placeholder":"​","style":"IPY_MODEL_92824531f9454214a7a62fe558c82923","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1284.56it/s, Materializing param=layernorm.weight]"}},"021b964be1ba418a9a47f64794135a62":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"0332432255534050b38a29f4b7b53d13":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"040d2085d2cf48c08a46a86f22d18fec":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"0588610082c54262846887ed5659274b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_349638bda0294701a30b737e086ab4d6","placeholder":"​","style":"IPY_MODEL_f718a4c440e64ccf9be71da6d147b2e6","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"0831b6e1adb04846b60e0fea68a426d9":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"0841d9ef77c94c6480de627a6cec9351":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"091972e206b14a5ba3bd6910ab949a98":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"092981518c634ef2a4c9e7b218a73653":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"098e401e767f44fa9e9b11889fc871b4":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"09a4a813987c4229891c90c878e28ed9":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_58e278f9725244a2875f213c72d3fcd2","IPY_MODEL_32e63974452140a7abad1ba34d174a26","IPY_MODEL_374d9d82e2d544548a2401061eb5af3c"],"layout":"IPY_MODEL_021b964be1ba418a9a47f64794135a62","tabbable":null,"tooltip":null}},"0ade1e9cbce24cea94e9f3c1c07a4de6":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_2e21a54039044d8cacbc6d031eeacc47","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_dc3c9d7d49964fbb87dddedef1f7397a","tabbable":null,"tooltip":null,"value":223.0}},"0c0cec13999642beafd0d2789bff4ae7":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"0c7a1681e8ae4ee2b60ea600ac959519":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_e8bc51cc8c274a089e4535a720b78785","placeholder":"​","style":"IPY_MODEL_f321d5b9749045bbaaf9938c4a781b82","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"0fc971f2c7794314aeb78714fc0bcbe3":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"1298d85abde143c1a0dd6ed2d324832d":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"138b2522d92e429f8ec6bf81dd4d1a4a":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_ec5eb5404778483c92e477f36680767a","placeholder":"​","style":"IPY_MODEL_091972e206b14a5ba3bd6910ab949a98","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 908.39it/s, Materializing param=layernorm.weight]"}},"13bb3e46566544909d5204ed782e0154":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_d69d95f50fd54e8bbd469fcca3d7e200","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_761c5f8cc3334119a60417a92913f4c9","tabbable":null,"tooltip":null,"value":223.0}},"151df2304ae844d6b4dc4b6a6b74f4c5":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"1686d6ee41cd44eb99411c537b478d39":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_8b2336f5d84e45ecbc2b83b84f93ed08","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_89043720d0334fa7975b5ee130b01454","tabbable":null,"tooltip":null,"value":223.0}},"17ec57e7ac2947d4b121380d87639ec3":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_add23e169d8242f985ba41e39075629d","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_0332432255534050b38a29f4b7b53d13","tabbable":null,"tooltip":null,"value":223.0}},"18d9ee9c0efa4e5c9af6094bc62775f0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_5961ee102fa84811a1ee095c8a1d76f2","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_151df2304ae844d6b4dc4b6a6b74f4c5","tabbable":null,"tooltip":null,"value":223.0}},"1aa7a52fd4e340cc96027dc532d2a03f":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"1c3108ec690d43afae5f80c3aa71552e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_0588610082c54262846887ed5659274b","IPY_MODEL_13bb3e46566544909d5204ed782e0154","IPY_MODEL_138b2522d92e429f8ec6bf81dd4d1a4a"],"layout":"IPY_MODEL_ea26272ea17c4380b94ccc38a2ae1069","tabbable":null,"tooltip":null}},"1c8a14c7a25f41bda64541033f70f6e1":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"1e2675d174ce48dfaa78abcbbc0f9378":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"1e850974fa1b410b9804b3bb8edd2f82":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_cd449b1c97884f0b90b598651d90f014","placeholder":"​","style":"IPY_MODEL_af7cbdc0aa754b169f79a7bae1c8ce62","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1286.86it/s, Materializing param=layernorm.weight]"}},"21747e1cc39b47b4853d309a99f25560":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"22558d5a2cdf4741ae7910c56e0f6a43":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2291dd898e83481f9ea1b61825c06e73":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_bdc9e49c7f1f44da9e5c27ced54c36aa","placeholder":"​","style":"IPY_MODEL_9ac55c8def5f4bd69e58e08f2e7bd92d","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 963.75it/s, Materializing param=layernorm.weight]"}},"256151ac1bf5415283111e702e978a34":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_943fc546c5c64b19949373e9fa04662f","placeholder":"​","style":"IPY_MODEL_0c0cec13999642beafd0d2789bff4ae7","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"2785e944c2584e31bb536a76878c7057":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"28349efca0a94ef0989c5b6008c07199":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_8af46bfb50ea4968926362ac1dba9ca4","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_6980ca3420494bd3b25631851f7c569c","tabbable":null,"tooltip":null,"value":223.0}},"28f307fbe34946b0999141c7bd158ee6":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_43f0d6d0e1b348c68027a7b4f6a8dda0","IPY_MODEL_9fa85b6aeb3a409785cf2e16fc8dbd5b","IPY_MODEL_6021d8f3e4f74d4e9f25e532fb0e20e0"],"layout":"IPY_MODEL_1c8a14c7a25f41bda64541033f70f6e1","tabbable":null,"tooltip":null}},"29b21e904c3b4e35bbec3b26250aba50":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"29d41bf819fe4bdd971da53c510b20b6":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2a010c0acd48410fa279b7fe12e169f7":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2ae160ae9ee84b039df30e33e0319e34":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2bcdd04d144a416f887caca2180de49c":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2ce7f3e18e864b5185adafe028eab54e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"2da739ba9a484e32a89c03a88ac13810":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"2e21a54039044d8cacbc6d031eeacc47":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"305d0a85aa1241a5997787ae9118f4bb":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_6a2d6a3a3d5649d389122fef42a2e6ab","IPY_MODEL_28349efca0a94ef0989c5b6008c07199","IPY_MODEL_7846d4b82a894aae93ce8f95bfcba349"],"layout":"IPY_MODEL_b2e372f325c34833824fbd68e5d3c39f","tabbable":null,"tooltip":null}},"309f5da7413146da87715b0349724e69":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_be2d1c37b52e46f1b151dee69b12c237","placeholder":"​","style":"IPY_MODEL_fd5e977919d64f5bbc252fe662aa9b02","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"310f80a1c88b46f099087377b81bc182":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"3155264e1ace45c1858785f1d8c403a2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_ae642053891447528ebf5041c773ca79","placeholder":"​","style":"IPY_MODEL_5d38bced11f74c39b0489df9b0809773","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1001.54it/s, Materializing param=layernorm.weight]"}},"32e63974452140a7abad1ba34d174a26":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_ef4e6cf4709246daa200675f85ce9c92","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_6ac6c178d6d44f1caea2298e4a8cb960","tabbable":null,"tooltip":null,"value":223.0}},"349638bda0294701a30b737e086ab4d6":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"34a5d06784d64f448f4115ce9eb7d220":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"374d9d82e2d544548a2401061eb5af3c":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_0831b6e1adb04846b60e0fea68a426d9","placeholder":"​","style":"IPY_MODEL_9b55a5a7dd8a40afae193e2fa44a80a3","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 973.19it/s, Materializing param=layernorm.weight]"}},"37e53c7f3dfb4e90b4c3a908f82807d9":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_4800904213e9448a812f16da4928d81d","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_9f1782e38d984fc8bd839e461ecd28e9","tabbable":null,"tooltip":null,"value":223.0}},"3e297e3625224bd3bb3a5e77bdffb009":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_0841d9ef77c94c6480de627a6cec9351","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_ea73d19716ad45e5a06c6acae95b85a5","tabbable":null,"tooltip":null,"value":223.0}},"403be9b1e90f46eb8a96c0528816d245":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"4346dac133e449f49eb2c450e0ddd37d":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"43e276c5b388420eac82f75ed70afcf5":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"43f0d6d0e1b348c68027a7b4f6a8dda0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_092981518c634ef2a4c9e7b218a73653","placeholder":"​","style":"IPY_MODEL_72d139fddb9e4d3a9063cb57b57b6844","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"46196c84467845af93b04bccc8feebc5":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"4800904213e9448a812f16da4928d81d":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"4c42733c1097455e98efb79c6d44ed1f":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"4f8ec6ae4d1e43988390bf20e90d30aa":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_a7ba79b7dc6843fca9b7922475af793b","placeholder":"​","style":"IPY_MODEL_f8a22973149a41398a9bd8f9c9357b69","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1062.38it/s, Materializing param=layernorm.weight]"}},"4fe602dd8ad44f8e90b2ae1620716f21":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"505c6aa30cf54733b16e931517a71f45":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"50eccf05cca046fe827d0a86a1293620":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_fea4f707e4164bfca00efa14e44a68cb","placeholder":"​","style":"IPY_MODEL_efd3ee95127f4906901cf33c850bdcf0","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"51b40f5d39984b2389d02e6befddfe50":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_505c6aa30cf54733b16e931517a71f45","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_7679cc18ff40413b93c015ac38046e16","tabbable":null,"tooltip":null,"value":223.0}},"531f18d7290b4aca9bbb3f8fc721c3d5":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"541b75945dba47f48656b1199455e1ac":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_8cc45b69f32b46d8a67d16141c11f6c8","IPY_MODEL_be1be741ea654bcb80367028d0ed6ed8","IPY_MODEL_fcd3567a70c644f98447c4af41d28342"],"layout":"IPY_MODEL_7ebd6c6cf1ae4aa3b8cd56060dfbe2af","tabbable":null,"tooltip":null}},"548a1ddd0df24a398d8f1fdb042492bb":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_1aa7a52fd4e340cc96027dc532d2a03f","placeholder":"​","style":"IPY_MODEL_5fec30be73014312a7a83d29a723468d","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"54e3e2cfa4d84eddb08cc1a0e5038547":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_a55452371a454b3bb5883776151d40ec","placeholder":"​","style":"IPY_MODEL_9cbdd01546244030ba270b960ff7fb1e","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"557705e640074dde83cd6a49bd33ad9b":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"55b6918cfe3849d59d454e993934cce9":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"57c42f2b78f04c84b85662ae4d4d6247":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_21747e1cc39b47b4853d309a99f25560","placeholder":"​","style":"IPY_MODEL_34a5d06784d64f448f4115ce9eb7d220","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"58e278f9725244a2875f213c72d3fcd2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_d4f4758f9e284b698b1bd40e2f86621a","placeholder":"​","style":"IPY_MODEL_ebf8767862aa474aa336e05653dca6da","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"5961ee102fa84811a1ee095c8a1d76f2":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"5a2cc19651744f9891adbd43ca85a287":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"5a3c2b02e9454ea2bbbbd6047e7a617c":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"5d38bced11f74c39b0489df9b0809773":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"5d69e617eae3487ea707d0be133ac20b":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"5def2e8e7b62402982c4130e6c3509ff":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"5f8462772b3544e2a7e973e14bf2b93c":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"5fec30be73014312a7a83d29a723468d":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"6021d8f3e4f74d4e9f25e532fb0e20e0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_64aed4df401c44a7972862051612ecf5","placeholder":"​","style":"IPY_MODEL_55b6918cfe3849d59d454e993934cce9","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1228.61it/s, Materializing param=layernorm.weight]"}},"618a7f9417e548e79fa1484e9181890a":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"61aa42a275be49ba9207495e8dd0e85a":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"62820c59b4ca44b2a31e1b5bf4aa9bc2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_77cc109869b345b8b0925bf4212cbeb4","placeholder":"​","style":"IPY_MODEL_842036720bde401698f6b81ffcaab2f2","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 968.23it/s, Materializing param=layernorm.weight]"}},"62ca2677569b4974bf9ce788e6aa4d68":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"64aed4df401c44a7972862051612ecf5":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"64f1ba55256f4d4580f78b694e75620d":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_bbeda6af492e4047b452eabce863e696","placeholder":"​","style":"IPY_MODEL_29b21e904c3b4e35bbec3b26250aba50","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 981.32it/s, Materializing param=layernorm.weight]"}},"660524c7e4f74c47b236a0f8d4e7fbc6":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_71c12e8c727544559189b8a5964ac1f2","IPY_MODEL_813606be114f427cae9b925c57eee5da","IPY_MODEL_97b472ea2a2e463ebef51148b7650196"],"layout":"IPY_MODEL_2ae160ae9ee84b039df30e33e0319e34","tabbable":null,"tooltip":null}},"678cdd2685d347e8aa9af6bb926c2f56":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_4346dac133e449f49eb2c450e0ddd37d","placeholder":"​","style":"IPY_MODEL_9cef162d4605405b8ea710d964466dd8","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1141.00it/s, Materializing param=layernorm.weight]"}},"68324bf8facb432ea17a24dffaa3fc9c":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"693951d2f70a40cbb7e9c8e5ae427a21":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_2a010c0acd48410fa279b7fe12e169f7","placeholder":"​","style":"IPY_MODEL_5a3c2b02e9454ea2bbbbd6047e7a617c","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1268.20it/s, Materializing param=layernorm.weight]"}},"6980ca3420494bd3b25631851f7c569c":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"6a2d6a3a3d5649d389122fef42a2e6ab":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_ced71d68de2444d892c7e6829d6a60ef","placeholder":"​","style":"IPY_MODEL_a5403ad89aa34a36a80ec2c20d559993","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"6a539d5b98bd416d8b8b090334b016db":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"6ac6c178d6d44f1caea2298e4a8cb960":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"6eca7796d91b4cec880eb8b2bc922c0c":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"70ddff67388646008a9e022ba1899ed2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"71129c0b1dce457bac4321b2e6191c5c":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"711ee5a71c154420b4fa3005ba3301b3":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"71c12e8c727544559189b8a5964ac1f2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_557705e640074dde83cd6a49bd33ad9b","placeholder":"​","style":"IPY_MODEL_e7411d9f535f45448e26e6a649c25992","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"72a2b75ee970443bb906c8f40c0d52ba":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_9e13f21ff90b4cba864abf95529333f7","placeholder":"​","style":"IPY_MODEL_d9b28fe3eb5d4dd2a8d232474b521b3b","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 992.62it/s, Materializing param=layernorm.weight]"}},"72d139fddb9e4d3a9063cb57b57b6844":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"761c5f8cc3334119a60417a92913f4c9":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"7679cc18ff40413b93c015ac38046e16":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"76c66d0f6fc5459c8cc4b413c496aa24":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"77cc109869b345b8b0925bf4212cbeb4":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"780a0b9dbfa24bc490bde118e71c77d2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_e52f77db6ad848e1988bff4d5d25c1ab","IPY_MODEL_1686d6ee41cd44eb99411c537b478d39","IPY_MODEL_ba069f02a11d44dbaeea14b904c4eac7"],"layout":"IPY_MODEL_62ca2677569b4974bf9ce788e6aa4d68","tabbable":null,"tooltip":null}},"7846d4b82a894aae93ce8f95bfcba349":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_f32f9e8082aa40008dec513e83b8fbf3","placeholder":"​","style":"IPY_MODEL_2ce7f3e18e864b5185adafe028eab54e","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 933.62it/s, Materializing param=layernorm.weight]"}},"78eb623b1579459aa9eae6f71ef03ec0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"7abc345feecd41cbaf0c9248c6a9ea2b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_85a9c59df810426f9a201581f7e9843d","IPY_MODEL_0ade1e9cbce24cea94e9f3c1c07a4de6","IPY_MODEL_4f8ec6ae4d1e43988390bf20e90d30aa"],"layout":"IPY_MODEL_d56acb8d12c94b87b8901401700a6779","tabbable":null,"tooltip":null}},"7af23bc3c780400bbfc18fe8f6299ffc":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_cd6c0ddd97454e77b2b8da84a82797e0","IPY_MODEL_51b40f5d39984b2389d02e6befddfe50","IPY_MODEL_bd0277edcd444ce38f2a39186a2abefe"],"layout":"IPY_MODEL_c3651fa1886d4e15815d8a72218812de","tabbable":null,"tooltip":null}},"7b6242f326064668bace644252fdc2a9":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"7ebd6c6cf1ae4aa3b8cd56060dfbe2af":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"813606be114f427cae9b925c57eee5da":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_531f18d7290b4aca9bbb3f8fc721c3d5","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_8e99359e85e44b1896d8c4c176e9f377","tabbable":null,"tooltip":null,"value":223.0}},"81553d1f82eb47cb9f086b1358e58376":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"8399eadcfc614a8a8d0051dff6455c69":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_0c7a1681e8ae4ee2b60ea600ac959519","IPY_MODEL_fd94c34f86d74a88a459dbcaf3f9a241","IPY_MODEL_2291dd898e83481f9ea1b61825c06e73"],"layout":"IPY_MODEL_098e401e767f44fa9e9b11889fc871b4","tabbable":null,"tooltip":null}},"84035e25e03a439795b21e3289ca1a85":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_98151093ca914f16b08f37ab26ba89c2","IPY_MODEL_9a9081eb326a459d8198bdf222f7150e","IPY_MODEL_64f1ba55256f4d4580f78b694e75620d"],"layout":"IPY_MODEL_2da739ba9a484e32a89c03a88ac13810","tabbable":null,"tooltip":null}},"842036720bde401698f6b81ffcaab2f2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"8486112545bb4f4eb76a5e623ca63083":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"85a9c59df810426f9a201581f7e9843d":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_8486112545bb4f4eb76a5e623ca63083","placeholder":"​","style":"IPY_MODEL_4fe602dd8ad44f8e90b2ae1620716f21","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"89043720d0334fa7975b5ee130b01454":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"8aacf8a3450b4a65af3d3e7ff9756c8b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_d2bc1d71c5004b9d95792f7ece6dddd0","placeholder":"​","style":"IPY_MODEL_1298d85abde143c1a0dd6ed2d324832d","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"8ae169023e0d482b9dec1d8bd58e3488":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_ba236abbadf04693b93768fb684cc7b5","IPY_MODEL_c0c778351fa841e5934ee61c30635a53","IPY_MODEL_c97c030331024ceda87c707e5d28898e"],"layout":"IPY_MODEL_99fc8d9576b94416a359fd28c36985cc","tabbable":null,"tooltip":null}},"8af46bfb50ea4968926362ac1dba9ca4":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"8b2336f5d84e45ecbc2b83b84f93ed08":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"8c0cb1bab950459a8a973642dd05776f":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"8cb5d2615668488fa6aa9233776fe16c":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_256151ac1bf5415283111e702e978a34","IPY_MODEL_96e386e89d0d4fa3bf7dd6bfc9470408","IPY_MODEL_62820c59b4ca44b2a31e1b5bf4aa9bc2"],"layout":"IPY_MODEL_5d69e617eae3487ea707d0be133ac20b","tabbable":null,"tooltip":null}},"8cc45b69f32b46d8a67d16141c11f6c8":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_61aa42a275be49ba9207495e8dd0e85a","placeholder":"​","style":"IPY_MODEL_78eb623b1579459aa9eae6f71ef03ec0","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"8e99359e85e44b1896d8c4c176e9f377":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"8ed236e17b7e4ec29e5ab5e84fada72f":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_f1de286111404978b175ca10e1c5b9f0","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_46196c84467845af93b04bccc8feebc5","tabbable":null,"tooltip":null,"value":223.0}},"92824531f9454214a7a62fe558c82923":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"9327b587e16d4ec9931c6617ae8db91b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"93f5294211a14c35853fd94b7eb5c5e3":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"943fc546c5c64b19949373e9fa04662f":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"94c435c5e5ec4b24b8ea80ac0337a8a5":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"9590a7bfd6b9466d8c23166ef477db9c":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"96e386e89d0d4fa3bf7dd6bfc9470408":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_76c66d0f6fc5459c8cc4b413c496aa24","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_7b6242f326064668bace644252fdc2a9","tabbable":null,"tooltip":null,"value":223.0}},"97b472ea2a2e463ebef51148b7650196":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_2785e944c2584e31bb536a76878c7057","placeholder":"​","style":"IPY_MODEL_403be9b1e90f46eb8a96c0528816d245","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 918.77it/s, Materializing param=layernorm.weight]"}},"98151093ca914f16b08f37ab26ba89c2":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_5def2e8e7b62402982c4130e6c3509ff","placeholder":"​","style":"IPY_MODEL_9590a7bfd6b9466d8c23166ef477db9c","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"99fc8d9576b94416a359fd28c36985cc":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"9a9081eb326a459d8198bdf222f7150e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_29d41bf819fe4bdd971da53c510b20b6","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_70ddff67388646008a9e022ba1899ed2","tabbable":null,"tooltip":null,"value":223.0}},"9ac55c8def5f4bd69e58e08f2e7bd92d":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"9b220ff83da24150b59dff98a3206a7e":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"9b55a5a7dd8a40afae193e2fa44a80a3":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"9cbdd01546244030ba270b960ff7fb1e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"9cef162d4605405b8ea710d964466dd8":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"9e13f21ff90b4cba864abf95529333f7":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"9f1782e38d984fc8bd839e461ecd28e9":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"9fa85b6aeb3a409785cf2e16fc8dbd5b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_fec0e595f7564fee9bb3c9c02f7b78c9","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_81553d1f82eb47cb9f086b1358e58376","tabbable":null,"tooltip":null,"value":223.0}},"a5403ad89aa34a36a80ec2c20d559993":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"a55452371a454b3bb5883776151d40ec":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"a730062df56441309c7ac6c88f7af85f":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"a7ba79b7dc6843fca9b7922475af793b":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"a9adbb758d2e4410b4f0d4562b8860c3":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_309f5da7413146da87715b0349724e69","IPY_MODEL_18d9ee9c0efa4e5c9af6094bc62775f0","IPY_MODEL_d6658df4e455409298df41ef7bc9f3e7"],"layout":"IPY_MODEL_0fc971f2c7794314aeb78714fc0bcbe3","tabbable":null,"tooltip":null}},"abf4f4d1e34447978dab549a48c65acc":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"add23e169d8242f985ba41e39075629d":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"ade55864b924486bb4d522559967370a":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"ae642053891447528ebf5041c773ca79":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"af7cbdc0aa754b169f79a7bae1c8ce62":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"b24204f51e834557b35c0d95462c3764":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"b2e372f325c34833824fbd68e5d3c39f":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"ba069f02a11d44dbaeea14b904c4eac7":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_e68935e92857449b9db367bdd7be97e1","placeholder":"​","style":"IPY_MODEL_bb017a3318f34cd09157da31f4f8c421","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 948.19it/s, Materializing param=layernorm.weight]"}},"ba236abbadf04693b93768fb684cc7b5":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_71129c0b1dce457bac4321b2e6191c5c","placeholder":"​","style":"IPY_MODEL_d44352aae4a947758e9096c869322ae1","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"bb017a3318f34cd09157da31f4f8c421":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"bbeda6af492e4047b452eabce863e696":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"bd0277edcd444ce38f2a39186a2abefe":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_e431817b54d143218f26eb7a5ca570c8","placeholder":"​","style":"IPY_MODEL_d63347b9aa8845ceb6bffdf2628b534f","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 952.08it/s, Materializing param=layernorm.weight]"}},"bdc9e49c7f1f44da9e5c27ced54c36aa":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"be1be741ea654bcb80367028d0ed6ed8":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_711ee5a71c154420b4fa3005ba3301b3","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_e319df5b41064aae9e820188d4387652","tabbable":null,"tooltip":null,"value":223.0}},"be2d1c37b52e46f1b151dee69b12c237":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"c0c778351fa841e5934ee61c30635a53":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_5f8462772b3544e2a7e973e14bf2b93c","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_b24204f51e834557b35c0d95462c3764","tabbable":null,"tooltip":null,"value":223.0}},"c3651fa1886d4e15815d8a72218812de":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"c5334ed811f34f3c88acf3c96bef2187":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"c97c030331024ceda87c707e5d28898e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_618a7f9417e548e79fa1484e9181890a","placeholder":"​","style":"IPY_MODEL_a730062df56441309c7ac6c88f7af85f","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1182.01it/s, Materializing param=layernorm.weight]"}},"ca9aaac986e944da8c48c2c4a18e9cbb":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_8aacf8a3450b4a65af3d3e7ff9756c8b","IPY_MODEL_37e53c7f3dfb4e90b4c3a908f82807d9","IPY_MODEL_678cdd2685d347e8aa9af6bb926c2f56"],"layout":"IPY_MODEL_d7077361dd6b458487e4df770f8d117d","tabbable":null,"tooltip":null}},"cd449b1c97884f0b90b598651d90f014":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"cd6c0ddd97454e77b2b8da84a82797e0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_e8b3b412c348406a88ac0c9fa6adf9ab","placeholder":"​","style":"IPY_MODEL_310f80a1c88b46f099087377b81bc182","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"ceaffd17e37642fc911a8164794c7c0e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_8c0cb1bab950459a8a973642dd05776f","placeholder":"​","style":"IPY_MODEL_43e276c5b388420eac82f75ed70afcf5","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"ced71d68de2444d892c7e6829d6a60ef":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"d2bc1d71c5004b9d95792f7ece6dddd0":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"d44352aae4a947758e9096c869322ae1":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"d4f4758f9e284b698b1bd40e2f86621a":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"d56acb8d12c94b87b8901401700a6779":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"d63347b9aa8845ceb6bffdf2628b534f":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"d6658df4e455409298df41ef7bc9f3e7":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_9b220ff83da24150b59dff98a3206a7e","placeholder":"​","style":"IPY_MODEL_ade55864b924486bb4d522559967370a","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 1001.11it/s, Materializing param=layernorm.weight]"}},"d69d95f50fd54e8bbd469fcca3d7e200":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"d7077361dd6b458487e4df770f8d117d":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"d899132d1ee540d88d86456122e8e2d3":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_57c42f2b78f04c84b85662ae4d4d6247","IPY_MODEL_f7226154487a4331afec4065d84a911f","IPY_MODEL_72a2b75ee970443bb906c8f40c0d52ba"],"layout":"IPY_MODEL_abf4f4d1e34447978dab549a48c65acc","tabbable":null,"tooltip":null}},"d9b28fe3eb5d4dd2a8d232474b521b3b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"da960caf4ce04f2fb35fb69ddf84d4e0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_ceaffd17e37642fc911a8164794c7c0e","IPY_MODEL_17ec57e7ac2947d4b121380d87639ec3","IPY_MODEL_693951d2f70a40cbb7e9c8e5ae427a21"],"layout":"IPY_MODEL_040d2085d2cf48c08a46a86f22d18fec","tabbable":null,"tooltip":null}},"dc3c9d7d49964fbb87dddedef1f7397a":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"e319df5b41064aae9e820188d4387652":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"e431817b54d143218f26eb7a5ca570c8":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"e52f77db6ad848e1988bff4d5d25c1ab":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_e94e6a6ae8bd4fb398ded5f14716091d","placeholder":"​","style":"IPY_MODEL_68324bf8facb432ea17a24dffaa3fc9c","tabbable":null,"tooltip":null,"value":"Loading weights: 100%"}},"e68935e92857449b9db367bdd7be97e1":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"e7411d9f535f45448e26e6a649c25992":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"e767837673f74cd69b55f34749347a4e":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"e8b3b412c348406a88ac0c9fa6adf9ab":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"e8bc51cc8c274a089e4535a720b78785":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"e94e6a6ae8bd4fb398ded5f14716091d":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"ea26272ea17c4380b94ccc38a2ae1069":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"ea73d19716ad45e5a06c6acae95b85a5":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"ProgressStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"ProgressStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","bar_color":null,"description_width":""}},"ebf8767862aa474aa336e05653dca6da":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"ec5eb5404778483c92e477f36680767a":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"ef4cd1fda2214f40bf642064e91d2b8e":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_54e3e2cfa4d84eddb08cc1a0e5038547","IPY_MODEL_f17ca556d58147de8908cd7516f8ebc0","IPY_MODEL_001fff04ef0b4aeca1e0881dc9f5d158"],"layout":"IPY_MODEL_6a539d5b98bd416d8b8b090334b016db","tabbable":null,"tooltip":null}},"ef4e6cf4709246daa200675f85ce9c92":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"efd3ee95127f4906901cf33c850bdcf0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"f17ca556d58147de8908cd7516f8ebc0":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_1e2675d174ce48dfaa78abcbbc0f9378","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_93f5294211a14c35853fd94b7eb5c5e3","tabbable":null,"tooltip":null,"value":223.0}},"f1de286111404978b175ca10e1c5b9f0":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"f2bf727f12d8468a993beb8692172f5b":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_50eccf05cca046fe827d0a86a1293620","IPY_MODEL_8ed236e17b7e4ec29e5ab5e84fada72f","IPY_MODEL_3155264e1ace45c1858785f1d8c403a2"],"layout":"IPY_MODEL_22558d5a2cdf4741ae7910c56e0f6a43","tabbable":null,"tooltip":null}},"f321d5b9749045bbaaf9938c4a781b82":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"f32f9e8082aa40008dec513e83b8fbf3":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"f405473105aa4fbe9bce831512096f51":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HBoxModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HBoxModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HBoxView","box_style":"","children":["IPY_MODEL_548a1ddd0df24a398d8f1fdb042492bb","IPY_MODEL_3e297e3625224bd3bb3a5e77bdffb009","IPY_MODEL_1e850974fa1b410b9804b3bb8edd2f82"],"layout":"IPY_MODEL_2bcdd04d144a416f887caca2180de49c","tabbable":null,"tooltip":null}},"f718a4c440e64ccf9be71da6d147b2e6":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"f7226154487a4331afec4065d84a911f":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_94c435c5e5ec4b24b8ea80ac0337a8a5","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_9327b587e16d4ec9931c6617ae8db91b","tabbable":null,"tooltip":null,"value":223.0}},"f8a22973149a41398a9bd8f9c9357b69":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"fcd3567a70c644f98447c4af41d28342":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"HTMLView","description":"","description_allow_html":false,"layout":"IPY_MODEL_c5334ed811f34f3c88acf3c96bef2187","placeholder":"​","style":"IPY_MODEL_5a2cc19651744f9891adbd43ca85a287","tabbable":null,"tooltip":null,"value":" 223/223 [00:00&lt;00:00, 899.64it/s, Materializing param=layernorm.weight]"}},"fd5e977919d64f5bbc252fe662aa9b02":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"HTMLStyleModel","state":{"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"HTMLStyleModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"StyleView","background":null,"description_width":"","font_size":null,"text_color":null}},"fd94c34f86d74a88a459dbcaf3f9a241":{"model_module":"@jupyter-widgets/controls","model_module_version":"2.0.0","model_name":"FloatProgressModel","state":{"_dom_classes":[],"_model_module":"@jupyter-widgets/controls","_model_module_version":"2.0.0","_model_name":"FloatProgressModel","_view_count":null,"_view_module":"@jupyter-widgets/controls","_view_module_version":"2.0.0","_view_name":"ProgressView","bar_style":"success","description":"","description_allow_html":false,"layout":"IPY_MODEL_e767837673f74cd69b55f34749347a4e","max":223.0,"min":0.0,"orientation":"horizontal","style":"IPY_MODEL_4c42733c1097455e98efb79c6d44ed1f","tabbable":null,"tooltip":null,"value":223.0}},"fea4f707e4164bfca00efa14e44a68cb":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}},"fec0e595f7564fee9bb3c9c02f7b78c9":{"model_module":"@jupyter-widgets/base","model_module_version":"2.0.0","model_name":"LayoutModel","state":{"_model_module":"@jupyter-widgets/base","_model_module_version":"2.0.0","_model_name":"LayoutModel","_view_count":null,"_view_module":"@jupyter-widgets/base","_view_module_version":"2.0.0","_view_name":"LayoutView","align_content":null,"align_items":null,"align_self":null,"border_bottom":null,"border_left":null,"border_right":null,"border_top":null,"bottom":null,"display":null,"flex":null,"flex_flow":null,"grid_area":null,"grid_auto_columns":null,"grid_auto_flow":null,"grid_auto_rows":null,"grid_column":null,"grid_gap":null,"grid_row":null,"grid_template_areas":null,"grid_template_columns":null,"grid_template_rows":null,"height":null,"justify_content":null,"justify_items":null,"left":null,"margin":null,"max_height":null,"max_width":null,"min_height":null,"min_width":null,"object_fit":null,"object_position":null,"order":null,"overflow":null,"padding":null,"right":null,"top":null,"visibility":null,"width":null}}},"version_major":2,"version_minor":0}}},"nbformat":4,"nbformat_minor":5}