PyPI cross-check
Does the method travel? The same protocol on a second ecosystem
Every number on this tab comes from the embargoed rolling backtest re-run on PyPI: the download-ranked packages with a GitHub repository, OSV advisories as the outcome, and only each package's own advisory history as input. Same folds, same 180-day horizon, same label embargo at every forecast date.
The package look-up shows where a package ranked at every forecast date and the OSV advisories that followed; the fold picker applies to the case study at the bottom. Try vllm for what advisory history cannot do: a package sits at the bottom of the ranking until its first advisory, then near the top.
Honest track record
Where notebook ranked before each window
Advisory history · Logistic Regression, embargoed rolling backtest: the rank this package held among every scored package at each forecast date, using only what was knowable then.
Scored at 26 forecast dates; on average in the top 0.16% of packages. An advisory followed within 180 days of 12 of those dates.
Download rank #470 on PyPI; 46,924,213 downloads a month; 31 OSV advisories on record. PyPI · GitHub
| Forecast date | Score | Rank | Share of panel | What followed |
|---|---|---|---|---|
| 2023-05 | 100.0% | 9 of 7,053 | top 0.11% | none within 180 days |
| 2023-06 | 100.0% | 9 of 7,053 | top 0.11% | none within 180 days |
| 2023-07 | 100.0% | 11 of 7,053 | top 0.14% | GHSA-44cc-43rp-5947 2024-01-19 |
| 2023-08 | 100.0% | 11 of 7,053 | top 0.14% | GHSA-44cc-43rp-5947 2024-01-19 |
| 2023-09 | 100.0% | 11 of 7,053 | top 0.14% | GHSA-44cc-43rp-5947 2024-01-19 |
| 2023-10 | 100.0% | 11 of 7,053 | top 0.14% | GHSA-44cc-43rp-5947 2024-01-19 |
| 2023-11 | 100.0% | 12 of 7,053 | top 0.16% | GHSA-44cc-43rp-5947 2024-01-19 |
| 2023-12 | 100.0% | 12 of 7,053 | top 0.16% | GHSA-44cc-43rp-5947 2024-01-19 |
| 2024-01 | 100.0% | 13 of 7,053 | top 0.17% | none within 180 days |
| 2024-02 | 100.0% | 13 of 7,053 | top 0.17% | GHSA-9q39-rmj3-p4r2 2024-08-29 |
| 2024-03 | 100.0% | 13 of 7,053 | top 0.17% | GHSA-9q39-rmj3-p4r2 2024-08-29 |
| 2024-04 | 100.0% | 13 of 7,053 | top 0.17% | GHSA-9q39-rmj3-p4r2 2024-08-29 |
| 2024-05 | 100.0% | 14 of 7,053 | top 0.18% | GHSA-9q39-rmj3-p4r2 2024-08-29 |
| 2024-06 | 100.0% | 15 of 7,053 | top 0.20% | GHSA-9q39-rmj3-p4r2 2024-08-29 |
| 2024-07 | 100.0% | 15 of 7,053 | top 0.20% | GHSA-9q39-rmj3-p4r2 2024-08-29 |
| 2024-08 | 100.0% | 15 of 7,053 | top 0.20% | none within 180 days |
| 2024-09 | 100.0% | 13 of 7,053 | top 0.17% | none within 180 days |
| 2024-10 | 100.0% | 13 of 7,053 | top 0.17% | none within 180 days |
| 2024-11 | 100.0% | 13 of 7,053 | top 0.17% | none within 180 days |
| 2024-12 | 100.0% | 12 of 7,053 | top 0.16% | none within 180 days |
| 2025-01 | 100.0% | 12 of 7,053 | top 0.16% | none within 180 days |
| 2025-02 | 100.0% | 12 of 7,053 | top 0.16% | none within 180 days |
| 2025-03 | 100.0% | 12 of 7,053 | top 0.16% | none within 180 days |
| 2025-04 | 100.0% | 13 of 7,053 | top 0.17% | none within 180 days |
| 2025-05 | 100.0% | 13 of 7,053 | top 0.17% | none within 180 days |
| 2025-06 | 100.0% | 14 of 7,053 | top 0.18% | none within 180 days |
The leak replicates
Stored labels inflate PyPI results the same way they inflated Jenkins
The PyPI advisory-history run was scored twice, with stored labels and under the embargo, exactly as the Jenkins runs were; the Jenkins pair with the same inputs is drawn beside it. Each pair is one configuration scored twice on identical test data. Stored labels let a training label be set by an advisory published inside the test window; embargoed rebuilds every training label from advisories known before the fold's forecast date. The gap is the leak. Hover a bar for the value.
ROC-AUC — separation of advisory-bound plugins from the rest; 0.5 is chance
Average precision — concentration of advisories at the top of the ranked list
The verdict inverts
ROC-AUC per fold, PyPI beside Jenkins, advisory history only
The same three advisory-history inputs (advisories to date, CVEs to date, worst CVSS to date) scored at the same thirteen forecast dates. On PyPI they clear the pre-registered criterion by a wide margin in every fold; on Jenkins the same inputs hover around the criterion and touch chance in some folds. Hover a point for the fold's outcome count.
What the development ranking looks like
Advisories caught vs packages reviewed
Reviewing the top 20% of packages by score would have caught 67% of the advisories that followed across the development folds (by fold: 63%, 65%, 65%, 64%, 65%, 67%, 73%, 71%, 73%, 72%, 71%, 69%, 66%), against a base rate of 1.2%. The thick line is the pooled ranking over every development fold; thin lines are the individual folds (183,378 package-months, 2,143 advisory outcomes). The dashed diagonal is random order. Right: the ROC curve of each fold for the first configuration.
Development-fold outcomes
Top-25 per development fold vs. what followed
Advisory history · Logistic Regression · advisory_only_logistic. Each package's best month inside the fold is shown; training labels were rebuilt from advisories known before the fold's forecast date, and the folds were run once.
Across the development folds shown, 13 of the 25 top-25 slots were followed by an advisory within 180 days. That is the honest precision at 25 for this configuration: a ranking signal with real lift over the base rate, not an operational triage list.
Development fold 2025-05 (13 of 25 confirmed)
| # | Package | Scored | Score | Confirmed | Severity | Advisory date | Lead time | Advisory |
|---|---|---|---|---|---|---|---|---|
| 1 | apache-airflow | 2025-05 | 100.0% | ✓ | Critical (9.8) | 2025-12-17 | 230 days | PYSEC-2025-87 |
| 2 | django | 2025-05 | 100.0% | ✓ | Critical (9.8) | 2025-10-01 | 153 days | PYSEC-2025-106 |
| 3 | mlflow | 2025-05 | 100.0% | ✓ | High (8.1) | 2025-10-29 | 181 days | GHSA-5cvj-7rg6-jggj |
| 9 | ansible | 2025-05 | 100.0% | ✓ | Medium (5.5) | 2025-12-04 | 217 days | GHSA-8ggh-xwr9-3373 |
| 10 | pillow | 2025-05 | 100.0% | ✓ | High (7.1) | 2025-07-01 | 61 days | GHSA-xg8h-j46f-w952 |
| 13 | vllm | 2025-06 | 100.0% | ✓ | High (8.8) | 2025-08-21 | 81 days | GHSA-79j6-g2m3-jgfw |
| 18 | pycti | 2025-06 | 100.0% | ✓ | Medium (5.4) | 2025-07-18 | 47 days | PYSEC-2025-181 |
| 19 | torch | 2025-05 | 100.0% | ✓ | High (7.5) | 2025-09-25 | 147 days | PYSEC-2025-203 |
| 20 | flask-appbuilder | 2025-06 | 100.0% | ✓ | Medium (6.5) | 2025-09-11 | 102 days | GHSA-765j-9r45-w2q2 |
| 21 | aiohttp | 2025-05 | 100.0% | ✓ | Low | 2025-07-14 | 74 days | GHSA-9548-qrrj-x5pj |
| 22 | jinja2 | 2025-05 | 100.0% | ✓ | High (7.1) | 2025-06-10 | 40 days | PYSEC-2025-74 |
| 23 | werkzeug | 2025-05 | 100.0% | ✓ | Moderate | 2025-12-02 | 215 days | GHSA-hgf8-39gv-g3f2 |
| 25 | transformers | 2025-06 | 100.0% | ✓ | High (7.8) | 2025-12-23 | 205 days | PYSEC-2025-211 |
The other 12 of the top 25
| # | Package | Scored | Score | Confirmed | Severity | Advisory date | Lead time | Advisory |
|---|---|---|---|---|---|---|---|---|
| 4 | opencv-contrib-python | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 5 | opencv-python | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 6 | tensorflow | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 7 | tensorflow-cpu | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 8 | gradio | 2025-06 | 100.0% | – | no advisory in the 180-day window | |||
| 11 | paddlepaddle | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 12 | litellm | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 14 | notebook | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 15 | opencv-contrib-python-headless | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 16 | opencv-python-headless | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 17 | langchain | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 24 | apache-airflow-providers-apache-hive | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
Advisory details come from OSV's PyPI advisory database (PYSEC and GHSA records); a confirmed row without details is one whose stored label was positive. Lead time = days from the scored month to publication. Package names link to their look-up on this tab.
How to read this
The mechanism replicates; the verdict on the features does not
Two things travel from Jenkins to PyPI unchanged: the label leak, which inflates stored-label results in both ecosystems by a similar margin, and the embargoed protocol that removes it. What does not travel is the verdict on any one feature family. A package's own advisory history is a strong signal on PyPI (pooled ROC-AUC 0.778) and barely clears the criterion on Jenkins (0.553). The likely reason is how advisories arrive: on PyPI they recur within the same widely used packages, while Jenkins advisories often come in ecosystem-wide batches that reach plugins with no prior record. Either way, a signal has to be re-validated in each ecosystem, which is why CANARY's contribution is the validation method rather than a feature list.
Scope of the cross-check: advisory-history inputs only (no activity clocks, contributor dynamics or install base were collected for PyPI); a download-ranked universe rather than a complete registry; and a monitoring setting, so the ranking is among packages the model has seen before, with no group split. Read the PyPI numbers as evidence that the protocol transfers, not as a PyPI risk model.