PyPI cross-check
Does the method travel? The same protocol on a second ecosystem
Every number on this tab comes from the embargoed rolling backtest re-run on PyPI: the download-ranked packages with a GitHub repository, OSV advisories as the outcome, and only each package's own advisory history as input. Same folds, same 180-day horizon, same label embargo at every forecast date.
The package look-up shows where a package ranked at every forecast date and the OSV advisories that followed; the fold picker applies to the case study at the bottom. Try vllm for what advisory history cannot do: a package sits at the bottom of the ranking until its first advisory, then near the top.
Honest track record
Where transformers ranked before each window
Advisory history · Logistic Regression, embargoed rolling backtest: the rank this package held among every scored package at each forecast date, using only what was knowable then.
Scored at 26 forecast dates; on average in the top 5.0% of packages. An advisory followed within 180 days of 24 of those dates.
Download rank #246 on PyPI; 113,570,167 downloads a month; 35 OSV advisories on record. PyPI · GitHub
| Forecast date | Score | Rank | Share of panel | What followed |
|---|---|---|---|---|
| 2023-05 | 31.4% | 6462 of 7,053 | top 92% | none within 180 days |
| 2023-06 | 84.7% | 228 of 7,053 | top 3.2% | GHSA-3863-2447-669p 2023-12-19 |
| 2023-07 | 83.9% | 231 of 7,053 | top 3.3% | GHSA-3863-2447-669p 2023-12-19 |
| 2023-08 | 83.9% | 236 of 7,053 | top 3.3% | GHSA-3863-2447-669p 2023-12-19 |
| 2023-09 | 82.9% | 242 of 7,053 | top 3.4% | GHSA-3863-2447-669p 2023-12-19 |
| 2023-10 | 82.9% | 246 of 7,053 | top 3.5% | GHSA-3863-2447-669p 2023-12-19 |
| 2023-11 | 82.2% | 254 of 7,053 | top 3.6% | GHSA-3863-2447-669p 2023-12-19 |
| 2023-12 | 82.2% | 257 of 7,053 | top 3.6% | GHSA-3863-2447-669p 2023-12-19 |
| 2024-01 | 99.0% | 74 of 7,053 | top 1.0% | GHSA-37q5-v5qm-c9v8 2024-04-10 |
| 2024-02 | 99.0% | 76 of 7,053 | top 1.1% | GHSA-37q5-v5qm-c9v8 2024-04-10 |
| 2024-03 | 98.9% | 78 of 7,053 | top 1.1% | GHSA-37q5-v5qm-c9v8 2024-04-10 |
| 2024-04 | 98.9% | 81 of 7,053 | top 1.1% | none within 180 days |
| 2024-05 | 99.8% | 48 of 7,053 | top 0.67% | PYSEC-2024-227 2024-11-22 |
| 2024-06 | 99.8% | 50 of 7,053 | top 0.69% | PYSEC-2024-227 2024-11-22 |
| 2024-07 | 99.8% | 57 of 7,053 | top 0.79% | PYSEC-2024-227 2024-11-22 |
| 2024-08 | 99.8% | 61 of 7,053 | top 0.85% | PYSEC-2024-227 2024-11-22 |
| 2024-09 | 99.7% | 65 of 7,053 | top 0.91% | PYSEC-2024-227 2024-11-22 |
| 2024-10 | 99.7% | 66 of 7,053 | top 0.92% | PYSEC-2024-227 2024-11-22 |
| 2024-11 | 99.7% | 67 of 7,053 | top 0.94% | PYSEC-2024-227 2024-11-22 |
| 2024-12 | 100.0% | 37 of 7,053 | top 0.51% | GHSA-6rvg-6v2m-4j46 2025-03-20 |
| 2025-01 | 99.9% | 37 of 7,053 | top 0.51% | GHSA-6rvg-6v2m-4j46 2025-03-20 |
| 2025-02 | 99.9% | 37 of 7,053 | top 0.51% | GHSA-6rvg-6v2m-4j46 2025-03-20 |
| 2025-03 | 99.9% | 37 of 7,053 | top 0.51% | GHSA-6rvg-6v2m-4j46 2025-03-20 |
| 2025-04 | 100.0% | 34 of 7,053 | top 0.47% | GHSA-fpwr-67px-3qhx 2025-04-29 |
| 2025-05 | 100.0% | 27 of 7,053 | top 0.37% | GHSA-qq3j-4f4f-9583 2025-05-19 |
| 2025-06 | 100.0% | 25 of 7,053 | top 0.34% | GHSA-489j-g2vx-39wf 2025-07-07 |
The leak replicates
Stored labels inflate PyPI results the same way they inflated Jenkins
The PyPI advisory-history run was scored twice, with stored labels and under the embargo, exactly as the Jenkins runs were; the Jenkins pair with the same inputs is drawn beside it. Each pair is one configuration scored twice on identical test data. Stored labels let a training label be set by an advisory published inside the test window; embargoed rebuilds every training label from advisories known before the fold's forecast date. The gap is the leak. Hover a bar for the value.
ROC-AUC — separation of advisory-bound plugins from the rest; 0.5 is chance
Average precision — concentration of advisories at the top of the ranked list
The verdict inverts
ROC-AUC per fold, PyPI beside Jenkins, advisory history only
The same three advisory-history inputs (advisories to date, CVEs to date, worst CVSS to date) scored at the same thirteen forecast dates. On PyPI they clear the pre-registered criterion by a wide margin in every fold; on Jenkins the same inputs hover around the criterion and touch chance in some folds. Hover a point for the fold's outcome count.
What the development ranking looks like
Advisories caught vs packages reviewed
Reviewing the top 20% of packages by score would have caught 67% of the advisories that followed across the development folds (by fold: 63%, 65%, 65%, 64%, 65%, 67%, 73%, 71%, 73%, 72%, 71%, 69%, 66%), against a base rate of 1.2%. The thick line is the pooled ranking over every development fold; thin lines are the individual folds (183,378 package-months, 2,143 advisory outcomes). The dashed diagonal is random order. Right: the ROC curve of each fold for the first configuration.
Development-fold outcomes
Top-25 per development fold vs. what followed
Advisory history · Logistic Regression · advisory_only_logistic. Each package's best month inside the fold is shown; training labels were rebuilt from advisories known before the fold's forecast date, and the folds were run once.
Across the development folds shown, 13 of the 25 top-25 slots were followed by an advisory within 180 days. That is the honest precision at 25 for this configuration: a ranking signal with real lift over the base rate, not an operational triage list.
Development fold 2025-05 (13 of 25 confirmed)
| # | Package | Scored | Score | Confirmed | Severity | Advisory date | Lead time | Advisory |
|---|---|---|---|---|---|---|---|---|
| 1 | apache-airflow | 2025-05 | 100.0% | ✓ | Critical (9.8) | 2025-12-17 | 230 days | PYSEC-2025-87 |
| 2 | django | 2025-05 | 100.0% | ✓ | Critical (9.8) | 2025-10-01 | 153 days | PYSEC-2025-106 |
| 3 | mlflow | 2025-05 | 100.0% | ✓ | High (8.1) | 2025-10-29 | 181 days | GHSA-5cvj-7rg6-jggj |
| 9 | ansible | 2025-05 | 100.0% | ✓ | Medium (5.5) | 2025-12-04 | 217 days | GHSA-8ggh-xwr9-3373 |
| 10 | pillow | 2025-05 | 100.0% | ✓ | High (7.1) | 2025-07-01 | 61 days | GHSA-xg8h-j46f-w952 |
| 13 | vllm | 2025-06 | 100.0% | ✓ | High (8.8) | 2025-08-21 | 81 days | GHSA-79j6-g2m3-jgfw |
| 18 | pycti | 2025-06 | 100.0% | ✓ | Medium (5.4) | 2025-07-18 | 47 days | PYSEC-2025-181 |
| 19 | torch | 2025-05 | 100.0% | ✓ | High (7.5) | 2025-09-25 | 147 days | PYSEC-2025-203 |
| 20 | flask-appbuilder | 2025-06 | 100.0% | ✓ | Medium (6.5) | 2025-09-11 | 102 days | GHSA-765j-9r45-w2q2 |
| 21 | aiohttp | 2025-05 | 100.0% | ✓ | Low | 2025-07-14 | 74 days | GHSA-9548-qrrj-x5pj |
| 22 | jinja2 | 2025-05 | 100.0% | ✓ | High (7.1) | 2025-06-10 | 40 days | PYSEC-2025-74 |
| 23 | werkzeug | 2025-05 | 100.0% | ✓ | Moderate | 2025-12-02 | 215 days | GHSA-hgf8-39gv-g3f2 |
| 25 | transformers | 2025-06 | 100.0% | ✓ | High (7.8) | 2025-12-23 | 205 days | PYSEC-2025-211 |
The other 12 of the top 25
| # | Package | Scored | Score | Confirmed | Severity | Advisory date | Lead time | Advisory |
|---|---|---|---|---|---|---|---|---|
| 4 | opencv-contrib-python | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 5 | opencv-python | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 6 | tensorflow | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 7 | tensorflow-cpu | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 8 | gradio | 2025-06 | 100.0% | – | no advisory in the 180-day window | |||
| 11 | paddlepaddle | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 12 | litellm | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 14 | notebook | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 15 | opencv-contrib-python-headless | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 16 | opencv-python-headless | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 17 | langchain | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 24 | apache-airflow-providers-apache-hive | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
Advisory details come from OSV's PyPI advisory database (PYSEC and GHSA records); a confirmed row without details is one whose stored label was positive. Lead time = days from the scored month to publication. Package names link to their look-up on this tab.
How to read this
The mechanism replicates; the verdict on the features does not
Two things travel from Jenkins to PyPI unchanged: the label leak, which inflates stored-label results in both ecosystems by a similar margin, and the embargoed protocol that removes it. What does not travel is the verdict on any one feature family. A package's own advisory history is a strong signal on PyPI (pooled ROC-AUC 0.778) and barely clears the criterion on Jenkins (0.553). The likely reason is how advisories arrive: on PyPI they recur within the same widely used packages, while Jenkins advisories often come in ecosystem-wide batches that reach plugins with no prior record. Either way, a signal has to be re-validated in each ecosystem, which is why CANARY's contribution is the validation method rather than a feature list.
Scope of the cross-check: advisory-history inputs only (no activity clocks, contributor dynamics or install base were collected for PyPI); a download-ranked universe rather than a complete registry; and a monitoring setting, so the ranking is among packages the model has seen before, with no group split. Read the PyPI numbers as evidence that the protocol transfers, not as a PyPI risk model.