PyPI cross-check
Does the method travel? The same protocol on a second ecosystem
Every number on this tab comes from the embargoed rolling backtest re-run on PyPI: the download-ranked packages with a GitHub repository, OSV advisories as the outcome, and only each package's own advisory history as input. Same folds, same 180-day horizon, same label embargo at every forecast date.
The package look-up shows where a package ranked at every forecast date and the OSV advisories that followed; the fold picker applies to the case study at the bottom. Try vllm for what advisory history cannot do: a package sits at the bottom of the ranking until its first advisory, then near the top.
The leak replicates
Stored labels inflate PyPI results the same way they inflated Jenkins
The PyPI advisory-history run was scored twice, with stored labels and under the embargo, exactly as the Jenkins runs were; the Jenkins pair with the same inputs is drawn beside it. Each pair is one configuration scored twice on identical test data. Stored labels let a training label be set by an advisory published inside the test window; embargoed rebuilds every training label from advisories known before the fold's forecast date. The gap is the leak. Hover a bar for the value.
ROC-AUC — separation of advisory-bound plugins from the rest; 0.5 is chance
Average precision — concentration of advisories at the top of the ranked list
The verdict inverts
ROC-AUC per fold, PyPI beside Jenkins, advisory history only
The same three advisory-history inputs (advisories to date, CVEs to date, worst CVSS to date) scored at the same thirteen forecast dates. On PyPI they clear the pre-registered criterion by a wide margin in every fold; on Jenkins the same inputs hover around the criterion and touch chance in some folds. Hover a point for the fold's outcome count.
What the development ranking looks like
Advisories caught vs packages reviewed
Reviewing the top 20% of packages by score would have caught 67% of the advisories that followed across the development folds (by fold: 63%, 65%, 65%, 64%, 65%, 67%, 73%, 71%, 73%, 72%, 71%, 69%, 66%), against a base rate of 1.2%. The thick line is the pooled ranking over every development fold; thin lines are the individual folds (183,378 package-months, 2,143 advisory outcomes). The dashed diagonal is random order. Right: the ROC curve of each fold for the first configuration.
Development-fold outcomes
Top-25 per development fold vs. what followed
Advisory history · Logistic Regression · advisory_only_logistic. Each package's best month inside the fold is shown; training labels were rebuilt from advisories known before the fold's forecast date, and the folds were run once.
Across the development folds shown, 13 of the 25 top-25 slots were followed by an advisory within 180 days. That is the honest precision at 25 for this configuration: a ranking signal with real lift over the base rate, not an operational triage list.
Development fold 2025-05 (13 of 25 confirmed)
| # | Package | Scored | Score | Confirmed | Severity | Advisory date | Lead time | Advisory |
|---|---|---|---|---|---|---|---|---|
| 1 | apache-airflow | 2025-05 | 100.0% | ✓ | Critical (9.8) | 2025-12-17 | 230 days | PYSEC-2025-87 |
| 2 | django | 2025-05 | 100.0% | ✓ | Critical (9.8) | 2025-10-01 | 153 days | PYSEC-2025-106 |
| 3 | mlflow | 2025-05 | 100.0% | ✓ | High (8.1) | 2025-10-29 | 181 days | GHSA-5cvj-7rg6-jggj |
| 9 | ansible | 2025-05 | 100.0% | ✓ | Medium (5.5) | 2025-12-04 | 217 days | GHSA-8ggh-xwr9-3373 |
| 10 | pillow | 2025-05 | 100.0% | ✓ | High (7.1) | 2025-07-01 | 61 days | GHSA-xg8h-j46f-w952 |
| 13 | vllm | 2025-06 | 100.0% | ✓ | High (8.8) | 2025-08-21 | 81 days | GHSA-79j6-g2m3-jgfw |
| 18 | pycti | 2025-06 | 100.0% | ✓ | Medium (5.4) | 2025-07-18 | 47 days | PYSEC-2025-181 |
| 19 | torch | 2025-05 | 100.0% | ✓ | High (7.5) | 2025-09-25 | 147 days | PYSEC-2025-203 |
| 20 | flask-appbuilder | 2025-06 | 100.0% | ✓ | Medium (6.5) | 2025-09-11 | 102 days | GHSA-765j-9r45-w2q2 |
| 21 | aiohttp | 2025-05 | 100.0% | ✓ | Low | 2025-07-14 | 74 days | GHSA-9548-qrrj-x5pj |
| 22 | jinja2 | 2025-05 | 100.0% | ✓ | High (7.1) | 2025-06-10 | 40 days | PYSEC-2025-74 |
| 23 | werkzeug | 2025-05 | 100.0% | ✓ | Moderate | 2025-12-02 | 215 days | GHSA-hgf8-39gv-g3f2 |
| 25 | transformers | 2025-06 | 100.0% | ✓ | High (7.8) | 2025-12-23 | 205 days | PYSEC-2025-211 |
The other 12 of the top 25
| # | Package | Scored | Score | Confirmed | Severity | Advisory date | Lead time | Advisory |
|---|---|---|---|---|---|---|---|---|
| 4 | opencv-contrib-python | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 5 | opencv-python | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 6 | tensorflow | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 7 | tensorflow-cpu | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 8 | gradio | 2025-06 | 100.0% | – | no advisory in the 180-day window | |||
| 11 | paddlepaddle | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 12 | litellm | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 14 | notebook | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 15 | opencv-contrib-python-headless | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 16 | opencv-python-headless | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 17 | langchain | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
| 24 | apache-airflow-providers-apache-hive | 2025-05 | 100.0% | – | no advisory in the 180-day window | |||
Advisory details come from OSV's PyPI advisory database (PYSEC and GHSA records); a confirmed row without details is one whose stored label was positive. Lead time = days from the scored month to publication. Package names link to their look-up on this tab.
How to read this
The mechanism replicates; the verdict on the features does not
Two things travel from Jenkins to PyPI unchanged: the label leak, which inflates stored-label results in both ecosystems by a similar margin, and the embargoed protocol that removes it. What does not travel is the verdict on any one feature family. A package's own advisory history is a strong signal on PyPI (pooled ROC-AUC 0.778) and barely clears the criterion on Jenkins (0.553). The likely reason is how advisories arrive: on PyPI they recur within the same widely used packages, while Jenkins advisories often come in ecosystem-wide batches that reach plugins with no prior record. Either way, a signal has to be re-validated in each ecosystem, which is why CANARY's contribution is the validation method rather than a feature list.
Scope of the cross-check: advisory-history inputs only (no activity clocks, contributor dynamics or install base were collected for PyPI); a download-ranked universe rather than a complete registry; and a monitoring setting, so the ranking is among packages the model has seen before, with no group split. Read the PyPI numbers as evidence that the protocol transfers, not as a PyPI risk model.