Web interface

CANARY Web Console

Security-advisory forecasting for Jenkins plugins from public signals, evaluated honestly: four validation layers, one pre-registered criterion, and every number on this site traceable to the layer that produced it.

PyPI cross-check

Does the method travel? The same protocol on a second ecosystem

Every number on this tab comes from the embargoed rolling backtest re-run on PyPI: the download-ranked packages with a GitHub repository, OSV advisories as the outcome, and only each package's own advisory history as input. Same folds, same 180-day horizon, same label embargo at every forecast date.

Embargoed labels at every fold
Packages in the panel7,053download-ranked, with a resolvable GitHub repository
Package-months scored183,37813 embargoed folds, 2023-05 to 2025-05
Advisory outcomes2,143base rate 1.17%
Pooled ROC-AUC, PyPI0.778Advisory history · Logistic Regression
Same features, Jenkins0.553Advisory history · XGBoost; PyPI with the same model: 0.720

The package look-up shows where a package ranked at every forecast date and the OSV advisories that followed; the fold picker applies to the case study at the bottom. Try vllm for what advisory history cannot do: a package sits at the bottom of the ranking until its first advisory, then near the top.

The leak replicates

Stored labels inflate PyPI results the same way they inflated Jenkins

The PyPI advisory-history run was scored twice, with stored labels and under the embargo, exactly as the Jenkins runs were; the Jenkins pair with the same inputs is drawn beside it. Each pair is one configuration scored twice on identical test data. Stored labels let a training label be set by an advisory published inside the test window; embargoed rebuilds every training label from advisories known before the fold's forecast date. The gap is the leak. Hover a bar for the value.

ROC-AUC — separation of advisory-bound plugins from the rest; 0.5 is chance

stored labels (leaky)embargoed (honest)chance 0.50PyPI · Advisory history · XGBoostrolling, 13 folds, 2,143 positivesPyPI · Advisory history · XGBoost stored: ROC-AUC 0.7730.773PyPI · Advisory history · XGBoost embargoed: ROC-AUC 0.7200.720Jenkins · Advisory history · XGBoostrolling, 13 folds, 760 positivesJenkins · Advisory history · XGBoost stored: ROC-AUC 0.6020.602Jenkins · Advisory history · XGBoost embargoed: ROC-AUC 0.5530.553

Average precision — concentration of advisories at the top of the ranked list

stored labels (leaky)embargoed (honest)PyPI · Advisory history · XGBoostrolling, 13 folds, 2,143 positivesPyPI · Advisory history · XGBoost stored: Average precision 0.2880.288PyPI · Advisory history · XGBoost embargoed: Average precision 0.1870.187Jenkins · Advisory history · XGBoostrolling, 13 folds, 760 positivesJenkins · Advisory history · XGBoost stored: Average precision 0.0370.037Jenkins · Advisory history · XGBoost embargoed: Average precision 0.0240.024

The verdict inverts

ROC-AUC per fold, PyPI beside Jenkins, advisory history only

The same three advisory-history inputs (advisories to date, CVEs to date, worst CVSS to date) scored at the same thirteen forecast dates. On PyPI they clear the pre-registered criterion by a wide margin in every fold; on Jenkins the same inputs hover around the criterion and touch chance in some folds. Hover a point for the fold's outcome count.

0.450.500.550.600.650.700.750.800.85chance 0.50pre-registered criterion 0.552023-052023-072023-092023-112024-012024-032024-052024-072024-092024-112025-012025-032025-05PyPI · Advisory history · Logistic Regression 2023-05 · ROC-AUC 0.763 · 144 advisory outcomesPyPI · Advisory history · Logistic Regression 2023-07 · ROC-AUC 0.778 · 159 advisory outcomesPyPI · Advisory history · Logistic Regression 2023-09 · ROC-AUC 0.763 · 179 advisory outcomesPyPI · Advisory history · Logistic Regression 2023-11 · ROC-AUC 0.765 · 200 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-01 · ROC-AUC 0.765 · 207 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-03 · ROC-AUC 0.794 · 193 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-05 · ROC-AUC 0.819 · 171 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-07 · ROC-AUC 0.808 · 129 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-09 · ROC-AUC 0.822 · 147 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-11 · ROC-AUC 0.816 · 137 advisory outcomesPyPI · Advisory history · Logistic Regression 2025-01 · ROC-AUC 0.813 · 148 advisory outcomesPyPI · Advisory history · Logistic Regression 2025-03 · ROC-AUC 0.798 · 153 advisory outcomesPyPI · Advisory history · Logistic Regression 2025-05 · ROC-AUC 0.788 · 176 advisory outcomesPyPI · Advisory history · XGBoost 2023-05 · ROC-AUC 0.647 · 144 advisory outcomesPyPI · Advisory history · XGBoost 2023-07 · ROC-AUC 0.665 · 159 advisory outcomesPyPI · Advisory history · XGBoost 2023-09 · ROC-AUC 0.698 · 179 advisory outcomesPyPI · Advisory history · XGBoost 2023-11 · ROC-AUC 0.695 · 200 advisory outcomesPyPI · Advisory history · XGBoost 2024-01 · ROC-AUC 0.731 · 207 advisory outcomesPyPI · Advisory history · XGBoost 2024-03 · ROC-AUC 0.741 · 193 advisory outcomesPyPI · Advisory history · XGBoost 2024-05 · ROC-AUC 0.759 · 171 advisory outcomesPyPI · Advisory history · XGBoost 2024-07 · ROC-AUC 0.745 · 129 advisory outcomesPyPI · Advisory history · XGBoost 2024-09 · ROC-AUC 0.776 · 147 advisory outcomesPyPI · Advisory history · XGBoost 2024-11 · ROC-AUC 0.753 · 137 advisory outcomesPyPI · Advisory history · XGBoost 2025-01 · ROC-AUC 0.765 · 148 advisory outcomesPyPI · Advisory history · XGBoost 2025-03 · ROC-AUC 0.771 · 153 advisory outcomesPyPI · Advisory history · XGBoost 2025-05 · ROC-AUC 0.780 · 176 advisory outcomesJenkins · Advisory history · XGBoost 2023-05 · ROC-AUC 0.553 · 132 advisory outcomesJenkins · Advisory history · XGBoost 2023-07 · ROC-AUC 0.525 · 90 advisory outcomesJenkins · Advisory history · XGBoost 2023-09 · ROC-AUC 0.595 · 67 advisory outcomesJenkins · Advisory history · XGBoost 2023-11 · ROC-AUC 0.558 · 55 advisory outcomesJenkins · Advisory history · XGBoost 2024-01 · ROC-AUC 0.583 · 40 advisory outcomesJenkins · Advisory history · XGBoost 2024-03 · ROC-AUC 0.568 · 22 advisory outcomesJenkins · Advisory history · XGBoost 2024-05 · ROC-AUC 0.608 · 23 advisory outcomesJenkins · Advisory history · XGBoost 2024-07 · ROC-AUC 0.623 · 32 advisory outcomesJenkins · Advisory history · XGBoost 2024-09 · ROC-AUC 0.575 · 42 advisory outcomesJenkins · Advisory history · XGBoost 2024-11 · ROC-AUC 0.576 · 39 advisory outcomesJenkins · Advisory history · XGBoost 2025-01 · ROC-AUC 0.525 · 66 advisory outcomesJenkins · Advisory history · XGBoost 2025-03 · ROC-AUC 0.540 · 75 advisory outcomesJenkins · Advisory history · XGBoost 2025-05 · ROC-AUC 0.515 · 77 advisory outcomesPyPI · Advisory history · Logistic RegressionPyPI · Advisory history · XGBoostJenkins · Advisory history · XGBoost

What the development ranking looks like

Advisories caught vs packages reviewed

Reviewing the top 20% of packages by score would have caught 67% of the advisories that followed across the development folds (by fold: 63%, 65%, 65%, 64%, 65%, 67%, 73%, 71%, 73%, 72%, 71%, 69%, 66%), against a base rate of 1.2%. The thick line is the pooled ranking over every development fold; thin lines are the individual folds (183,378 package-months, 2,143 advisory outcomes). The dashed diagonal is random order. Right: the ROC curve of each fold for the first configuration.

0%0%20%20%40%40%60%60%80%80%100%100%random orderAdvisory history · Logistic Regression review top 1% of plugins → 30% of advisoriesAdvisory history · Logistic Regression review top 2% of plugins → 43% of advisoriesAdvisory history · Logistic Regression review top 3% of plugins → 51% of advisoriesAdvisory history · Logistic Regression review top 5% of plugins → 61% of advisoriesAdvisory history · Logistic Regression review top 8% of plugins → 62% of advisoriesAdvisory history · Logistic Regression review top 10% of plugins → 63% of advisoriesAdvisory history · Logistic Regression review top 15% of plugins → 65% of advisoriesAdvisory history · Logistic Regression review top 20% of plugins → 67% of advisoriesAdvisory history · Logistic Regression review top 25% of plugins → 68% of advisoriesAdvisory history · Logistic Regression review top 30% of plugins → 70% of advisoriesAdvisory history · Logistic Regression review top 40% of plugins → 73% of advisoriesAdvisory history · Logistic Regression review top 50% of plugins → 77% of advisoriesAdvisory history · Logistic Regression review top 60% of plugins → 81% of advisoriesAdvisory history · Logistic Regression review top 70% of plugins → 86% of advisoriesAdvisory history · Logistic Regression review top 80% of plugins → 91% of advisoriesAdvisory history · Logistic Regression review top 90% of plugins → 96% of advisoriesAdvisory history · Logistic Regression review top 100% of plugins → 100% of advisoriestop 10% → 63% caughttop 20% → 67% caughttop 30% → 70% caughtshare of packages reviewedAdvisory history · Logistic Regression
0.00.00.20.20.40.40.60.60.80.81.01.0false-positive rate2023-05 (144 outcomes) · AUC 0.7632023-07 (159 outcomes) · AUC 0.7782023-09 (179 outcomes) · AUC 0.7632023-11 (200 outcomes) · AUC 0.7652024-01 (207 outcomes) · AUC 0.7652024-03 (193 outcomes) · AUC 0.7942024-05 (171 outcomes) · AUC 0.8192024-07 (129 outcomes) · AUC 0.8082024-09 (147 outcomes) · AUC 0.8222024-11 (137 outcomes) · AUC 0.8162025-01 (148 outcomes) · AUC 0.8132025-03 (153 outcomes) · AUC 0.7982025-05 (176 outcomes) · AUC 0.788

Development-fold outcomes

Top-25 per development fold vs. what followed

Advisory history · Logistic Regression · advisory_only_logistic. Each package's best month inside the fold is shown; training labels were rebuilt from advisories known before the fold's forecast date, and the folds were run once.

Embargoed

Across the development folds shown, 13 of the 25 top-25 slots were followed by an advisory within 180 days. That is the honest precision at 25 for this configuration: a ranking signal with real lift over the base rate, not an operational triage list.

Development fold 2025-05 (13 of 25 confirmed)

Scored: 2025-05 and 2025-06  |  Labels known as of: 2025-06  |  Fold ROC-AUC: 0.788  |  Top-25 precision: 13/25 (52%)  |  Base rate: 1.2% (41.7×)
#PackageScoredScoreConfirmedSeverityAdvisory dateLead timeAdvisory
1apache-airflow2025-05100.0%✓Critical (9.8)2025-12-17230 daysPYSEC-2025-87
2django2025-05100.0%✓Critical (9.8)2025-10-01153 daysPYSEC-2025-106
3mlflow2025-05100.0%✓High (8.1)2025-10-29181 daysGHSA-5cvj-7rg6-jggj
9ansible2025-05100.0%✓Medium (5.5)2025-12-04217 daysGHSA-8ggh-xwr9-3373
10pillow2025-05100.0%✓High (7.1)2025-07-0161 daysGHSA-xg8h-j46f-w952
13vllm2025-06100.0%✓High (8.8)2025-08-2181 daysGHSA-79j6-g2m3-jgfw
18pycti2025-06100.0%✓Medium (5.4)2025-07-1847 daysPYSEC-2025-181
19torch2025-05100.0%✓High (7.5)2025-09-25147 daysPYSEC-2025-203
20flask-appbuilder2025-06100.0%✓Medium (6.5)2025-09-11102 daysGHSA-765j-9r45-w2q2
21aiohttp2025-05100.0%✓Low2025-07-1474 daysGHSA-9548-qrrj-x5pj
22jinja22025-05100.0%✓High (7.1)2025-06-1040 daysPYSEC-2025-74
23werkzeug2025-05100.0%✓Moderate2025-12-02215 daysGHSA-hgf8-39gv-g3f2
25transformers2025-06100.0%✓High (7.8)2025-12-23205 daysPYSEC-2025-211
The other 12 of the top 25
#PackageScoredScoreConfirmedSeverityAdvisory dateLead timeAdvisory
4opencv-contrib-python2025-05100.0%–no advisory in the 180-day window
5opencv-python2025-05100.0%–no advisory in the 180-day window
6tensorflow2025-05100.0%–no advisory in the 180-day window
7tensorflow-cpu2025-05100.0%–no advisory in the 180-day window
8gradio2025-06100.0%–no advisory in the 180-day window
11paddlepaddle2025-05100.0%–no advisory in the 180-day window
12litellm2025-05100.0%–no advisory in the 180-day window
14notebook2025-05100.0%–no advisory in the 180-day window
15opencv-contrib-python-headless2025-05100.0%–no advisory in the 180-day window
16opencv-python-headless2025-05100.0%–no advisory in the 180-day window
17langchain2025-05100.0%–no advisory in the 180-day window
24apache-airflow-providers-apache-hive2025-05100.0%–no advisory in the 180-day window

Advisory details come from OSV's PyPI advisory database (PYSEC and GHSA records); a confirmed row without details is one whose stored label was positive. Lead time = days from the scored month to publication. Package names link to their look-up on this tab.

How to read this

The mechanism replicates; the verdict on the features does not

Two things travel from Jenkins to PyPI unchanged: the label leak, which inflates stored-label results in both ecosystems by a similar margin, and the embargoed protocol that removes it. What does not travel is the verdict on any one feature family. A package's own advisory history is a strong signal on PyPI (pooled ROC-AUC 0.778) and barely clears the criterion on Jenkins (0.553). The likely reason is how advisories arrive: on PyPI they recur within the same widely used packages, while Jenkins advisories often come in ecosystem-wide batches that reach plugins with no prior record. Either way, a signal has to be re-validated in each ecosystem, which is why CANARY's contribution is the validation method rather than a feature list.

Scope of the cross-check: advisory-history inputs only (no activity clocks, contributor dynamics or install base were collected for PyPI); a download-ranked universe rather than a complete registry; and a monitoring setting, so the ranking is among packages the model has seen before, with no group split. Read the PyPI numbers as evidence that the protocol transfers, not as a PyPI risk model.