Web interface

CANARY Web Console

Security-advisory forecasting for Jenkins plugins from public signals, evaluated honestly: four validation layers, one pre-registered criterion, and every number on this site traceable to the layer that produced it.

PyPI cross-check

Does the method travel? The same protocol on a second ecosystem

Every number on this tab comes from the embargoed rolling backtest re-run on PyPI: the download-ranked packages with a GitHub repository, OSV advisories as the outcome, and only each package's own advisory history as input. Same folds, same 180-day horizon, same label embargo at every forecast date.

Embargoed labels at every fold
Packages in the panel7,053download-ranked, with a resolvable GitHub repository
Package-months scored183,37813 embargoed folds, 2023-05 to 2025-05
Advisory outcomes2,143base rate 1.17%
Pooled ROC-AUC, PyPI0.778Advisory history · Logistic Regression
Same features, Jenkins0.553Advisory history · XGBoost; PyPI with the same model: 0.720

The package look-up shows where a package ranked at every forecast date and the OSV advisories that followed; the fold picker applies to the case study at the bottom. Try vllm for what advisory history cannot do: a package sits at the bottom of the ranking until its first advisory, then near the top.

Honest track record

Where opencv-python ranked before each window

Advisory history · Logistic Regression, embargoed rolling backtest: the rank this package held among every scored package at each forecast date, using only what was knowable then.

Recorded, not re-scored

Scored at 26 forecast dates; on average in the top 0.06% of packages. An advisory followed within 180 days of 10 of those dates.

Download rank #533 on PyPI; 37,656,040 downloads a month; 32 OSV advisories on record. PyPI · GitHub

0255075100top 20%2023-052023-082023-112024-022024-052024-082024-112025-022025-052023-05 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2023-06 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2023-07 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2023-08 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2023-09 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2023-10 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2023-11 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2023-12 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2024-01 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2024-02 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2024-03 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2024-04 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2024-05 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2024-06 · rank 5 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2024-07 · rank 6 of 7053 · 100th percentile · score 1.000 advisory followed within 180 days2024-08 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2024-09 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2024-10 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2024-11 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2024-12 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2025-01 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2025-02 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2025-03 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2025-04 · rank 6 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2025-05 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 days2025-06 · rank 5 of 7053 · 100th percentile · score 1.000 no advisory within 180 daysadvisory followed within 180 daysno advisory in window
Forecast dateScoreRankShare of panelWhat followed
2023-05100.0%5 of 7,053top 0.06%PYSEC-2023-183 2023-09-29
2023-06100.0%5 of 7,053top 0.06%PYSEC-2023-183 2023-09-29
2023-07100.0%5 of 7,053top 0.06%PYSEC-2023-183 2023-09-29
2023-08100.0%5 of 7,053top 0.06%PYSEC-2023-183 2023-09-29
2023-09100.0%5 of 7,053top 0.06%none within 180 days
2023-10100.0%5 of 7,053top 0.06%none within 180 days
2023-11100.0%5 of 7,053top 0.06%none within 180 days
2023-12100.0%5 of 7,053top 0.06%none within 180 days
2024-01100.0%5 of 7,053top 0.06%none within 180 days
2024-02100.0%5 of 7,053top 0.06%GHSA-qr4w-53vh-m672 2024-08-30
2024-03100.0%5 of 7,053top 0.06%GHSA-qr4w-53vh-m672 2024-08-30
2024-04100.0%5 of 7,053top 0.06%GHSA-qr4w-53vh-m672 2024-08-30
2024-05100.0%5 of 7,053top 0.06%GHSA-qr4w-53vh-m672 2024-08-30
2024-06100.0%5 of 7,053top 0.06%GHSA-qr4w-53vh-m672 2024-08-30
2024-07100.0%6 of 7,053top 0.07%GHSA-qr4w-53vh-m672 2024-08-30
2024-08100.0%6 of 7,053top 0.07%none within 180 days
2024-09100.0%6 of 7,053top 0.07%none within 180 days
2024-10100.0%6 of 7,053top 0.07%none within 180 days
2024-11100.0%6 of 7,053top 0.07%none within 180 days
2024-12100.0%6 of 7,053top 0.07%none within 180 days
2025-01100.0%6 of 7,053top 0.07%none within 180 days
2025-02100.0%6 of 7,053top 0.07%none within 180 days
2025-03100.0%6 of 7,053top 0.07%none within 180 days
2025-04100.0%6 of 7,053top 0.07%none within 180 days
2025-05100.0%5 of 7,053top 0.06%none within 180 days
2025-06100.0%5 of 7,053top 0.06%none within 180 days

The leak replicates

Stored labels inflate PyPI results the same way they inflated Jenkins

The PyPI advisory-history run was scored twice, with stored labels and under the embargo, exactly as the Jenkins runs were; the Jenkins pair with the same inputs is drawn beside it. Each pair is one configuration scored twice on identical test data. Stored labels let a training label be set by an advisory published inside the test window; embargoed rebuilds every training label from advisories known before the fold's forecast date. The gap is the leak. Hover a bar for the value.

ROC-AUC — separation of advisory-bound plugins from the rest; 0.5 is chance

stored labels (leaky)embargoed (honest)chance 0.50PyPI · Advisory history · XGBoostrolling, 13 folds, 2,143 positivesPyPI · Advisory history · XGBoost stored: ROC-AUC 0.7730.773PyPI · Advisory history · XGBoost embargoed: ROC-AUC 0.7200.720Jenkins · Advisory history · XGBoostrolling, 13 folds, 760 positivesJenkins · Advisory history · XGBoost stored: ROC-AUC 0.6020.602Jenkins · Advisory history · XGBoost embargoed: ROC-AUC 0.5530.553

Average precision — concentration of advisories at the top of the ranked list

stored labels (leaky)embargoed (honest)PyPI · Advisory history · XGBoostrolling, 13 folds, 2,143 positivesPyPI · Advisory history · XGBoost stored: Average precision 0.2880.288PyPI · Advisory history · XGBoost embargoed: Average precision 0.1870.187Jenkins · Advisory history · XGBoostrolling, 13 folds, 760 positivesJenkins · Advisory history · XGBoost stored: Average precision 0.0370.037Jenkins · Advisory history · XGBoost embargoed: Average precision 0.0240.024

The verdict inverts

ROC-AUC per fold, PyPI beside Jenkins, advisory history only

The same three advisory-history inputs (advisories to date, CVEs to date, worst CVSS to date) scored at the same thirteen forecast dates. On PyPI they clear the pre-registered criterion by a wide margin in every fold; on Jenkins the same inputs hover around the criterion and touch chance in some folds. Hover a point for the fold's outcome count.

0.450.500.550.600.650.700.750.800.85chance 0.50pre-registered criterion 0.552023-052023-072023-092023-112024-012024-032024-052024-072024-092024-112025-012025-032025-05PyPI · Advisory history · Logistic Regression 2023-05 · ROC-AUC 0.763 · 144 advisory outcomesPyPI · Advisory history · Logistic Regression 2023-07 · ROC-AUC 0.778 · 159 advisory outcomesPyPI · Advisory history · Logistic Regression 2023-09 · ROC-AUC 0.763 · 179 advisory outcomesPyPI · Advisory history · Logistic Regression 2023-11 · ROC-AUC 0.765 · 200 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-01 · ROC-AUC 0.765 · 207 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-03 · ROC-AUC 0.794 · 193 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-05 · ROC-AUC 0.819 · 171 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-07 · ROC-AUC 0.808 · 129 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-09 · ROC-AUC 0.822 · 147 advisory outcomesPyPI · Advisory history · Logistic Regression 2024-11 · ROC-AUC 0.816 · 137 advisory outcomesPyPI · Advisory history · Logistic Regression 2025-01 · ROC-AUC 0.813 · 148 advisory outcomesPyPI · Advisory history · Logistic Regression 2025-03 · ROC-AUC 0.798 · 153 advisory outcomesPyPI · Advisory history · Logistic Regression 2025-05 · ROC-AUC 0.788 · 176 advisory outcomesPyPI · Advisory history · XGBoost 2023-05 · ROC-AUC 0.647 · 144 advisory outcomesPyPI · Advisory history · XGBoost 2023-07 · ROC-AUC 0.665 · 159 advisory outcomesPyPI · Advisory history · XGBoost 2023-09 · ROC-AUC 0.698 · 179 advisory outcomesPyPI · Advisory history · XGBoost 2023-11 · ROC-AUC 0.695 · 200 advisory outcomesPyPI · Advisory history · XGBoost 2024-01 · ROC-AUC 0.731 · 207 advisory outcomesPyPI · Advisory history · XGBoost 2024-03 · ROC-AUC 0.741 · 193 advisory outcomesPyPI · Advisory history · XGBoost 2024-05 · ROC-AUC 0.759 · 171 advisory outcomesPyPI · Advisory history · XGBoost 2024-07 · ROC-AUC 0.745 · 129 advisory outcomesPyPI · Advisory history · XGBoost 2024-09 · ROC-AUC 0.776 · 147 advisory outcomesPyPI · Advisory history · XGBoost 2024-11 · ROC-AUC 0.753 · 137 advisory outcomesPyPI · Advisory history · XGBoost 2025-01 · ROC-AUC 0.765 · 148 advisory outcomesPyPI · Advisory history · XGBoost 2025-03 · ROC-AUC 0.771 · 153 advisory outcomesPyPI · Advisory history · XGBoost 2025-05 · ROC-AUC 0.780 · 176 advisory outcomesJenkins · Advisory history · XGBoost 2023-05 · ROC-AUC 0.553 · 132 advisory outcomesJenkins · Advisory history · XGBoost 2023-07 · ROC-AUC 0.525 · 90 advisory outcomesJenkins · Advisory history · XGBoost 2023-09 · ROC-AUC 0.595 · 67 advisory outcomesJenkins · Advisory history · XGBoost 2023-11 · ROC-AUC 0.558 · 55 advisory outcomesJenkins · Advisory history · XGBoost 2024-01 · ROC-AUC 0.583 · 40 advisory outcomesJenkins · Advisory history · XGBoost 2024-03 · ROC-AUC 0.568 · 22 advisory outcomesJenkins · Advisory history · XGBoost 2024-05 · ROC-AUC 0.608 · 23 advisory outcomesJenkins · Advisory history · XGBoost 2024-07 · ROC-AUC 0.623 · 32 advisory outcomesJenkins · Advisory history · XGBoost 2024-09 · ROC-AUC 0.575 · 42 advisory outcomesJenkins · Advisory history · XGBoost 2024-11 · ROC-AUC 0.576 · 39 advisory outcomesJenkins · Advisory history · XGBoost 2025-01 · ROC-AUC 0.525 · 66 advisory outcomesJenkins · Advisory history · XGBoost 2025-03 · ROC-AUC 0.540 · 75 advisory outcomesJenkins · Advisory history · XGBoost 2025-05 · ROC-AUC 0.515 · 77 advisory outcomesPyPI · Advisory history · Logistic RegressionPyPI · Advisory history · XGBoostJenkins · Advisory history · XGBoost

What the development ranking looks like

Advisories caught vs packages reviewed

Reviewing the top 20% of packages by score would have caught 67% of the advisories that followed across the development folds (by fold: 63%, 65%, 65%, 64%, 65%, 67%, 73%, 71%, 73%, 72%, 71%, 69%, 66%), against a base rate of 1.2%. The thick line is the pooled ranking over every development fold; thin lines are the individual folds (183,378 package-months, 2,143 advisory outcomes). The dashed diagonal is random order. Right: the ROC curve of each fold for the first configuration.

0%0%20%20%40%40%60%60%80%80%100%100%random orderAdvisory history · Logistic Regression review top 1% of plugins → 30% of advisoriesAdvisory history · Logistic Regression review top 2% of plugins → 43% of advisoriesAdvisory history · Logistic Regression review top 3% of plugins → 51% of advisoriesAdvisory history · Logistic Regression review top 5% of plugins → 61% of advisoriesAdvisory history · Logistic Regression review top 8% of plugins → 62% of advisoriesAdvisory history · Logistic Regression review top 10% of plugins → 63% of advisoriesAdvisory history · Logistic Regression review top 15% of plugins → 65% of advisoriesAdvisory history · Logistic Regression review top 20% of plugins → 67% of advisoriesAdvisory history · Logistic Regression review top 25% of plugins → 68% of advisoriesAdvisory history · Logistic Regression review top 30% of plugins → 70% of advisoriesAdvisory history · Logistic Regression review top 40% of plugins → 73% of advisoriesAdvisory history · Logistic Regression review top 50% of plugins → 77% of advisoriesAdvisory history · Logistic Regression review top 60% of plugins → 81% of advisoriesAdvisory history · Logistic Regression review top 70% of plugins → 86% of advisoriesAdvisory history · Logistic Regression review top 80% of plugins → 91% of advisoriesAdvisory history · Logistic Regression review top 90% of plugins → 96% of advisoriesAdvisory history · Logistic Regression review top 100% of plugins → 100% of advisoriestop 10% → 63% caughttop 20% → 67% caughttop 30% → 70% caughtshare of packages reviewedAdvisory history · Logistic Regression
0.00.00.20.20.40.40.60.60.80.81.01.0false-positive rate2023-05 (144 outcomes) · AUC 0.7632023-07 (159 outcomes) · AUC 0.7782023-09 (179 outcomes) · AUC 0.7632023-11 (200 outcomes) · AUC 0.7652024-01 (207 outcomes) · AUC 0.7652024-03 (193 outcomes) · AUC 0.7942024-05 (171 outcomes) · AUC 0.8192024-07 (129 outcomes) · AUC 0.8082024-09 (147 outcomes) · AUC 0.8222024-11 (137 outcomes) · AUC 0.8162025-01 (148 outcomes) · AUC 0.8132025-03 (153 outcomes) · AUC 0.7982025-05 (176 outcomes) · AUC 0.788

Development-fold outcomes

Top-25 per development fold vs. what followed

Advisory history · Logistic Regression · advisory_only_logistic. Each package's best month inside the fold is shown; training labels were rebuilt from advisories known before the fold's forecast date, and the folds were run once.

Embargoed

Across the development folds shown, 13 of the 25 top-25 slots were followed by an advisory within 180 days. That is the honest precision at 25 for this configuration: a ranking signal with real lift over the base rate, not an operational triage list.

Development fold 2025-05 (13 of 25 confirmed)

Scored: 2025-05 and 2025-06  |  Labels known as of: 2025-06  |  Fold ROC-AUC: 0.788  |  Top-25 precision: 13/25 (52%)  |  Base rate: 1.2% (41.7×)
#PackageScoredScoreConfirmedSeverityAdvisory dateLead timeAdvisory
1apache-airflow2025-05100.0%✓Critical (9.8)2025-12-17230 daysPYSEC-2025-87
2django2025-05100.0%✓Critical (9.8)2025-10-01153 daysPYSEC-2025-106
3mlflow2025-05100.0%✓High (8.1)2025-10-29181 daysGHSA-5cvj-7rg6-jggj
9ansible2025-05100.0%✓Medium (5.5)2025-12-04217 daysGHSA-8ggh-xwr9-3373
10pillow2025-05100.0%✓High (7.1)2025-07-0161 daysGHSA-xg8h-j46f-w952
13vllm2025-06100.0%✓High (8.8)2025-08-2181 daysGHSA-79j6-g2m3-jgfw
18pycti2025-06100.0%✓Medium (5.4)2025-07-1847 daysPYSEC-2025-181
19torch2025-05100.0%✓High (7.5)2025-09-25147 daysPYSEC-2025-203
20flask-appbuilder2025-06100.0%✓Medium (6.5)2025-09-11102 daysGHSA-765j-9r45-w2q2
21aiohttp2025-05100.0%✓Low2025-07-1474 daysGHSA-9548-qrrj-x5pj
22jinja22025-05100.0%✓High (7.1)2025-06-1040 daysPYSEC-2025-74
23werkzeug2025-05100.0%✓Moderate2025-12-02215 daysGHSA-hgf8-39gv-g3f2
25transformers2025-06100.0%✓High (7.8)2025-12-23205 daysPYSEC-2025-211
The other 12 of the top 25
#PackageScoredScoreConfirmedSeverityAdvisory dateLead timeAdvisory
4opencv-contrib-python2025-05100.0%–no advisory in the 180-day window
5opencv-python2025-05100.0%–no advisory in the 180-day window
6tensorflow2025-05100.0%–no advisory in the 180-day window
7tensorflow-cpu2025-05100.0%–no advisory in the 180-day window
8gradio2025-06100.0%–no advisory in the 180-day window
11paddlepaddle2025-05100.0%–no advisory in the 180-day window
12litellm2025-05100.0%–no advisory in the 180-day window
14notebook2025-05100.0%–no advisory in the 180-day window
15opencv-contrib-python-headless2025-05100.0%–no advisory in the 180-day window
16opencv-python-headless2025-05100.0%–no advisory in the 180-day window
17langchain2025-05100.0%–no advisory in the 180-day window
24apache-airflow-providers-apache-hive2025-05100.0%–no advisory in the 180-day window

Advisory details come from OSV's PyPI advisory database (PYSEC and GHSA records); a confirmed row without details is one whose stored label was positive. Lead time = days from the scored month to publication. Package names link to their look-up on this tab.

How to read this

The mechanism replicates; the verdict on the features does not

Two things travel from Jenkins to PyPI unchanged: the label leak, which inflates stored-label results in both ecosystems by a similar margin, and the embargoed protocol that removes it. What does not travel is the verdict on any one feature family. A package's own advisory history is a strong signal on PyPI (pooled ROC-AUC 0.778) and barely clears the criterion on Jenkins (0.553). The likely reason is how advisories arrive: on PyPI they recur within the same widely used packages, while Jenkins advisories often come in ecosystem-wide batches that reach plugins with no prior record. Either way, a signal has to be re-validated in each ecosystem, which is why CANARY's contribution is the validation method rather than a feature list.

Scope of the cross-check: advisory-history inputs only (no activity clocks, contributor dynamics or install base were collected for PyPI); a download-ranked universe rather than a complete registry; and a monitoring setting, so the ranking is among packages the model has seen before, with no group split. Read the PyPI numbers as evidence that the protocol transfers, not as a PyPI risk model.