Bitcoin’s AI security sprint found 6,700 issues in 55 hours, but no one knows how many are real
What happened: a turbocharged security sprint
In a wild little experiment, an AI-assisted security push aimed at Bitcoin-related projects churned out roughly 6,700 issue reports in about 55 hours. Those flags covered roughly 425 repositories, and the team labeled around 1,029 of them as high- or critical-severity. Earlier checkpoints showed the effort ramping up quickly: at about 27.5 hours it had scanned near 390 projects and produced about 4,962 findings, then kept on expanding.
The workflow looked like a tag-team: automated models blasted through code to find potential trouble spots while human specialists tuned prompts, interpreted results, tried to reproduce problems, and decided what was worth telling maintainers. In short, it wasn’t a pure robot takeover — it was an assembly line where machines did the heavy lifting and people did the picky confirmation work.
Why the big numbers are misleading (and what’s missing)
Those headline figures are exciting, but they’re also frustratingly opaque. The public updates didn’t include key follow-ups: how many reports were reliably reproduced, how many maintainers agreed a finding was a real vulnerability, how many were downgraded or rejected, and how many actually led to patches. Without those denominators and clear definitions for severity, you can’t tell signal from noise.
Some specifics were shared — for example, the scan found that about 19.5% of projects had a SECURITY.md file and around 13.1% included an email contact — but we don’t know the exact set of repositories or how those percentages were measured. There were also notes about spending (roughly $10,000 to scan 100+ repos early on, and roughly $20,000 later while scanning more repos), and mentions that a separate Coldcard issue helped kick the wider push into motion. Importantly, the sprint wasn’t credited with discovering that Coldcard flaw.
Criticism popped up too: some observers said the campaign couldn’t realistically triage its own torrent of alerts. That critique is meaningful if the team can’t turn alerts into verified fixes, but public posts didn’t include a clear list of confirmed problems or fixes, so it’s hard to measure whether the criticism sticks.
If someone wanted to be truly helpful and transparent, they’d publish how many findings were reproduced, how many were acknowledged by maintainers, how many were downgraded or rejected, and how many were fixed — plus explicit definitions of each category. That breakdown would let us judge how much of the AI-generated traffic became actual security work.
Bottom line: fast scanner, slow cure
The sprint proved one thing loud and clear: AI can flood an ecosystem-scale review pipeline in no time. But speed isn’t the same as impact. The real value depends on the portion of those 6,700 alerts that humans can validate, disclose responsibly, and convert into real patches. Think of the models as a very eager scanner and the human team as the careful gardener who decides which sprouts are weeds and which are prize bonsai.
Short version — impressive demo of scale, unclear proof of outcome. Until we get the post-sprint accounting with reproduced cases, maintainer responses, and fix counts, the headline number remains an intriguing snapshot rather than a security report card.
