NextFin News - The New York Stock Exchange told Congress on Wednesday that it used Anthropic's Project Glasswing — a controlled program built around an unreleased, vulnerability-hunting frontier model — to find and fix flaws in its cybersecurity systems, marking the first time a regulated operator of global equity-trading infrastructure has publicly confirmed that frontier artificial intelligence is now scanning its live defenses. "We're an early participant in Project Glasswing," Lynn Martin, president of NYSE Group, told the House Financial Services Committee. "We have been able to find a variety of items that we've been able to very quickly address because of the state-of-the-art technology." The disclosure is more than a cybersecurity footnote: it is the clearest signal yet that the defensive promise of frontier AI has crossed from research labs into the production environment of the world's most systemically important financial markets — and it lands just weeks before Anthropic's expected October initial public offering, where Glasswing's real-world results will be read as commercial evidence as much as a security milestone.
The Disclosure and the Program Behind It
The testimony connects two public facts. Intercontinental Exchange, the Fortune 500 parent of the NYSE, announced in June that it had deployed Anthropic's Claude Mythos Preview across all of its businesses — the NYSE and other exchanges, clearinghouses, data services, and its mortgage-technology platform — with deployment, security architecture and governance managed by ICE itself. Ben Jackson, president of ICE, framed the move as a defensive necessity: "The systems we run are the backbone of global financial markets. As part of Project Glasswing, we're advancing the use and sophistication of AI across our cyber security in a manner that is secure, auditable, and designed for regulated industries... we can detect vulnerabilities at scale and deliver the highest quality services to our customers." Martin's congressional appearance on Wednesday was the first time the exchange operator put a name and a result to that deployment.
Project Glasswing itself is a tightly controlled research program. Anthropic launched it on April 7, 2026, bringing together 11 launch partners — Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks — plus more than 40 additional organizations that build or maintain critical software. Access is invitation- and vetting-based, not a product that can be purchased. Anthropic committed up to $100 million in model-usage credits and $4 million in donations to open-source security organizations, including $2.5 million to Alpha-Omega and OpenSSF through the Linux Foundation and $1.5 million to the Apache Software Foundation. Within a month, Anthropic said its roughly 50 partners had collectively found more than 10,000 high- or critical-severity vulnerabilities; in a follow-on update the company reported scanning more than 1,000 open-source projects and estimating 6,202 high- or critical-severity flaws in them out of 23,019 total findings. By early June, third-party tracking put the program at roughly 200 organizations across more than 15 countries.
The engine behind the effort is Claude Mythos Preview, a general-purpose frontier model that Anthropic has deliberately kept unreleased. The company's stated reason is a stark one: "AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities." Mythos has already found thousands of high-severity vulnerabilities, including some in every major operating system and web browser — and some that survived decades of human review. Participants can access the model after the research preview at $25 per million input tokens and $125 per million output tokens across the Claude API, Amazon Bedrock, Google Cloud's Vertex AI and Microsoft Foundry.
Why the NYSE Case Is Different From a Routine Penetration Test
The important question is not whether the NYSE found bugs — every large operator runs penetration tests, bug bounties and code audits. The important question is what changed in the economics and tempo of finding them. For decades, vulnerability discovery was bottlenecked by scarce human expertise. A handful of elite security researchers could read code, reason about edge cases, chain partial weaknesses into a working exploit, and write it up. That scarcity is what kept the median time from first disclosure to first observed exploitation long enough — 771 days in 2018, by security-industry estimates — for defenders to patch. By 2024, that window had collapsed to single-digit hours.
Frontier models change the supply side of that equation. A model that can read an entire codebase, reason across modules, and propose reproducible exploits does not sleep, does not bill by the hour, and does not run out of senior staff. The commercial pentesting market — valued at roughly $2.4 billion to $2.7 billion in 2025 and projected to reach $5.5 billion to $7.4 billion by the early 2030s, according to market-research estimates — was built on the assumption that expert-grade attack simulation was expensive and slow. Autonomous systems are already testing that assumption. XBOW, an autonomous penetration-testing platform founded by Semmle and GitHub veteran Oege de Moor, reached the top rank on HackerOne in 2025, submitting nearly 1,060 vulnerabilities, and raised $75 million to scale a product now used by large banks and technology firms. When an autonomous system can submit more vulnerabilities in a year than most human researchers submit in a career, the market is not being disrupted at the margin — it is being repriced.
For a venue like the NYSE, the stakes are asymmetric. A vulnerability in trading, clearing or market-data infrastructure is not just a data leak; it is a potential interruption to the price-discovery mechanism for the world's largest equity market. ICE's statement that the deployment is "secure, auditable, and designed for regulated industries" is doing heavy lifting here: a regulated exchange cannot hand its code to an external black box. The governance model — ICE controls the deployment, the architecture and the audit trail — is the template that will determine whether other financial-market utilities follow.
The Second-Order Effect: Defense Is Also a Product Launch
There is a second-order consequence that the market is already pricing. Anthropic is widely expected to go public as early as October 2026, in what could be one of the largest technology listings in years. The company closed a funding round earlier this year at a $380 billion private valuation, with annualized revenue run rate reportedly around $30 billion — up from $1 billion at the end of 2024, according to company disclosures and reporting. Every Glasswing disclosure from a marquee operator like the NYSE is a reference customer that a pre-IPO company can point to when arguing that its most powerful model has real-world, mission-critical utility.
That creates an uncomfortable tension. Glasswing is framed as a defensive consortium — a way to put dangerous capabilities to work before they proliferate to actors who will not deploy them safely. But the same program also demonstrates, in public, exactly what the unreleased model can do: find flaws in every major operating system and browser, including ones that decades of human review missed. For the 200 organizations inside Glasswing, that is reassurance. For every attacker watching the congressional testimony, it is a capability advertisement. The defensive race and the offensive race are now running on the same model family, and the side that iterates faster wins.
The pricing signal matters too. At $25/$125 per million tokens, Mythos is positioned as a premium tool, not a commodity scanner. That pricing, combined with the invitation-only access, tells us how Anthropic sees the risk: this is not a capability to be democratized yet. It is a controlled capability, loaned to vetted defenders, with the company betting that the defensive value — and the commercial proof — outweighs the risk of the technique spreading.
The Counter-Thesis: AI Vulnerability Hunting Is Overhyped
The strongest argument against reading too much into the NYSE disclosure is that AI-assisted vulnerability discovery has been overhyped before, and the gap between "found a variety of items" and "materially reduced systemic risk" is wide. Security researchers have warned that frontier models produce false positives at scale, that they struggle with context outside the code they can see, and that organic model guardrails are inconsistent — the same request framed differently can produce different outcomes. A congressional soundbite about "items we've been able to very quickly address" does not tell investors how many of those items were critical, how many were low-severity noise, or whether the fixes will survive the next audit cycle.
There is also a deeper structural objection, and it has a named critic. Bruce Schneier, the veteran cryptographer and security researcher, wrote in early June that although Anthropic's program is finding large numbers of vulnerabilities, almost none of them had been patched, and that the company's refusal to release supporting data — asking the industry to "trust us" — was "a big problem." The public record so far supports the skepticism on pace: only one publicly disclosed vulnerability can be directly tied to Glasswing by name, CVE-2026-4747, a FreeBSD NFS remote-code-execution flaw that Anthropic described as fully autonomously identified and exploited. Three more — a 27-year-old OpenBSD flaw, a 16-year-old FFmpeg bug, and Linux kernel privilege-escalation chains — remained under embargo pending patches. Finding a flaw is only the first step; the security value is realized only when the flaw is fixed, and the patching pipeline, not the finding engine, may be the binding constraint.
That objection is serious, but it does not negate the directional shift. Even if AI only compresses the cycle, the defender still gets the first look at their own code — and in a world where exploitation windows run in hours, the first look is the only look that matters. Cloudflare, a Glasswing partner, reported finding 2,000 bugs — 400 of them high- or critical-severity — across its critical-path systems, with a false-positive rate its own team considered better than human testers. The NYSE's point was not that AI eliminated risk; it was that the exchange learned about its exposures faster than it otherwise would have.
What to Watch Next
Three signals will tell us whether this is a structural shift or a one-off proof of concept. First, the 90-day report Anthropic promised: within 90 days of launch, the company said it would report publicly on what it learned, including the vulnerabilities fixed and improvements made that can be disclosed — a report expected in early July 2026, with a fuller accounting of CVEs and patching progress to follow. Second, watch whether other regulated market utilities — clearinghouses, depositories, payment networks — announce their own Glasswing or equivalent deployments. If the NYSE is alone, this is a pilot. If the model becomes standard infrastructure for market utilities, it is a regime change. Third, watch Anthropic's IPO filing: the S-1 will show whether Glasswing participants convert into paying customers at the $25/$125 per-million-token price, or whether the program remains a loss-leading research exercise.
The falsifying signal is specific: if Anthropic's public accounting shows only a handful of directly attributable CVEs months after launch, with major partners reporting little material patching and high false-positive rates, then the structural-shift thesis weakens and Glasswing looks more like a pre-IPO marketing exercise than a security inflection point. Conversely, if the 90-day report shows a large volume of verified, patched vulnerabilities across critical infrastructure, the case that frontier AI has permanently lowered the cost of defense becomes much harder to argue against.
Short term, expect more congressional attention: Martin's testimony is likely to prompt questions about whether AI-vetted infrastructure should become a regulatory expectation for systemically important financial utilities. Medium term, the pentesting and managed-security vendors will have to reposition around AI-augmented workflows or lose share to autonomous platforms. Long term, the question is whether frontier models become a permanent layer of the security stack — like antivirus was in the 1990s — or whether the capability diffuses so widely that it becomes a commodity and the advantage shifts to whoever patches fastest.
The central judgment: this is structural, not cyclical. The scarcity of human expertise that protected software for decades has ended, and it will not return. The NYSE did not just find some bugs; it confirmed that the most important financial infrastructure on earth now treats frontier AI as a core defensive control. The market has not finished pricing what that means for the security industry, for critical-infrastructure operators, or for Anthropic's valuation.
"We're an early participant in Project Glasswing. We have been able to find a variety of items that we've been able to very quickly address because of the state-of-the-art technology." — Lynn Martin, president of NYSE Group, testifying before the House Financial Services Committee, September 2, 2026
The kicker: in cybersecurity, the defender only needs to be faster until the attacker is faster once. The NYSE bet that Anthropic's model buys it that speed advantage — and the rest of the market is now watching to see whether the bet pays off before the attackers get the same tool.
Explore more exclusive insights at nextfin.ai.
