The Business of Cyber Security

The AI Labs in Cyber

Related: AI Security, AI for Offense, MCP & Agent Identity, The AI Deal Machine, AI Security Standards, Fable/Mythos & AI Export Controls.

In April 2026 Anthropic launched Project Glasswing and, in doing so, named a set of partners. The founding partner roster — AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, Palo Alto Networks — amounted to a list of the companies a leading model maker considered central to AI-era defense. Four weeks later OpenAI launched Daybreak, built around GPT-5.5 and its Codex agent harness, and three of the marquee names (Cisco, CrowdStrike, Palo Alto Networks) signed onto both. The frontier labs are no longer only tool suppliers to the security industry; they now also act as king-makers and category-validators, and — through the same vulnerability-discovery capability — as latent offensive actors. What follows sets out what the labs do in cyber, how they capture value, and why their moves are a leading indicator for M&A.

The three roles a frontier lab plays in cyber

A frontier lab now plays three distinct roles in cybersecurity simultaneously, and conflating them obscures the strategy. As a supplier, it sells model capacity (tokens, agent harnesses, guardrails) into every security product — the commoditized input layer. As a king-maker / category-validator, it chooses partners, funds programs, and publishes capability benchmarks that anoint incumbents and legitimize whole sub-segments — an act of distribution power no vendor can replicate. As a principal, it owns a model that can autonomously find and weaponize zero-days, which makes the lab itself a latent offensive actor and the subject of government control (see 20f). The M&A signal lives almost entirely in the middle role: when a lab crowns CrowdStrike and Palo Alto, it is telling the market where defensive budget — and acquisition currency — will concentrate.

A frontier lab plays three cyber roles at once — the M&A signal is in the middle one Supplier (commoditized) → King-maker (distribution power) → Principal (latent weapon) 1 · Supplier tokens · agent harnesses guardrails sold into every security product commoditized input — value leaks to the buyer 2 · King-maker picks partners · funds programs · publishes capability benchmarks anoints incumbents + validates categories → M&A 3 · Principal owns a model that finds & weaponizes zero-days autonomously latent weapon → government control (20f) The same vulnerability-discovery capability powers all three — defense, distribution, and weapon are one engine. Source: Anthropic Project Glasswing; OpenAI Daybreak; BIS Fable/Mythos directive. Exhibit: The Business of Cyber Security.
The lab's commercial leverage in cyber is the middle box — distribution power expressed as partner selection and category validation. That is the role that moves acquisition currency and origination timing. See [The AI Deal Machine](20c-ai-deal-machine.md) and [The Platform Wars](03m-platform-wars.md).

Anthropic — Project Glasswing (the defensive consortium)

What it is. An industry-wide AI cyber-defense program built on the Claude Mythos frontier model, which had autonomously found thousands of zero-days across every major OS and browser before public discussion. Anthropic granted partners access to Claude Mythos Preview — an unreleased model it described as having crossed the threshold where AI can surpass all but the most skilled humans at finding and exploiting software vulnerabilities — and committed up to $100M in usage credits plus $4M to open-source security organizations.

What it produced. Anthropic and roughly 50 partners used Mythos Preview to identify more than 10,000 high- or critical-severity vulnerabilities in critical software. The program then expanded (announced Jun 2, 2026) to 150 organizations across 15+ countries, spanning power, water, healthcare, communications, and hardware — and more than 40 additional infrastructure maintainers were granted model access.

The founding roster is the strategic payload: AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, Palo Alto Networks. A model vendor effectively crowned CrowdStrike and Palo Alto as the AI-era cyber leaders and validated the entire "security for AI / AI for security" thesis — while privately warning U.S. officials that uncontrolled Mythos release could make large-scale attacks "significantly more likely." That warning is the hinge into 20f.

OpenAI — Daybreak, and Aardvark → Codex Security

Daybreak (announced May 11, 2026) is OpenAI's answer: its models plus a tiered access framework plus Codex as the agentic harness plus a security-partner set. It prioritizes high-impact threats, generates and tests risks inside the enterprise with scoped access, and produces audit-ready remediation evidence — explicitly "defender-first" positioning to differentiate from the offensive-capability narrative around Mythos. Reporting notes Daybreak and Glasswing post nearly identical benchmarks and share three anchor partners (Cisco, CrowdStrike, Palo Alto) — the security incumbents are hedging across both labs rather than betting on one.

Aardvark → "Codex Security." OpenAI's agentic security researcher Aardvark became Codex Security, delivering continuous protection as code evolves and rolling out to ChatGPT Enterprise/Business/Edu. Zscaler partnered with OpenAI (Zero Trust Exchange + OpenAI-powered AI Asset Analysis for MCP/agent risk), giving OpenAI a network-security distribution beachhead that mirrors Anthropic's endpoint/platform anchors.

The Jun 23 2026 expansion — "fix, not just find," plus a channel. OpenAI's largest Daybreak push since launch reframes the program around remediation and distribution. It ships GPT-5.5-Cyber (its strongest model for finding and patching vulns — tracing reachability across large codebases, validating in controlled environments, developing/testing patches, preparing human-review evidence), an upgraded Codex Security (whole-codebase scanning → threat models → validated findings → generated patches → export to existing VM pipelines via SARIF/CodeQL), and — strategically the most important piece — a Daybreak Cyber Partner Program that lets security vendors embed GPT-5.5 with "Trusted Access for Cyber" inside their own products and services. A companion "Patch the Planet" initiative (with Trail of Bits and HackerOne) funds open-source maintainers to move from findings to fixes (30+ projects committed, incl. cURL, Go, Python, Sigstore, pyca/cryptography). The deliberate refocus from discovery to patching sharpens the defender-first contrast with the Mythos offensive narrative — and the partner program turns the king-maker role into an actual supply relationship the crowned incumbents now resell, tightening the falsifiable-bear-case squeeze below (do the labs stay suppliers, or move down the stack?).

GPT-5.6 (launched Jul 9, 2026) — cyber as the headline capability, and a government-reviewed release path. OpenAI's GPT-5.6 family — Sol (flagship), Terra (mid-tier), Luna (low-cost) — reached ChatGPT, Codex, and the API on Jul 9, 2026, with the company calling it its "strongest cybersecurity model yet, achieving frontier performance with significantly fewer tokens"; supported defensive uses include threat modeling, code review and patching, and blue teaming. OpenAI's launch benchmarks are pointed directly at Anthropic: per the company's reading of the Artificial Analysis Coding Agent Index, Sol scores 80 — 2.8 points above Claude Fable 5, using under half the output tokens at roughly one-third lower cost — with Terra placed just above Fable 5 and Luna above Opus 4.8 (vendor-published claims; no independent head-to-head yet). Pricing per million tokens: Sol $5 input / $30 output, Terra $2.50 / $15, Luna $1 / $6. The release path is as significant as the model: at the administration's request under the June 2026 AI-security executive order's voluntary prerelease-review framework (EO 14409, Regulation), OpenAI held GPT-5.6 to a limited "trusted partner" preview in late June while federal reviewers — reportedly including the Commerce Department's Center for AI Standards and Innovation — conducted additional testing before approving the wide release. A frontier model whose flagship claim is cybersecurity performance, cleared for the public through a government review channel, is the Fable/Mythos precedent (20f) operating as routine process rather than emergency directive. (TechCrunch, Jul 9 2026 · TechCrunch — limited rollout after government request, Jun 26 2026 · Engadget · OpenAI — previewing GPT-5.6 Sol)

GPT-5.6-Cyber and the two-tier Daybreak split (announced Aug 10, 2026) — an offense-grade model gated to vetted defenders. OpenAI introduced GPT-5.6-Cyber, a model built on GPT-5.6-Sol and trained for specialized offensive tasks such as finding zero-days and building exploit chains. On the vendor's own testing it completed prompts involving exploit-chain development, privilege escalation, and authentication bypass at a 95% rate, against 1.5% for GPT-5.6-Sol — a capability jump OpenAI describes as putting frontier cyber intelligence in the hands of trusted defenders before attackers can field the equivalent. Because the model carries reduced refusal safeguards for authorized exploit development, it is not offered broadly: access runs only through an expanded Daybreak program split into two tiers — Daybreak Blue (Sol and other general frontier models with guardrails tuned for defensive work) and Daybreak Red (cyber-specific models, including GPT-5.6-Cyber). Named launch partners include the large services and audit firms (Accenture, Capgemini, EY, IBM, KPMG, PwC) and security platform vendors (Palo Alto Networks, Sophos, CrowdStrike, Fortinet, Akamai, Cloudflare). The release is the clearest instance yet of a frontier lab productizing offensive cyber capability under a controlled-distribution model — the dual-use posture Anthropic frames through Mythos (20f), here delivered as a gated commercial tier rather than a research preview, and it extends the supplier-to-the-incumbents relationship (the same platform vendors resell Daybreak access) one step further up the capability curve. The first named downstream deployment followed two days later: on Aug 12, 2026 Palo Alto Networks said its Unit 42 consulting arm would run the Daybreak models inside customer environments under a service called Frontier AI Exposure Analysis, with a multi-model harness routing each task to the best-performing model and consultants validating the output — the point at which gated lab capability becomes a billable engagement rather than a partner listing (04f, 03k). (SecurityWeek, Aug 11 2026 · The Hacker News · TechCrunch, Aug 10 2026 · Axios)

Astra flagged as a potential first "critical" cyber model (Aug 7, 2026). Days before the GPT-5.6-Cyber launch, OpenAI disclosed that internal evaluations of its upcoming Astra model showed large gains in agentic coding and offensive-cybersecurity capability, and that it was treating Astra as potentially the first model to reach the "critical" cybersecurity tier of its Preparedness Framework — the level at which a model could independently discover and weaponize zero-day exploits against hardened production systems, or plan and execute an end-to-end cyberattack from a high-level objective without human direction. OpenAI stated that testing was ongoing and that it had not confirmed Astra crossed the threshold, but that in the interim it had paused non-compliant internal Astra activity and imposed additional controls (isolated testing environments, restricted access, real-time monitoring). GPT-5.6-Sol, by contrast, had peaked at the framework's "high" tier. It is the first public instance of a frontier developer applying a self-defined "critical" cyber-capability trip-wire to restrict its own model, and it sits at the boundary between the AI-for-offense capability curve tracked here and the model-control and regulatory thread on 20f and 16 — the same autonomous-offense threshold the Mythos program described crossing (above) and the trigger around which the EO 14409 prerelease-review framework is built. (SecurityWeek · CNBC, Aug 10 2026 · Forbes, Aug 9 2026)

The hyperscaler agent-security land grab

The cloud platforms are not waiting for the labs — they are racing to own the governance layer for AI agents, the highest-aggregation-risk slice of Security for AI:

Platform Move (2026) What it captures
Microsoft Agent 365 (govern AI agents); MDASH (100+ specialized agents for vuln discovery, integrated with Defender); Purview controls for coding agents (Claude Code, Copilot, Codex, OpenClaw); reportedly preparing Project Perception (see below) The agent-governance control plane, bundled into E5 — the structural bear case for every AI-security pure-play
Google Cloud Post-Wiz ($32B, completed Mar 11 2026) Threat Hunting + Detection Engineering agents (preview); Wiz integrates across agent studios Cloud-native AI security distribution at hyperscaler scale
Cloudflare Mesh (private networking for AI agents); zero-trust egress for agents The network/identity boundary for agent-to-agent traffic
AWS Security agents + Bedrock Guardrails; Glasswing founding partner Guardrails at near-zero marginal cost, bundled with inference

The bundling threat made concrete. Every capability a startup sells as "AI security" — posture, guardrails, agent identity, red-teaming — the hyperscalers and platforms are shipping as a feature at near-zero marginal cost. That is why the AI-security pool is simultaneously the highest-multiple and the highest-aggregation-risk part of the market (see AI Security and 03l).

Microsoft — Project Perception (announced Jul 18, 2026; public preview Aug 3, 2026)

The Information reported on Jul 17, 2026 that Microsoft was preparing an enterprise security product, "Project Perception," and Microsoft officially launched it on Jul 18, 2026. The product operates similarly to Anthropic's Mythos: deployed within an organization's IT environment to identify vulnerabilities and provide fixes. The product routes individual security tasks across AI models from Microsoft, OpenAI, and Anthropic, matching each task to a model rather than running everything on a single frontier model — a deliberate multi-model router strategy. Its stated positioning is price: an expected cost well below Mythos, whose estimated API pricing runs roughly double that of Opus-class models (about 100% above Opus 4.8 and about 82% above OpenAI's comparable tier, per Anthropic's published pricing as read by trade coverage). (The Information, Jul 17 2026 · TechRepublic, Jul 17 2026)

Project Perception entered public preview on Aug 3, 2026, integrated into the Microsoft Defender workflow. As shipped, it is an agentic system coordinating red, blue, and green agents — red agents run attack simulation and defense testing, blue agents triage and investigate signals, and green agents implement fixes and harden defenses — built on the MAI-Cyber-1-Flash model with humans retained in the decision loop. The preview places it alongside the agentic-SOC cohort (Agentic SOC) as a hyperscaler-shipped instance of the "AI makes products better" thread, bundled into the Defender/E5 surface rather than sold as a standalone. (Microsoft, Jul 27 2026 · Redmond, Jul 27 2026)

The product bears on two critical threads. First, it is the clearest instance to date of a platform building a Mythos-adjacent security product on top of multiple labs' models — treating frontier models as interchangeable, price-competed inputs behind a router rather than as exclusive foundations, which is the supplier-commoditization leg of the falsifiable bear case below. This is exactly the "frontier models become interchangeable infrastructure" scenario that would compress Anthropic's or OpenAI's market power as security vendors; their offensive negotiating strength (control of the unique model) dissolves if a platform can swap which model routes each task. Second, it introduces direct price competition into the frontier-AI vulnerability-discovery-and-remediation category that Mythos currently defines, alongside OpenAI's GPT-5.6 cost positioning (above) — consistent with the buyer-side expectation, recorded in a 2026 Wall Street CISO survey, that AI platform providers become security vendors within one to two years. A regulatory asymmetry is in play: a product assembled from already-released models may face a lighter path than a single restricted frontier model, though the EO 14409 prerelease-review framework (Regulation) could reach it if benchmarked cyber capability becomes the trigger.

How the labs capture value in cyber

The labs make comparatively little direct money in cybersecurity today; the programs run on credits and goodwill. Their real return is strategic leverage:

  1. Distribution power. By choosing partners and publishing benchmarks, a lab steers where enterprise security budget — and platform acquisition currency — flows. This is aggregation theory applied one layer up: the lab aggregates the capability the whole industry now depends on.
  2. Category validation. When Anthropic and OpenAI both stand up cyber programs, they convert "AI security" from a speculative line item into a board-level mandate, pulling forward demand across 03l, Agentic SOC, and 20b.
  3. Talent and data flywheel. Vulnerability findings from these programs feed back into model training and into the labs' own security posture — a compounding moat.
  4. Optionality on becoming the platform. The unresolved question (the falsifiable thesis below): do the labs stay suppliers, or do they move down the stack into security products and disintermediate the very incumbents they crowned?

Falsifiable bear case

The king-maker dynamic is bullish for incumbent platforms and AI-security M&A only if three things hold, and each is contestable. (1) The labs stay suppliers. If Anthropic or OpenAI ships its own security product surface, the crowned incumbents become resellers of a commoditizing input — the labs disintermediate them. (2) The capability stays scarce. Arora's "five years of bugs in six weeks" cuts both ways; if Mythos-class offensive capability proliferates to open models (he estimates Chinese open models ~6 months behind), the defensive premium compresses and the consortium's edge erodes. (3) Governance doesn't freeze the capability. The Fable/Mythos export action (20f) shows the government can switch the engine off; a capability that can be disabled by directive is a fragile foundation for a category. If any of the three breaks, the labs' "validation" reprices from a durable moat to a temporary subsidy.

/ angle

Read the roster as a buyer-intent map. A lab's partner list and a hyperscaler's agent-security shipping cadence are pre-declared acquisition appetite. CrowdStrike, Palo Alto, Cisco, Microsoft, and Google have each told the market — via these programs — exactly where their AI-security gaps are. That is a buy-side targeting signal you can act on before the deal is sourced.

Sell-side origination. Founder-led AI-security pure-plays (posture, guardrails, agent identity, red-team) sitting in a category a frontier lab just validated, with declared strategic acquirers and a closing scarcity window, are textbook sell-side mandates — the window shuts as the platforms aggregate. See The Graduating Class.


Sources: Anthropic — Project Glasswing · Anthropic — Expanding Project Glasswing · CyberScoop — tech giants launch Glasswing · CSO Online — 10,000 vulnerabilities · CNBC — Mythos expands to 150 orgs (Jun 2 2026) · OpenAI Daybreak · OpenAI — Daybreak: securing the world (Jun 23 2026 expansion) · Help Net Security — OpenAI expands Daybreak (Jun 23 2026) · SecurityWeek — OpenAI refocuses on patching over discovery · The New Stack — Daybreak vs Glasswing shared partners · OpenAI — Aardvark · Zscaler + OpenAI · All-In — Nikesh Arora "5 years of bugs in 6 weeks"

China matches the frontier on cyber bug-finding (Zhipu GLM-5.2)

The Wall Street Journal (Raffaele Huang, Robert McMillan, Amrith Ramkumar; Jun 28, 2026 — "China Has Matched Anthropic in Cybersecurity, Resetting AI Race") reported that Chinese AI has matched U.S. frontier models on a security-critical task. China's Zhipu AI (Z.ai) released GLM-5.2 in June 2026 that matches the latest U.S. models at finding security bugs — in some benchmarks besting Anthropic's Claude Opus 4.8, and, with further prompting, GLM-5.2 and Opus 4.8 can approach Mythos at bug-finding (GLM-5.2 still lags U.S. models on other tasks). The framing: the U.S. clampdown on a leading U.S. lab (the Fable/Mythos export directive — see 20f) plus continued AI-chip exports to China is fueling concern that Washington is handing Beijing a cyber-offense advantage — resetting rather than protecting the AI race.

The parity claim was contested — read it narrowly. The headline outran the evidence. The WSJ's only substantive comparison was a single line — "when given further instructions, Opus 4.8 and GLM-5.2 can match Mythos in bug-finding ability, according to researchers" — i.e. bug-finding on a narrow benchmark with prompting, not a head-to-head with the restricted model and not general capability. Researchers pushed back publicly: AI-policy analyst Miles Brundage called the equivalence unproven and staked a 100-to-1 wager on then-forthcoming UK AISI cyber-range results, which GLM-5.2 had not yet publicly faced.

Two independent evaluations have since quantified the gap. The UK AI Security Institute published its first public open-weight cyber assessment on Jul 17, 2026, finding GLM-5.2 the most cyber-capable open-weight model it had tested: on its narrow cyber-task suite the model performed comparably to closed frontier models released about four months earlier (Claude Opus 4.6, GPT-5.3-Codex), and on its long-horizon cyber ranges it reached roughly the level of Opus 4.5 (released under seven months before it) — an overall lag of 4 to 7 months, narrower than the 6-to-10-month gap AISI had measured through most of 2025. At advertised prices a 100-million-token cyber-range run cost about $46 on GLM-5.2 versus about $85 on the Opus comparators. On Aug 2, 2026 the non-profit SaferAI published a separate evaluation (run from the public API with no developer cooperation) placing GLM-5.2's offensive-cyber capability 2 to 4 months behind the frontier: near-saturation on the Cybench benchmark, within the 95% confidence intervals of Opus 4.7 and GPT-5.5, with its CyberGym reproduction rate rising from about 37% to 76% as the per-task token budget increased from 2M to 50M — consistent with AISI's finding that cyber capability scales with inference budget. Both evaluations make the same secondary point, which matters more than the exact gap: because GLM-5.2 is open-weight, its cyber capabilities are not gated behind the content filtering and refusal training that closed frontier models apply — SaferAI reported that Opus 4.7 declined its CyberGym tasks while GLM-5.2 refused none — so an open-weight model of comparable capability carries higher misuse risk than an API-gated one, and that gap cannot be re-closed once weights are public. The disciplined read: "matched the frontier" overstates it, but a capable offensive-cyber model now runs on open weights a few months behind the closed frontier and without removable safeguards (UK AISI, Jul 17 2026 · SaferAI, Aug 2 2026).

Why it matters (business / M&A): (1) offensive-cyber capability in frontier models is now contested and multipolar, not a U.S. monopoly — weakening the "U.S.-allied vs. sovereign-substitute" bifurcation thesis from the supply side (cf. [28 book — AI as Strategic Asset], 14d national cyber powers); (2) it hardens the Security-for-AI demand case (defenders must assume capable adversary models) and the AI-for-Security arms race ("AI must fight AI"); (3) it elevates export-control exposure and model provenance as a diligence axis for any AI-dependent cyber target. Source: WSJ via AIC (Jun 28 2026).


Updated 2026-08-16 18:13 UTC · © El Dorado Capital · el-doradocapital.com · Market intelligence for informational purposes only; not investment advice.