CWG, an Israeli offensive security firm, has launched ZEUS, an agentic penetration testing platform designed to keep human testers in control of the decisions that pure automation keeps missing. The product arrives after a sharp drop in confidence in fully autonomous AI testing.
According to the Cobalt State of Pentesting Report 2026, the percentage of information security professionals who believe autonomous AI tests are reliable fell from 29% to 9% within a year. Forty-seven percent now prefer a human-in-the-loop model. Seventy-eight percent of respondents said automatic AI tools missed at least one critical vulnerability.
CWG built ZEUS on data from about 1,500 real-world penetration tests that identified more than 24,000 security vulnerabilities. In those engagements, 44% of critical findings were discovered only after pentesters moved off the automated tool path. Sixty-three percent of critical findings involved business logic flaws, multi-step attacks, or complex attack chains. Seventy-one percent of tests produced at least one significant finding that the initial scanning phase did not catch.
The company says ZEUS is meant to automate recon, asset discovery, enumeration, and initial analysis while leaving attack planning, chain building, business context analysis, validation, and final decisions with human experts. CWG projects the platform can cut recon time by 80%, shorten overall test duration by 50%, expand asset coverage by 3.5 times, and automate 85% of information gathering and early scanning work.
“ZEUS is not intended to replace the pentester, but to empower him to focus on tasks where human thinking still provides a significant advantage,” the company stated.
The launch reflects a broader shift in offensive security away from the earlier promise of fully autonomous testing and toward agentic systems that keep humans accountable for the steps that determine whether a critical vulnerability is found or missed.
Conditions Driving the Launch
Industry confidence in fully autonomous penetration testing has collapsed. The Cobalt State of Pentesting Report 2026 found that the share of security professionals who consider autonomous AI tests reliable fell from 29% to 9% in a single year.
Preference has shifted toward human oversight. Forty-seven percent of respondents in the same Cobalt report now say they prefer a human-in-the-loop model over fully autonomous testing.
Automatic tools are routinely missing critical issues. Seventy-eight percent of security professionals reported that AI-driven tools failed to identify at least one critical vulnerability in tests they observed or ran.
Complex attack patterns remain outside the reach of pure automation. Business logic flaws, multi-step attack chains, and context-dependent exploitation continue to surface primarily when human testers deviate from scripted or automated paths.
CWG’s own field data reinforced the same pattern. Across roughly 1,500 penetration tests that uncovered more than 24,000 vulnerabilities, 44% of critical findings appeared only after testers left the automated tool path.
Sixty-three percent of those critical findings involved business logic weaknesses or multi-step attack chains that initial automated scanning did not catch.
Seventy-one percent of CWG’s tests produced at least one significant finding that the first scanning phase missed, showing that volume automation alone does not equal discovery quality.
Remediation pressure is rising at the same time. Average time to fix vulnerabilities tied to LLM systems has nearly doubled compared with the prior year, increasing the cost of missed findings.
Hundreds of millions of dollars flowed into autonomous penetration testing startups over the past two years on the premise that AI could replace large parts of human testing. Results have not matched that premise for many buyers.
Offensive security teams still need speed and broader coverage. Organizations want faster recon, wider asset discovery, and shorter overall test cycles without giving up the ability to catch the issues that only experienced pentesters consistently find.
The market is moving from “replace the pentester” to “augment the pentester.” Vendors and practitioners are increasingly treating agentic systems as tools that handle scale while humans retain authority over planning, validation, and final risk judgment.
What AI Security Looked Like Before
For the last two years, a large share of investment and product messaging in offensive security centered on full autonomy. The premise was straightforward: AI agents would run penetration tests from reconnaissance through exploitation with limited human involvement. Vendors raised hundreds of millions of dollars on that vision. Buyers were told that automation would shrink test cycles, expand coverage, and reduce dependence on scarce human pentesters.
In practice, many of those systems excelled at the early, repetitive layers of work. Asset discovery, basic scanning, and enumeration improved. What did not improve at the same rate was the discovery of issues that depend on business context, multi-step logic, or creative chaining of weaknesses. Security teams increasingly reported that automated tools produced long lists of findings while still missing the vulnerabilities that mattered most.
Industry data began to reflect the gap. Confidence that autonomous AI tests were reliable fell sharply. A large share of practitioners reported that automatic tools had missed at least one critical vulnerability. Average remediation times for certain AI-related weaknesses lengthened rather than shortened. The result was a growing mismatch between the marketing claim of end-to-end autonomy and the operational reality that the highest-value findings still required experienced humans to leave the automated path.
By early 2026, the dominant model in many environments was still “run the autonomous tool first, then have people chase what it missed.” That sequence treated human expertise as a cleanup step rather than as the authority layer for the decisions that determine whether a critical risk is found.
What AI Security Looks Like Now
The market is adjusting. CWG’s launch of ZEUS is one concrete example of a broader shift toward agentic systems that keep humans in control of the steps that pure automation has repeatedly failed to handle well.
Under the new model, AI takes on the high-volume, lower-judgment work: reconnaissance, asset discovery, enumeration, information gathering, and initial scanning. Human pentesters retain responsibility for attack planning, chain construction, business-context analysis, validation of findings, and final decisions about what constitutes real risk. The platform is positioned as an accelerator for the parts of testing that scale, not as a replacement for the parts that require judgment.
The supporting evidence for this shift is no longer only anecdotal. Cobalt’s 2026 data showed trust in fully autonomous tests collapsing to 9 percent, with nearly half of respondents preferring a human-in-the-loop approach. CWG’s own dataset from roughly 1,500 engagements showed that a large share of critical findings only appeared after testers moved off automated paths, and that most of those critical issues involved business logic or multi-step chains.
The practical change is in the division of labor. Organizations that adopt this model are no longer asking whether AI can finish a penetration test alone. They are asking which parts of the test should be automated so that human attention can stay on the decisions that still determine whether the most serious vulnerabilities are identified. ZEUS is built for that division: heavy automation of recon and early analysis, explicit human authority over planning, validation, and final judgment.
Our Take
AI Security Take
The collapse in confidence in fully autonomous penetration testing is not a temporary reaction. It is evidence that the highest-value part of offensive security still depends on human judgment over business logic, multi-step attack paths, and context that pure automation has not reliably captured.
CWG’s ZEUS launch is useful because it makes the division of labor explicit. Recon, discovery, enumeration, and initial scanning are good candidates for heavy automation. Attack planning, chain construction, business-context analysis, validation, and final risk decisions are not. Treating those latter steps as optional human “cleanup” is how critical findings get missed.
Security teams evaluating agentic testing tools should ask three practical questions before buying:
Which specific phases does the system automate, and which phases remain under human authority by design?
Does the platform surface the business-logic and multi-step findings that automated scanners have historically under-detected, or does it mainly accelerate the same early-phase work?
When a critical issue is found only after a human leaves the automated path, is that treated as an expected part of the process or as an exception?
The market is moving from “can AI finish the test” to “which parts of the test should AI own so humans can stay accountable for the decisions that still determine whether serious risk is found.” Tools that blur that line will keep producing volume. Tools that enforce it will produce better coverage of the vulnerabilities that matter.
Organizations that continue to treat full autonomy as the goal, rather than controlled augmentation, will keep discovering the same gap the Cobalt data and CWG’s own engagements already documented: automation scales the easy layers and still underperforms on the hard ones.