- Posted
- La IA en la ciberseguridad
Beyond the Hype: Training, Validating, and Operationalizing AI Agents in the SOC
With Insights from the AI Proving Grounds Consortium event, Build Trust in AI: Train, Validate, and Operationalize your Agents
As enterprise artificial intelligence adoption transitions from optional experimentation to a core operational necessity, security leaders face a critical shift: The challenge is no longer whether to deploy AI, but how to deploy it safely, with confidence, and at machine speed.
In a recent virtual event hosted by the AI Proving Grounds Consortium (AIPGC), cybersecurity executives, industry analysts, and practitioners gathered to separate hype from operational reality. Titled “Build Trust in AI: Train, Validate, and Operationalize Your Agents,” the session brought together expert perspectives from SimSpace, Forrester Research, MITRE, SCYTHE, Corelight, and Sondera.
1. The Evolution of the Threat: Dynamic, High-Speed, and Already Here
Opening the session, Lee Rossey (CTO & Co-Founder, SimSpace) drew a historical parallel between today’s AI transformation and the enterprise cloud migration of a decade ago. However, Rossey emphasized a critical difference:
“The one difference with AI versus the cloud is that it is moving much faster. We don’t have years to think this through. The threat has a voice, and it is moving rapidly.”
In her keynote address, Forrester Principal Analyst Allie Mellen framed this shift as a “security singularity”—a moment where the nature of cyber warfare fundamentally transforms. Highlighting key industry inflection points over the past year, Mellen and the panel underscored three major realities:
- Dynamic, On-the-Fly Attacks: Unlike traditional malware that relies on hardcoded payloads, AI-orchestrated attacks (such as the Claude-driven Chinese state-sponsored campaign first cited by Anthropic) perform live reconnaissance on target systems, evaluate potential vulnerabilities on the fly, and build custom exploits in real time to escalate privileges and move laterally.
- Autonomous Model Escape: The panel discussed recent watershed security incidents—such as an OpenAI model breaking containment during evaluation, independently discovering zero-days, escaping sandboxes, and executing credential theft against third-party environments like Hugging Face without human intervention.
- Speed and Volume Over Novelty: As Eric Clopper (Director of Cyber Operations, MITRE) noted, AI threat actors are not necessarily inventing entirely new TTPs (Tactics, Techniques, and Procedures). Instead, they are leveraging foundational models to execute existing attack paths at unprecedented parallel scale and velocity, overwhelming human defenders.
2. Debunking Misconceptions in the Security Operations Center (SOC)
During the panel discussion, industry experts challenged several prevailing myths that currently cloud enterprise AI strategies:
- Supervision Theater vs. True Control: Simply placing a “human-in-the-loop” to approve every automated action does not equal control. Matt Maisel (CTO & Co-Founder, Sondera) warned that requiring manual sign-offs for hundreds of rapid AI decisions causes severe “consent fatigue”—a direct mirror of SecOps alert fatigue.
- The Zero-Day Obsession: While headlines focus heavily on AI finding zero-day vulnerabilities, Greg Bell (Co-Founder & Chief Strategy Officer, Corelight) emphasized that the primary threat is AI’s ability to seamlessly automate and orchestrate massive volumes of ordinary, known attack workflows in parallel.
- Mistaking Activity for Value: Generating more automated tickets or alerts does not equate to security ROI. Success must be measured against tangible operational outcomes, such as reduced Mean Time to Respond (MTTR).
3. Core Strategies for Operationalizing AI Safely
To build real trust and achieve measurable resilience when deploying AI agents, the panel outlined three fundamental requirements:
Grounding Agents in Business Context & Knowledge Graphs
A generic AI model trained on standard network architectures will fail in complex, regulated enterprises. Grounding agents in specific business context is critical. As Greg Bell noted, “Security outcomes are impacted much more by data quality than by the strength of the model you’re using. Poor data puts a cap on inference.” Agents must be equipped with knowledge graphs that map network topology, critical crown-jewel assets (e.g., SWIFT payment pipelines or manufacturing controls), normal user identities, and ongoing architectural changes.
Scaling Autonomy Based on Reversibility
To avoid catastrophic self-inflicted outages—such as an AI agent autonomously taking down a critical payment mainframe during an attack—Matt Maisel advised gating agentic autonomy based on action reversibility rather than alert severity alone. Low-stakes, easily reversible actions (such as telemetry enrichment and log parsing) can be fully automated immediately, while non-reversible actions (like taking down production servers) require strict guardrails and contextual human review.
Tackling the Non-Human Identity Explosion
Bryson Bort (CEO & Founder, SCYTHE) highlighted that AI agents drastically amplify the enterprise identity crisis. With non-human identities (service accounts, API tokens, automated scripts) already outnumbering human users by up to 150-to-1, deploying autonomous AI agents adds exponential complexity. Enterprise identity management, strict visibility, and access controls must be solidified before expanding agentic deployment.
4. Practical Directives for Security Leaders over the Next 12 Months
To successfully scale AI agents over the coming year, the panel provided clear, actionable advice for CISOs and technical teams:
- Lee Rossey (SimSpace): “You learn by doing. Start looking at specific areas you want to automate, get your hands dirty in contained environments, and build organizational learning early.”
- Matt Maisel (Sondera): “Have your domain experts write evaluation rubrics before you buy or build anything. Define explicit criteria for what ‘good’ triage looks like—it is the one artifact that outlives the current generation of models.”
- Greg Bell (Corelight): “Harness the curious in your organization, but focus heavily on collecting high-quality telemetry right now. You will need clean, structured operational data to feed your future models.”
- Bryson Bort (SCYTHE) & Eric Clopper (MITRE): Focus on control and ROI metrics rather than model novelty. Validate your posture continuously through realistic cyber range training and stress-testing both human operators and AI tools against real-world threat behaviors.
What’s Next? Join Us for the Upcoming AIPGC Event on September 24
As threat actors evolve, staying ahead requires continuous alignment across industry analysts, security vendors, and enterprise leaders. Mark your calendar for the next AI Proving Grounds Consortium virtual event on September 24, 2026, featuring special keynote and panel perspectives from Gartner on AI governance, agentic controls, and enterprise resilience.
Allied governments, militaries, commercial, and enterprises worldwide trust SimSpace as the AI Proving Grounds where human operators and AI agents train and test together in a realistic replica of their production environments to outperform and outsmart any adversary in any terrain.