6 min readfrom VentureBeat

Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities

Our take

Google continues to advance its AI capabilities with the release of Gemini 3.8 Flash, offering distinct models tailored for specific needs. The standard 3.8 Flash excels at agentic tasks and software development, demonstrating significant performance improvements over its predecessor and rivaling larger models at a reduced cost. Notably, Flash Cyber represents a substantial leap in cybersecurity, autonomously identifying and patching vulnerabilities with impressive efficiency—already securing Google's own code. For those exploring enterprise AI, consider “Forward-deployed engineering is how enterprise AI learns” for deeper insights.
Google’s Gemini 3.8 Flash is built for agents, while its Cyber twin hunts vulnerabilities

Google’s relentless iteration on its Flash models continues to reshape the landscape of AI-powered productivity, and the unveiling of Gemini 3.8 Flash and its specialized Cyber variant is a significant development. The sheer speed of these releases – three versions in just six weeks – underscores the company’s commitment to rapidly advancing its AI capabilities. This flurry of innovation is particularly relevant given the escalating challenges in cybersecurity, as highlighted by recent reports like the one detailing Amazon’s new scam-detection feature for Alexa PSA: Amazon’s shopping AI can now tell you if that message is a scam. The advancements in agentic tasks, software development, and crucially, vulnerability detection, demonstrate a clear focus on practical applications and tangible user benefits, aligning with our brand's ethos of accessible and empowering AI solutions. The rapid progress also builds upon broader trends in forward-deployed engineering, where continuous refinement and real-world feedback drive iterative improvements in AI models Forward-deployed engineering is how enterprise AI learns.

The performance gains touted by Pichai, particularly the outperformance of 3.8 Flash on coding benchmarks at a lower cost, are compelling. The introduction of Flash Cyber, optimized for cybersecurity, is a particularly noteworthy development. Its ability to discover vulnerabilities and generate patches at scale, significantly outperforming larger commercial models, represents a potential paradigm shift in how organizations approach threat mitigation. The fact that Google is already leveraging Flash Cyber to secure its own code provides a powerful testament to its efficacy. The story of the 13-year-old vulnerability discovered in Chromium and Chrome is a stark reminder of the challenges even seasoned engineers face, and highlights the potential for AI to augment human capabilities in this critical area. This isn't about replacing security professionals; it's about equipping them with tools that dramatically enhance their efficiency and effectiveness, something we champion as a core value. The limitations on initial access via the Fairwind Program, prioritizing government and critical infrastructure partners, are understandable given the model's enhanced prompt injection robustness and the need to carefully manage its deployment.

Beyond the immediate technical specifications, the broader implications of this release are substantial. The accelerated pace of AI model development, exemplified by Google’s Flash series, is creating a dynamic and competitive environment. The shift towards smaller, more specialized models, like Flash Cyber, represents a pragmatic approach to addressing specific challenges, rather than relying solely on massive, general-purpose models. The described capabilities of Gemini 3.8 Flash – building interactive games, generating DOS versions of Google Maps, and creating 3D visualizers – showcase the breadth of its potential applications. These examples underscore the power of AI to democratize access to sophisticated tools and capabilities, empowering users across a wide range of disciplines. The improved performance in specialized knowledge domains, as demonstrated by its success on benchmarks for finance and law, further expands its utility and reinforces its value proposition. As Anthropic continues to refine its own models, the competition will only intensify, pushing the boundaries of what’s possible Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads.

Looking ahead, the emergence of AI agents capable of autonomously identifying and patching vulnerabilities raises a critical question: how will organizations adapt their security strategies to account for this new reality? While Google emphasizes its commitment to defensive capabilities, the potential for AI to be used for malicious purposes remains a significant concern. The “vulnerability apocalypse” described by Google's engineering director suggests that the pace of discovery is accelerating, and defenders must be prepared to respond accordingly. The continued refinement of these models, coupled with increasing adoption across industries, will undoubtedly shape the future of cybersecurity and redefine the roles of human security professionals. The balance between enabling powerful AI capabilities and safeguarding against potential misuse will be a defining challenge in the years to come.

Google keeps cranking out Flash models: the company on Wednesday announced two versions of a new 3.8 Flash

The variants include a standard Flash, a “workhorse” model for agentic tasks, software development, and multi-step reasoning, and Flash Cyber optimized for vulnerability detection and mitigation.

Google CEO Sundar Pichai said in an X post that 3.8 Flash delivers “significant leaps” from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. For instance, it outperformed many large frontier models on the DeepSWE coding benchmark, at far lower cost. 

Meanwhile, Flash Cyber is the company’s “most capable” cybersecurity model, Pichai said; it also matches frontier-level performance when it comes to discovering vulnerabilities and patching them at scale. The model achieved 86.2% on the CyberGym cybersecurity benchmark and 47.2% on CWE-Bench, which evaluates AI patching abilities. In an internal Google benchmark, the model achieved a more than 70% success rate discovering vulnerabilities across 20 programming languages, Pichai said. 

3.8 is Google’s third Flash release in six weeks and comes quickly on the heels of version 3.7. 

3.8 working "harder" with "greater diligence"

3.8 Flash is available now in Gemini Enterprise; devs can try it out in the Gemini API via Google AI Studio, Google Antigravity, Android Studio, or generate UIs in Stitch. It is priced at $0.75 per million input tokens and $3.75 per million output tokens — the same introductory pricing as Gemini 3.7 Flash — and users can customize and adjust model effort levels based on their needs around quality, cost, and latency.

For instance, when compute efficiency is a priority, they can adjust to lower token overhead, or simply continue working with 3.7 Flash, which is “fully supported for efficiency-first workloads,” Google senior product director Tulsee Doshi and Gemini security lead Raluca Ada Popa wrote in a blog post

“3.8 Flash works harder,” exhibiting “greater diligence” with complex tasks like executing extra reasoning steps, although at times it may use more tokens to maximize performance, Doshi and Popa note. The model has a 1M-token input window and a 64K-token output limit, and can ingest text as well as images, audio, video, and PDF files. 

3.8 Flash was evaluated across numerous benchmarks testing coding, multimodal capabilities, computer use, long-context and knowledge work, and scientific reasoning. Google says it also does well in specialized knowledge domains requiring more in-depth analysis and reporting. For instance, the model outperformed its predecessor and other frontier models on benchmarks like Vals Finance Agent V2 for finance, and Harvey's Legal Agent Benchmark for law; it also scored 54.9% on Humanity’s Last Exam (HLE)-Verified, reflecting its ability to take on multi-step reasoning tasks across subjects like math, science, and humanities. 

In one example shared by Google, Gemini 3.8 Flash built a game with a simple prompt using looping techniques in Google’s Antigravity platform. The game uses puzzles, storytelling that changes based on the environment, and images and textures from Nano Banana to create a 3D experience (in this case a wizard navigating a castle). 

In other instances, the model created a fully-functional DOS version of Google Maps featuring interactive locations, directions, and street views; a 3D visualizer that automatically decomposed devices into layers for inspection with a slider capability; and a topographic map of famous geographical sites based on real datasets from the U.S. Geological Survey, complete with real-time cross-sections, 2D projections, and scientific explanations. 

According to Arena.ai, 3.8 Flash landed at No. 14 in Agent Arena, ranking above DeepSeek-V4-Pro, and showed a significant jump over Gemini 3.7 Flash (which sits all the way down at No. 32). It debuted at No. 7 in Text Arena, ahead of Claude Opus 5 and Gemini 3.7 Flash. It improved over 3.7 Flash in several areas: multi-turn requests, writing, literature, and language, longer queries, hard prompts, coding, instruction following, software and IT services, and business, management and financial ops. 

Flash Cyber is already securing Google's code

Flash Cyber is initially being rolled out to “trusted defenders” through Google’s Fairwind Program, which prioritizes government authorities, critical-infrastructure operators, and other partners looking for advanced cyber defense capabilities. Organizations can apply for access. 

Google says the model version has undergone “rigorous training” in the cybersecurity domain and represents a “significant leap in prompt injection robustness.” It is particularly adept at autonomous vulnerability discovery — at least, based on internal Gemini benchmarks — and automated patching. It is also very good at coding, Popa said in a video. 

The goal was to equip defenders with expert-level capabilities to give them a leg up over threat actors (whether malicious, fellow AI agents, or human hackers). “We have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation,” Doshi and Popa explain. 

The model ships a more permissive set of mitigations for cybersecurity safeguards — which is why, for now, it is only being shared with limited partners — and safeguards against misuse in cyber offense and areas like chemical, biological, radiological, and nuclear (CBRN).

Google is already using 3.8 Flash Cyber to secure its own code; it produced 2.6 times more correct patches in Chrome vulnerabilities versus much larger commercial models. 

Wiz — which Google acquired earlier this year at a historic $32 billion — reported that 3.8 Flash Cyber had 7.5% to 9.7% higher recall of real-world vulnerabilities on an internal penetration testing benchmark at 2.3 to 5.2 times lower cost than leading frontier models. Similarly, Google’s Cloud Vulnerability Research found a critical foundational vulnerability in less than 2 hours with 3.8 Flash Cyber. Typically, that research and discovery would take months, Google claims. 

AI agents are “incredibly skilled” at finding and exploiting vulnerabilities, Popa said. Scanning large codebases with big AI models is expensive, and defenders are overwhelmed. “In cybersecurity, attackers need only find one significant flaw over millions of lines of code. Defenders have to remove every one of those flaws to be able to defend against attackers.” 

Doug Turner, engineering director for Chrome, described a “vulnerability apocalypse” in recent months due to generative AI. “Simply overnight, we saw a hockey stick increase in the number of software vulnerabilities reported through our vulnerability research program,” he said in a video. 

One interesting vulnerability 3.8 Flash Cyber discovered had been in Chromium and Chrome for 13 years, he explained. It was a “very subtle bug” that dozens, if not hundreds, of engineers looked at but never flagged. “Gemini 3.8 Flash Cyber is going to allow us to create better suggested fixes so that developers’ lives can get a lot easier.” 

Read on the original site

Open the publisher's page for the full experience

View original article