Google continues its rapid iteration of its Gemini family of large language models with the Wednesday announcement of two powerful new variants: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. This latest release marks Google’s third Flash model deployment in just six weeks, underscoring the company’s aggressive pace in advancing AI capabilities. The standard 3.8 Flash model is engineered as a versatile "workhorse," designed to excel in complex agentic tasks, intricate software development, and multi-step reasoning processes. Complementing this, Flash Cyber is a specialized iteration meticulously optimized for the critical domains of vulnerability detection and proactive mitigation within cybersecurity.
Sundar Pichai, CEO of Google, heralded the arrival of 3.8 Flash on X, stating that it delivers "significant leaps" in performance compared to its predecessor, 3.7 Flash. He highlighted improvements across software engineering, agentic tasks, and multi-step reasoning, noting that 3.8 Flash has outperformed numerous large frontier models on the demanding DeepSWE coding benchmark, all while operating at a substantially lower cost. This cost-effectiveness is a crucial factor in democratizing access to advanced AI capabilities for developers and organizations.
The Flash Cyber model, according to Pichai, stands as Google’s "most capable" cybersecurity model to date. It achieves frontier-level performance in identifying and rectifying vulnerabilities at scale. Its prowess is evidenced by its impressive scores: 86.2% on the CyberGym cybersecurity benchmark and 47.2% on CWE-Bench, a benchmark that specifically evaluates AI’s ability to automate code patching. Furthermore, within an internal Google benchmark encompassing 20 distinct programming languages, 3.8 Flash Cyber demonstrated a success rate exceeding 70% in discovering vulnerabilities, showcasing its broad applicability and deep understanding of diverse coding ecosystems.
The 3.8 Flash model is now accessible within Gemini Enterprise and can be explored by developers through various Google platforms, including the Gemini API, Google AI Studio, Google Antigravity, and Android Studio. It also facilitates UI generation within Stitch. The introductory pricing mirrors that of Gemini 3.7 Flash, set at $0.75 per million input tokens and $3.75 per million output tokens. A key feature of the Gemini Flash models is their adaptability, allowing users to customize and adjust model effort levels to balance quality, cost, and latency according to their specific project requirements. For scenarios where compute efficiency is paramount, users can opt for lower token overhead. Alternatively, they can continue to leverage 3.7 Flash, which remains fully supported for efficiency-first workloads, as explained by Tulsee Doshi, Google’s senior product director, and Raluca Ada Popa, Gemini’s security lead, in their official blog post.
Doshi and Popa elaborated on the enhanced capabilities of 3.8 Flash, describing it as working "harder" with "greater diligence" on complex tasks. This translates to the model’s capacity to execute additional reasoning steps, which, while maximizing performance, may occasionally result in higher token consumption. The model boasts an expansive 1 million token input window and a 64,000 token output limit. Its multimodal capabilities are significant, allowing it to process not only text but also images, audio, video, and PDF files, offering a comprehensive data ingestion pipeline for diverse applications.
The performance of 3.8 Flash has been rigorously evaluated across a wide spectrum of benchmarks. These evaluations cover critical areas such as coding proficiency, multimodal understanding, computer usage simulations, long-context comprehension, knowledge-intensive work, and scientific reasoning. Google asserts that the model also excels in specialized knowledge domains requiring in-depth analysis and comprehensive reporting. Notably, it surpassed both its predecessor and other leading frontier models on benchmarks like Vals Finance Agent V2 (Vals.ai/benchmarks/fabv2) for financial analysis and Harvey’s Legal Agent Benchmark (Vals.ai/benchmarks/hlab) for legal applications. Its ability to tackle complex, multi-step reasoning tasks across subjects like mathematics, science, and humanities was further validated by a score of 54.9% on Humanity’s Last Exam (HLE)-Verified.
To illustrate the practical applications of 3.8 Flash, Google shared several compelling examples. In one instance, utilizing Google’s Antigravity platform, Gemini 3.8 Flash was prompted to build a game. The resulting creation featured interactive puzzles, dynamic storytelling influenced by the in-game environment, and integrated visuals and textures from Nano Banana to construct a 3D wizard navigation experience within a castle setting. In another demonstration, the model generated a fully functional DOS version of Google Maps, complete with interactive locations, navigation capabilities, and street view functionality. It also produced a 3D visualizer capable of decomposing devices into layered components for inspection, complete with a slider for granular analysis. Furthermore, it constructed a detailed topographic map of renowned geographical sites, drawing on real-world datasets from the U.S. Geological Survey. This map included dynamic cross-sections, 2D projections, and in-depth scientific explanations, showcasing its capacity for sophisticated data visualization and scientific interpretation.
The impact of 3.8 Flash on AI agent performance is already evident in industry rankings. According to Arena.ai, 3.8 Flash secured the No. 14 spot in the Agent Arena, outranking DeepSeek-V4-Pro and representing a substantial improvement over Gemini 3.7 Flash, which is positioned at No. 32. In the Text Arena, it debuted at No. 7, ahead of prominent models like Claude Opus 5 and Gemini 3.7 Flash. Its advancements over 3.7 Flash were observed across a variety of critical areas, including multi-turn requests, writing quality, literary analysis, language comprehension, handling longer queries, complex prompt execution, coding efficiency, instruction adherence, software and IT services, and business, management, and financial operations.
The specialized Gemini 3.8 Flash Cyber model is currently undergoing a phased rollout to "trusted defenders" through Google’s Fairwind Program. This initiative prioritizes government authorities, critical infrastructure operators, and other key partners seeking advanced cyber defense capabilities. Organizations interested in leveraging this technology can apply for access. Google emphasizes that Flash Cyber has undergone "rigorous training" within the cybersecurity domain, representing a "significant leap in prompt injection robustness." Its core strengths lie in autonomous vulnerability discovery and automated patching, areas where internal Gemini benchmarks show exceptional performance. Raluca Ada Popa further highlighted the model’s strong coding capabilities in a recent video presentation.
The overarching objective behind Flash Cyber’s development is to equip cybersecurity defenders with expert-level capabilities, thereby providing them with a strategic advantage against threat actors, whether they be malicious AI agents or human hackers. Doshi and Popa emphasized Google’s commitment to prioritizing vulnerability fixing over offensive capabilities like exploitation from the outset of the project. To ensure responsible deployment and mitigate potential misuse, Flash Cyber ships with a more permissive set of cybersecurity safeguards. This is why its current distribution is limited to select partners, with specific safeguards in place to prevent its application in cyber offense and sensitive areas such as chemical, biological, radiological, and nuclear (CBRN) threats.
Google is already integrating 3.8 Flash Cyber into its own security infrastructure. In its application to Chrome vulnerabilities, the model produced 2.6 times more correct patches compared to significantly larger commercial models, demonstrating its practical efficacy in securing critical software. Wiz, a cybersecurity firm acquired by Google earlier this year for a substantial $32 billion, reported that 3.8 Flash Cyber achieved a 7.5% to 9.7% higher recall of real-world vulnerabilities on an internal penetration testing benchmark, all while operating at a 2.3 to 5.2 times lower cost than leading frontier models. Similarly, Google’s Cloud Vulnerability Research team utilized 3.8 Flash Cyber to identify a critical foundational vulnerability in Chromium and Chrome in under two hours, a task that would typically take months for human researchers.
Popa underscored the "incredibly skilled" nature of AI agents in both finding and exploiting vulnerabilities. She noted the significant cost associated with scanning large codebases using large AI models and the overwhelming challenge faced by defenders. "In cybersecurity," Popa stated, "attackers need only find one significant flaw over millions of lines of code. Defenders have to remove every one of those flaws to be able to defend against attackers." This highlights the crucial need for advanced, efficient, and cost-effective tools like Flash Cyber.
Doug Turner, engineering director for Chrome, described a recent "vulnerability apocalypse" driven by the proliferation of generative AI. He observed a dramatic, "hockey stick increase" in the number of software vulnerabilities reported through Google’s vulnerability research program. He recounted a particularly noteworthy vulnerability discovered by 3.8 Flash Cyber, which had remained undetected in Chromium and Chrome for an astonishing 13 years. This "very subtle bug" had been examined by dozens, if not hundreds, of engineers over the years without being flagged. Turner expressed optimism about the future, stating, "Gemini 3.8 Flash Cyber is going to allow us to create better suggested fixes so that developers’ lives can get a lot easier." This sentiment encapsulates the transformative potential of advanced AI models in bolstering cybersecurity defenses and streamlining the development lifecycle.

