Google continues its relentless pace of innovation in the realm of artificial intelligence, unveiling two powerful new iterations of its Gemini Flash models: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. Announced on Wednesday, these releases represent a significant leap forward, offering enhanced capabilities for a wide array of applications, from complex agentic tasks and software development to the critical domain of cybersecurity. The company’s commitment to democratizing advanced AI, while simultaneously pushing the boundaries of performance, is clearly demonstrated with these latest additions to its burgeoning AI portfolio.
Sundar Pichai, CEO of Google, took to X (formerly Twitter) to herald the arrival of 3.8 Flash, proclaiming that it delivers "significant leaps" over its predecessor, Gemini 3.7 Flash. These advancements are particularly pronounced in crucial areas such as software engineering, agentic task execution, and multi-step reasoning. In a testament to its potent capabilities, 3.8 Flash reportedly outperformed numerous large frontier models on the demanding DeepSWE coding benchmark, achieving these impressive results at a substantially lower operational cost. This cost-efficiency, coupled with enhanced performance, positions Gemini 3.8 Flash as a compelling option for developers and businesses seeking to maximize their AI investments.
Complementing the general-purpose 3.8 Flash is Gemini 3.8 Flash Cyber, a specialized model engineered with a singular focus on cybersecurity. Pichai described Flash Cyber as Google’s "most capable" cybersecurity model to date, capable of matching frontier-level performance in the critical tasks of vulnerability detection and mitigation at scale. Its efficacy is underscored by its performance on industry-standard benchmarks: 86.2% on the CyberGym cybersecurity benchmark and an impressive 47.2% on CWE-Bench, which specifically evaluates an AI’s ability to generate patches for software vulnerabilities. Furthermore, an internal Google benchmark revealed that 3.8 Flash Cyber achieved a success rate exceeding 70% in discovering vulnerabilities across a broad spectrum of 20 programming languages. This remarkable proficiency suggests a profound impact on the future of digital defense.
This latest release marks Google’s third Flash model deployment in a mere six weeks, a cadence that highlights the company’s accelerated development cycle and its strategic focus on this particular class of AI models. The rapid succession of 3.7 Flash followed by 3.8 Flash and its specialized cyber variant underscores a commitment to iterative improvement and rapid deployment of cutting-edge AI technologies.
Gemini 3.8 Flash: Working "Harder" with "Greater Diligence"
The general-purpose Gemini 3.8 Flash is now readily accessible within Gemini Enterprise. Developers can also leverage its power through the Gemini API, integrating it into their workflows via Google AI Studio, Google Antigravity, and Android Studio. The model can also be utilized to generate user interfaces within Stitch. The introductory pricing remains consistent with Gemini 3.7 Flash, set at $0.75 per million input tokens and $3.75 per million output tokens. Crucially, Google has empowered users with the ability to customize and adjust model effort levels, allowing for a fine-tuned balance between quality, cost, and latency to suit specific project requirements.
For scenarios where compute efficiency is paramount, users can opt for lower token overhead. Alternatively, they can continue to utilize Gemini 3.7 Flash, which Google Senior Product Director Tulsee Doshi and Gemini Security Lead Raluca Ada Popa stated in a blog post is "fully supported for efficiency-first workloads." This flexibility ensures that users can select the model and configuration that best aligns with their operational constraints and performance objectives.
Doshi and Popa elaborated on the enhanced capabilities of 3.8 Flash, noting that it "works harder" and exhibits "greater diligence" when tackling complex tasks. This is achieved through the execution of additional reasoning steps, which, while potentially leading to slightly higher token utilization in some instances, ultimately maximizes performance. The model boasts an expansive 1 million token input window and a 64 thousand token output limit, demonstrating its capacity to process and generate substantial amounts of information. Furthermore, its multimodal capabilities extend beyond text to include images, audio, video, and PDF files, opening up a wealth of new application possibilities.
The performance of Gemini 3.8 Flash has been rigorously evaluated across a diverse array of benchmarks. These evaluations cover critical areas such as coding proficiency, multimodal understanding, computer interaction, long-context comprehension, knowledge-intensive tasks, and scientific reasoning. Google reports that the model excels in specialized knowledge domains that demand in-depth analysis and comprehensive reporting. For instance, on benchmarks like Vals Finance Agent V2 for financial analysis and Harvey’s Legal Agent Benchmark for legal applications, 3.8 Flash consistently outperformed its predecessor and other leading frontier models. Its ability to handle multi-step reasoning tasks across disciplines like mathematics, science, and humanities was further validated by a score of 54.9% on Humanity’s Last Exam (HLE)-Verified.
To illustrate the practical applications of Gemini 3.8 Flash, Google provided several compelling examples. In one instance, using Google’s Antigravity platform, the model generated a game from a simple prompt, incorporating looping techniques. This game featured puzzles, dynamic storytelling influenced by the game’s environment, and visual elements from Nano Banana to create an immersive 3D experience, depicting a wizard navigating a castle. In another demonstration, the model constructed a fully functional DOS version of Google Maps, complete with interactive locations, navigation capabilities, and street-view functionality. It also created a 3D visualizer capable of decomposing devices into layers for inspection with a slider feature, and generated a topographic map of renowned geographical sites based on real-world data from the U.S. Geological Survey, incorporating real-time cross-sections, 2D projections, and scientific explanations.
The impact of Gemini 3.8 Flash is also being recognized by independent AI benchmarking platforms. According to Arena.ai, 3.8 Flash secured the 14th position in Agent Arena, outranking DeepSeek-V4-Pro and marking a substantial improvement over Gemini 3.7 Flash, which ranked much lower at 32nd. In Text Arena, it debuted at 7th place, surpassing Claude Opus 5 and Gemini 3.7 Flash. Its performance improvements over 3.7 Flash were noted across several key areas, including multi-turn requests, writing, literature and language comprehension, longer queries, complex prompts, coding, instruction following, software and IT services, and business, management, and financial operations.
Gemini 3.8 Flash Cyber: Securing Google’s Code and Beyond
Gemini 3.8 Flash Cyber is initially being rolled out to a select group of "trusted defenders" through Google’s Fairwind Program. This initiative prioritizes government authorities, critical infrastructure operators, and other partners who require advanced cyber defense capabilities. Organizations interested in gaining access can apply through the program.
Google emphasizes that Flash Cyber has undergone "rigorous training" specifically within the cybersecurity domain, representing a "significant leap in prompt injection robustness." Its core strengths lie in autonomous vulnerability discovery and automated patching, as evidenced by internal Gemini benchmarks. Raluca Ada Popa highlighted the model’s proficiency in coding as well, underscoring its versatility. The overarching objective behind Flash Cyber’s development is to equip cybersecurity defenders with expert-level capabilities, providing them with a crucial advantage over increasingly sophisticated threat actors, whether they be malicious AI agents or human hackers. Doshi and Popa clarified, "We have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation."
This specialized model is released with a more permissive set of mitigations for cybersecurity safeguards, which is why its initial distribution is limited to trusted partners. These safeguards are designed to prevent misuse in cyber offense and in sensitive areas such as chemical, biological, radiological, and nuclear (CBRN) applications.
Google is already leveraging Gemini 3.8 Flash Cyber to enhance its own internal security. In tests concerning Chrome vulnerabilities, the model produced 2.6 times more correct patches compared to significantly larger commercial models. Wiz, a cybersecurity firm acquired by Google earlier this year for a historic $32 billion, reported that on an internal penetration testing benchmark, Gemini 3.8 Flash Cyber demonstrated a 7.5% to 9.7% higher recall of real-world vulnerabilities, while operating at 2.3 to 5.2 times lower cost than leading frontier models. Similarly, Google’s Cloud Vulnerability Research team identified a critical foundational vulnerability in less than two hours using 3.8 Flash Cyber – a task that typically takes months of research and discovery, according to Google.
Popa underscored the growing threat landscape, stating that AI agents are "incredibly skilled" at both finding and exploiting vulnerabilities. The sheer scale of modern codebases makes scanning them with large AI models an expensive undertaking, leaving defenders at a disadvantage. "In cybersecurity, attackers need only find one significant flaw over millions of lines of code. Defenders have to remove every one of those flaws to be able to defend against attackers," she explained.
Doug Turner, Engineering Director for Chrome, characterized the current landscape as a "vulnerability apocalypse" amplified by the advent of generative AI. "Simply overnight, we saw a hockey stick increase in the number of software vulnerabilities reported through our vulnerability research program," he stated in a video. He further detailed a particularly insightful vulnerability discovered by 3.8 Flash Cyber, which had resided within Chromium and Chrome for an astonishing 13 years. Turner described it as a "very subtle bug" that had eluded the attention of dozens, if not hundreds, of engineers. He concluded with optimism, stating, "Gemini 3.8 Flash Cyber is going to allow us to create better suggested fixes so that developers’ lives can get a lot easier." This sentiment reflects the profound impact these advanced AI models are poised to have on software development and cybersecurity, promising a future with more secure and efficient digital systems.

