Microsoft has launched its first cybersecurity-specific mannequin inside MDASH, its multi-model vulnerability identification and remediation harness.
The corporate says MDASH, utilizing MAI-Cyber-1-Flash and GPT-5.4, scored 95.95% on CyberGym. It additionally claims the configuration prices 50% lower than its present greatest MDASH mixture of GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. Entry is proscribed to accredited MDASH clients by an Azure AI Foundry non-public preview.
MAI-Cyber-1-Flash is designed to deal with as much as 90% of MDASH duties, with GPT-5.4 reserved for the toughest 10%. It’s out there solely inside MDASH, not as a standalone public mannequin or general-purpose software programming interface.
The headline rating belongs to MDASH operating MAI-Cyber-1-Flash alongside GPT-5.4, to not the brand new mannequin by itself. CyberGym Stage 1 is a known-vulnerability copy check. It provides an agent a vulnerability description and the corresponding unpatched supply code, then checks whether or not it will probably produce a working proof of idea. It doesn’t measure blind vulnerability discovery or whether or not a generated patch is right.
CyberGym’s public leaderboard didn’t checklist Microsoft’s 95.95% end result when checked on July 28, 2026. It nonetheless listed Microsoft’s Could 12 MDASH submission at 88.4%. Microsoft’s public supplies don’t say whether or not the end result was submitted for itemizing.
Microsoft’s earlier 96.55% MDASH end result doesn’t resolve the comparability. That June determine counted any crash, together with non-target vulnerabilities. The July supplies don’t say whether or not the 95.95% end result makes use of the identical criterion, so the 2 scores can’t safely be learn as a before-and-after efficiency development.
In accordance with Microsoft’s mannequin card, MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137 billion whole parameters, 5 billion lively parameters, and a 256,000-token context window. It’s a cybersecurity fine-tune of MAI-Code-1-Flash, which was developed from a MAI-Pondering-1 mid-training checkpoint.
The mannequin card says the evaluated configuration changed 80% of MDASH’s current fashions and raised the reported CyberGym end result from 88.4% to 95.95%. That 80% determine is the share of fashions changed. The separate 90% determine is the utmost share of duties Microsoft says the smaller mannequin can deal with.
Taken collectively, the disclosed design factors to routing because the central technical declare: MAI-Cyber-1-Flash is meant to deal with most duties, GPT-5.4 takes the toughest the rest, and Microsoft experiences the result on the MDASH system degree.
Microsoft’s launch announcement defines the 50% saving in opposition to its present greatest MDASH mannequin mixture of GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. The product web page individually describes the system as delivering “comparable efficiency at 50% of the price of main fashions.” The announcement and mannequin card don’t disclose the token use, name quantity, latency, activity combine, or compute allocation behind that comparability, so the determine can’t but be independently reproduced or normalised in opposition to different techniques.
“The mannequin is one enter, the system round it’s the product.”
Taesoo Kim, Microsoft’s vice chairman of agentic safety, used that distinction when describing MDASH in June. Underneath a light-weight terminal harness, the mannequin card experiences scores of 0.314 on CVEBench, 0.553 on CyberSecEval4 menace intelligence, 0.33 on its malware-analysis check, and 0.651 on CRSBench at POV=1200.
The mannequin scored zero throughout the kernel, userspace, and browser classes of ExploitGym, which asks brokers to show equipped vulnerabilities and crashing inputs into working code-execution exploits. These outcomes come from completely different duties and scoring scales, so none is a standalone CyberGym rating for MAI-Cyber-1-Flash.
Microsoft stated all benchmark testing happened in a network-isolated atmosphere with no entry to manufacturing techniques, the general public web, or exterior providers. The mannequin card additionally warns that generated textual content and code could also be inaccurate or incomplete and needs to be reviewed earlier than consequential use.
Software program vulnerability administration utilizing MAI-Cyber-1-Flash inside MDASH is the primary state of affairs Microsoft has introduced for Challenge Notion, its broader system for coordinating defensive safety brokers. Challenge Notion is scheduled to enter public preview on August 3, with Microsoft planning to increase the mannequin past software program vulnerability work to further safety workflows.



