Microsoft Threat Intelligence has issued a warning about a new campaign by the threat group Storm 1175. This group is now targeting organizations that use autonomous agents in their main business processes. These agents perform complex tasks by interacting with various software and databases. Storm 1175 uses a new type of exploit designed to change how these automated systems make decisions. By intercepting instructions sent to the agents, the attackers can redirect critical actions to serve their own purposes. This development shows the new security challenges businesses face as they move from traditional software to increasingly dynamic self-managing digital systems.  

Deconstructing the Technique of Logic Manipulation 

Stomp 1175 mainly uses a method called instruction injection. In this attack, the hacker adds harmful commands to the data that a digital agent is configured to handle. Since these agents are built to be helpful, they might interpret hidden commands as genuine user requests. For instance, a customer service agent could be tricked into sending out sensitive records while trying to help with a normal question. The agent simply follows its programming, not realizing the command came from an attacker. This method bypasses standard security systems, which usually look for malicious code rather than malicious instructions.  

Storm 1175 also targets the CAP knowledge and the CAP retrieval systems that these agents use. Most autonomous systems get their information from a CAP internal CAP library to answer questions or perform tasks. Attackers try to disrupt these libraries with false information or logic traps. When an agent uses this breached data, it can cause a series of actions that benefit the attacker, such as lowering security settings or giving the attacker temporary admin access. Since the agent behaves like a regular user, these actions often go undetected by standard security tools.  

Exploiting Autonomous Connectivity and Integration 

One major risk in the Storm 1175 campaign is the high level of connectivity between digital assistants, which grants access to business applications such as email, finance, and project management tools. While this linkage is helpful, it means that if one agent is compromised, damage can spread widely. Storm 1175 can use an agent’s real credentials to move across the network, acting as a trusted insider and causing harm without triggering standard malware detection.  

Microsoft’s research shows that Storm 1175 is especially focused on the supply chain of these automated systems. They go after third-party companies that create the logic framework for digital agents. If they compromise just one provider, Storm 1175 could affect hundreds of organizations at once. This hub-and-spoke attack method is highly efficient for them, enabling them to reach many targets with little effort. Companies should check the security of their automation partners as carefully as they do for standard software vendors.  

Strengthening The Defensive Parameter For Automated Platforms 

To defend against Storm 1175, Microsoft recommends a zero-trust approach to agent permissions. Digital agents should only get the minimum access they need to do their jobs. This principle of least privilege means that even if an agent is compromised, it cannot cause much harm. Also, any high-impact actions by an agent should require a clear human confirmation. This extra step helps catch risky actions, such as deleting data or transferring large files, so the system does not follow harmful instructions without someone checking first.  

Adding instruction filtering at the gateway is another important defense. This means using a separate, tightly controlled system to check the inputs sent to the main agent. The filter looks for signs of instruction injection and blocks suspicious commands before they reach the agent’s core logic. Microsoft also recommends setting up behavioral baselines for each agent. For example, if an agent that usually handles HR tasks suddenly tries to access financial files, the system should immediately trigger a security lockdown. This quick response helps catch compromised logic before it leads to a serious breach.  

Monitoring The Developing Threat Horizon 

Storm 1175’s actions signal a shift from targeting people to targeting machines with social engineering. Instead of tricking users, attackers now manipulate automated agents. This development requires new tools for monitoring how agents make decisions. Old logs showing file access can’t show the full picture. Security teams must trace the logic flow behind each action to identify which instructions were used to compromise the system.  

Microsoft is working with global partners to create a standard registry of known logic exploits for real-time sharing of Storm 1175’s tactics. Like a virus definition file, this registry helps automated systems spot and block harmful instructions, aiming to build collective immunity so attacks become harder and costlier, deterring groups like Storm 1175.  

Establishing a Standardized Registry for Autonomous Defense 

As digital systems become more integrated into corporate infrastructure, organizations are evolving their security structures to adapt. Network environments now require constant monitoring and defense. Security is increasingly defined by the strength and consistency of logical protections, not just by password security. In the future, effective logic-based defenses will reduce fears about hidden attackers. Security will rely on dependable processes that maintain system integrity. Companies will benefit from persistent, logic-driven security that continuously verifies and protects digital operations.

Source: Storm-1175 focuses gaze on vulnerable web-facing assets in high-tempo Medusa ransomware operations 

Organizations must now maintain essential data and operational control in the cloud amid new regulations, higher resilience standards, and rapid technology evolution.  

In June 2025, Microsoft CEO Satya Nadella introduced solutions through Microsoft Sovereign Cloud to address these challenges. We continually strengthen our approach to sovereignty, ensuring we meet customer needs and comply with regulations for both our sovereign public and private clouds. Today, we announce new features that enhance our security and digital sovereignty controls, offer advanced AI, and provide broader scale, supported by local partner experts. Key updates include:  

  • End-to-end AI data processing in Europe as part of the EU (European Union) data boundary, which means data processed by artificial intelligence stays completely within the borders of the European Union.  
  • Microsoft 365 Copilot now offers in-country processing for Copilot interactions in 15 countries. Details are available on the Microsoft 365 blog.  
  • Expansion of the Sovereign Landing Zones service, which are pre-configured secure cloud environments set up according to specific sovereignty requirements. Now, Microsoft Azure Local (a locally operated version of Azure) also supports disconnected operations, allowing these systems to run without an active internet connection.  
  • Microsoft 365 Local is now generally available.  
  • Azure Local, a version of Microsoft’s cloud platform operated in specific locations for greater data control, now supports a greater maximum number of servers, external SAN (storage area network, a type of shared data storage), and the latest NVIDIA GPUs (graphics processing units used for complex computing tasks like AI).  
  • Our partner digital sovereignty specialization is now available.  

Microsoft Sovereign Cloud: Continuous Innovation 

Our latest updates deliver new digital sovereignty features in AI, security, and productivity. More enhancements are coming soon to better support customers’ sovereign cloud needs.  

We know that ongoing innovation is important, and we have started putting many of our promises into action. As of this month, we have:  

  • Established a European board of directors composed of European nationals exclusively overseeing all data center operations in compliance with European law, thereby putting Europe’s cloud infrastructure into the hands of Europeans.  
  • Increased European data center capacity with recent launches in Austria and an upcoming launch in Belgium this month  
  • Expanded open source investment through funding secure open source software (OSS) projects and collaborations, as well as publishing  
  • AI access principles that widen safe, responsible access to advanced AI, helping European developers, startups, and enterprises compete more effectively across the region  
  • Advance our European security program by providing AI-powered intelligence and cybersecurity capacity-building initiatives to strengthen Europe’s digital resilience against threat actors.  

Building on our sovereign efforts, we are now launching new Sovereign Public Cloud and AI capabilities to further strengthen compliance and control. 

Organizations need comprehensive sovereignty solutions that enable compliance and control from the start of their planning.  

EU Data Boundary Includes AI Data Processing Residency 

We are keeping our promises regarding AI data processing by ensuring that data processed by AI services for EU customers remains within the European Union unless the customer asks otherwise.  

This means that all customer data, whether stored or in transit, will be kept and processed only in the EU. We use strict controls and clear processes to meet EU customer requirements.  

Expanding Microsoft 365 Copilot In-Country Data Processing To 15 Countries. 

After years of investing in global infrastructure and strong data residency, Microsoft will now provide in-country data processing for Microsoft 365 Copilot interactions in 15 countries worldwide.  

By the end of 2025, customers in Australia, India, Japan, and the United Kingdom will be able to have their Microsoft 365 Copilot interactions processed in their own country. In 2026, we will add this option for customers in 11 more countries, including Canada, Germany, Italy, Malaysia, Poland, South Africa, Spain, Sweden, Switzerland, the United Arab Emirates, and the United States.  

New Sovereign Landing Zone (SLZ) Foundation 

We are also launching an updated sovereign landing zone (SLZ) built on the trusted Azua landing zone (ALZ) foundation.  

The sovereign landing zone is our recommended setup for customers who want to use sovereign controls in the Azure private cloud.  

The refresh of the sovereign landing zone includes:  

  • Updated management group hierarchy and supporting Azure policy definitions, initiatives, and assignments to help implement the sovereign public cloud controls.  
  • We provide guidance on where to deploy Azure Key Vault managed by HSM (Hardware Security Module, a dedicated device for securely managing cryptographic keys), if needed, as part of level two sovereign controls.  
  • Deployment is easier now with the Azure Landing Zone Accelerator and Azure Landing Zone Library. For more details, see the Sovereign Landing Zone (SLZ) implementation options.  

In the coming months, we will add more built-in Azure policy definitions, initiatives, and assignments to the sovereign landing zone. This will help customers set up sovereign controls in the public cloud more quickly.  

Using sovereign landing zones gives customers a clear structure that speeds compliance with local sovereignty rules and simplifies policy management. It also helps organizations scale their workloads across Azure regions while remaining aligned with regulations and maintaining consistent operations.  

New Sovereign Private Cloud and AI Capabilities 

As organizations prioritize sovereignty, balancing compliance and innovation is crucial. Our updates merge advanced AI and scalable infrastructure across public and private clouds.  

Supporting Thousands Of AI Models On Azure Local With NVIDIA RTX GPUs. 

We are improving our sovereign private cloud with Azure Local, introducing a new Azure option that leverages the latest NVIDIA RTX PRO 6000 Blackwell Server Edition GPU for high-performance AI workloads in secure environments.  

This GPU can run over 2,000 models, including GPT, OSS, DeepSeek V3, Mistral, NeMo, and Llama 4 Maverick. It enables organizations to accelerate AI projects securely in a private cloud, supporting innovation and the adoption of advanced solutions while ensuring strong data protection and compliance.  

Customers can access thousands of ready-to-use open source AI models for tasks such as generative AI, analytics, and real-time decision-making, all with strong governance.  

Increasing Azure Local Scale to Hundreds of Servers 

Previously, Azure Local supported clusters of up to 16 servers. With our latest updates, it can now handle hundreds of servers. This change helps organizations with large or growing needs run bigger and more complex workloads, scale easily, and meet security and sovereignty requirements in Europe and worldwide.  

SAN Support On Azure Local 

One important update is that Azure Local now supports storage area networks (SANs), specialized, high-speed networks that provide access to consolidated, block-level data storage. Customers can securely connect their current on-premise storage to Azure Local, making it easier to use their existing storage while taking advantage of cloud services. This helps keep data in the right location and gives European businesses greater flexibility to comply with local data rules without sacrificing performance or control.  

Microsoft 365 Local: General Availability of Key Workloads 

Another key update is that Microsoft 365 Local is now generally available. This brings core tools like Exchange Server (for email), SharePoint Server (for document management), and Skype for Business Server (for communications) directly to Azure Local. Starting in December, customers can use these tools on Azure Local in connected mode, with a fully isolated option coming early in 2026. This setup lets organizations maintain full control while meeting strict compliance and data residency requirements.  

Disconnected Operations: General Availability 

Microsoft’s Sovereign Private Cloud brings sovereignty principles to dedicated environments for organizations with strict compliance and control needs using Azure Local. Azure Local lets government agencies, global companies, and regulated groups keep local control while still using Microsoft’s global cloud platform.  

Disconnected operations for Azure Local, available in early 2026, let customers manage multiple on-premise clusters from a single control system. Organizations can securely run private cloud operations independently, ensuring business continuity even in remote settings.  

New Partner: Digital Sovereignty Specialization Now Available 

We are launching the Digital Sovereignty Specialization in the Microsoft AI Cloud Partner Program. This specialization enables partners to demonstrate expertise in secure, compliant, and sovereign cloud solutions for Azure and Microsoft 365. Partners who earn this badge show they can meet strict data privacy and regulatory standards, supporting customer control and innovation. The specialization includes rigorous audits and offers benefits such as increased visibility, special recognition, and priority access to sovereign cloud projects.  

Looking Ahead: Advancing Sovereignty Through Greater Controls 

The Microsoft Sovereign Cloud Roadmap will introduce new capabilities to address evolving customer needs, including:  

Sovereign Private Cloud 

  • Enhanced change controls: We will introduce a set of configurable policies and approval workflows that empower organizations to exercise explicit oversight over changes propagating from the cloud to the edge, strengthening governance and compliance.  
  • Site-to-site disaster recovery: Azure site recovery in Azure local helps maintain business continuity by keeping business apps and workloads running during outages.  
  • Moving from hybrid to fully disconnected: Azure Local enables customers to transition workloads from hybrid to fully disconnected operations, providing flexibility for business continuity.  

National Partner Clouds 

National partner clouds are a key part of our sovereign cloud strategy. They offer independent cloud environments that deliver Microsoft Azure and Microsoft 365, all under local ownership and control.  

  • Delos Cloud is designed to meet the German government’s BSI cloud platform requirements.  
  • Bleu is designed to meet the French government’s ANSSI SecNumCloud requirements.  

For many public sector organizations, ERP is a critical workload that requires modernization to cloud environments. SAP is planning to deploy its RISE with SAP offering on Microsoft Azure for both Bleu and Delos cloud customers. In addition to supporting RISE with SAP for customers using Microsoft Azure public cloud deployments.  

Learn More About Microsoft’s Sovereign Solutions 

Microsoft offers leading sovereign solutions, including a flexible public cloud, a private cloud that grows with your business, and national partner clouds built for specific compliance needs. We are committed to ongoing investment and innovation so our customers can achieve sovereignty without compromise.  

Find out more about the latest in cloud innovation this November at Microsoft Ignite. Learn more and sign up today.

Source: Microsoft strengthens sovereign cloud capabilities with new services 

NASA is moving forward faster than ever on developing autonomous rovers capable of operating on the lunar surface, paving the way for independent robotic operations in future lunar missions. To achieve this goal, NASA is outfitting these new rovers with next-level artificial intelligence capabilities that enable them to navigate terrain, identify features, and execute mission tasks with minimal human intervention.  

This work aligns with NASA’s broader objectives to build out the Artemis program and establish a long-term human presence on the Moon. It is anticipated that autonomous robotic systems will play an integral role in achieving this goal, enabling exploration, data gathering, and infrastructure assembly in environments where humans will have limited ability to control them.  

Moving Toward Autonomous Lunar Exploration  

Traditional rovers have operated primarily via commands from Earth, with operators remotely controlling the rovers’ movement and conducting scientific activities. Although relatively short, the time required for a command to travel from Earth to the moon and back imposes limits on the rover’s real-time response.  

NASA intends to equip rovers with AI-driven autonomy so they can make decisions based on their immediate environment. For instance, this will allow the rover to determine if there are obstacles along its path and adjust it accordingly, as well as to select individual science targets to prioritize without waiting for a “go ahead” from mission control.  

NASA’s movement toward implementing autonomous systems within the next decade is consistent with the broader evolution of robotics in exploration today, as robots become increasingly independent due to the increasing complexity of missions.  

AI-Driven Navigation and Terrain Analysis  

The use of machine learning verifiably evaluates terrain on the moon in real time as it travels, detecting potential hazards that could obstruct its path. As it travels, a rover will use onboard cameras and other sensors to evaluate the conditions of the surface it is traveling on. It can then use this information to update and reroute itself accordingly.    

The Rover’s ability to dynamically change its route as it travels is essential, given the moon’s tough, often unpredictable terrain, which can change dramatically from one area to another over just a few hundred meters.  

Supporting Human Missions Through Robotics  

Autonomous rovers serve two purposes: they function as exploration tools that lay the groundwork for future human missions. The autonomous rovers will first explore potential landing sites to identify resource locations and establish infrastructure before crewed spacecraft arrive, as they will provide support to astronauts. 

The discovery of water-ice deposits by rover missions will enable humans to establish a sustainable, long-term presence on the Moon. Rovers will provide assistance through their capabilities to transport supplies, execute repairs, and track environmental conditions. 

NASA envisions an environment where humans and autonomous rovers work together to accomplish mission goals.  

Reducing Dependence on Earth-Based Control  

Reduced dependence on continual communication with Earth is one of the main benefits of autonomous systems. Although the delay between Earth and the moon is shorter than for missions to farther destinations, limited communication windows and available bandwidth create communication issues in both cases.  

The ability to conduct long-duration writing and autonomous rover activities, no matter where they are located, means there will be no delay in undertaking their mission activities, regardless of length or location.  

The ability to sustain continuous operations will enable greater productivity and effectiveness in the overall mission compared to traditional means. 

Integration with Broader Lunar Infrastructure  

NASA is developing autonomous rover systems that can work together as part of the broader lunar infrastructure, which includes orbiting spacecraft, surface base habitats, and communications networks. The rover systems will enable rovers on the lunar surface to communicate, exchange information, and coordinate their actions with each other and with other spacecraft and surface-based missions’ goals simultaneously.  

Networks will also improve the efficiency of lunar exploration by enabling the rovers to send their collected scientific data back to Earth or to orbiting resupply platforms. Similarly, data collected by the moon’s surface and by other surrounding spacecraft through the lunar infrastructural network(s) will send commands and updates necessary for operating in that lunar environment, as well as improved rover operations on the moon and elsewhere in space.  

NASA is designing this new lunar infrastructure to ensure continued sustainability, so the robotic systems that support the lunar exploration effort will continue to support future exploration efforts.  

Challenges in Autonomous Space Robotics  

While progress has been made toward the development of autonomous rovers, several challenges remain before they are fully successful. The environment on the moon is very hostile; temperatures range from extreme heat to extreme cold. There is high exposure to radiation and fine dust that can disrupt mechanical and electrical components.  

AI systems must also be highly robust; any failure in navigation or decision-making can negatively impact the success of future lunar missions. For AI systems to succeed, they must undergo extensive testing and validation to demonstrate they can consistently achieve the desired results in the real world.  

NASA is continuing to improve its technology through simulation, field testing, and incremental mission deployments.  

The Role of AI in Space Exploration  

AI has emerged as an important part of most current space exploration projects. They are used to process and interpret data, adapt to environmental changes, and carry out complex tasks without direct human involvement.  

For AI-enabled robotic vehicles designed for planetary surfaces, AI will assist with both navigation and the scientific analysis of terrestrial materials. A robotic system will be able to identify points of interest (POIs) and conduct scientific experiments much more efficiently than a system without AI. This will greatly increase the ability to obtain scientific data from each robotic mission.  

NASA is investing in AI to achieve its long-term exploration objectives.  

Conclusion: Robots Leading the Way to the Moon  

NASA’s development of autonomous rover systems is a key milestone in our efforts to explore the Moon’s surface. Through its ability to operate autonomously, the rover will lay the groundwork for a more prolonged human presence on the Moon and greater mission efficiency.  

Increasingly, AI-based systems will become essential for exploring and understanding the Moon and the solar system.

Source: NASA News Release 

As organizations rapidly adopt AI, safeguarding these advances is mission-critical. Google Cloud empowers you to securely develop and deploy AI, addressing compliance and privacy from the start.  

Today, we’re introducing a solution to manage risk throughout the AI lifecycle. AI protection is a tool set designed to secure your AI workloads and data across any cloud or model, regardless of platform.  

AI protection helps teams manage AI risk in several ways:  

  • It discovers AI assets in your environment and checks them for possible vulnerabilities.  
  • It secures AI assets using controls, policies, and guardrails.  
  • It manages threats to AI systems with tools for direction, investigation, and response.  

AI Protection integrates with the Security Command Center to manage security risks across clouds. This provides security teams with a unified view for monitoring AI and cloud risks.  

Discovering AI Inventory 

Managing AI risk begins with knowing where and how AI is used. Our tools automatically find and catalog models, applications, data, and their connections.  

Understanding the data supporting AI applications and protecting that data is critical. Sensitive Data Protection (SDP) identifies and secures sensitive information, now automating data discovery for Vertex AI datasets. SDP displays sensitivity and types of training data, as well as data profiles for deeper insights.  

Once sensitive data locations are identified, AI Protection leverages SCC’s virtual red teaming to detect risky combinations and potential attack paths, and to recommend steps to strengthen security.  

Securing AI Assets 

Model Armor, an AI protection feature, is now available. Model Armor protects AI models against certain attack types, including prompt injection (manipulating AI responses by inserting malicious input), jailbreak (bypassing restrictions on AI behavior), data loss, malicious URLs (web addresses leading to harmful sites), and offensive content. Model Armor works with many models across different clouds, so you get consistent protection for your models and platforms, even if your needs change later.  

Developers can now add Model Armor’s prompt and response screening automatic checks for inappropriate, harmful, or unsafe inputs and outputs to their applications using a REST API (a way for applications to communicate over the web) or by integrating with Apigee (an API management platform). Soon, you’ll be able to use Model Armor inline without changing your apps, thanks to upcoming integrations with Vertex AI and our cloud networking products.  

We are using Model Armor not only because it provides robust protection against prompt injections, jailbreaks, and sensitive data leaks, but also because it helps us achieve a unified security posture through the Security Command Center. We can quickly identify, prioritize, and respond to potential vulnerabilities without impacting the experience of our development teams or the apps themselves. We view Model Armor as critical to safeguarding our AI applications and to centralizing the monitoring of AI security threats alongside our other security findings within SCC. It is a game changer,” said Jay DePaul, Chief Cybersecurity and Technology Risk Officer, Dun & Bradstreet.  

Organizations can use AI protection to enhance the security of Vertex AI applications by applying security postures in the Security Command Center. These controls are built on a deep understanding of Vertex AI’s design, helping you set secure configurations and prevent unwanted changes.  

Managing AI Threats 

AI protection uses security intelligence and research from Google and Mandiant to help protect your AI systems. Security Command Center detectors can spot initial access attempts, privilege escalation, and persistence threats in AI workloads. New detectors based on the latest intelligence, including those for model hijacking, will be available soon.  

“As AI-driven solutions become increasingly commonplace, securing AI systems is paramount and surpasses basic data protection. AI security – by its virtue – necessitates a holistic strategy that includes model integrity, data provenance, compliance, and robust governance,” said Dr. Grace Trinidad, Research Director, IDC.  

Piecemeal solutions can leave critical vulnerabilities exposed, rendering organizations susceptible to threats such as adversarial attacks or data poisoning, and adding to the overwhelming security challenges that security teams already face. A comprehensive lifecycle-focused approach enables organizations to effectively mitigate the multifaceted risks posed by generative AI and manage increasingly complex security workloads. By taking a holistic approach to AI protection, Google Cloud simplifies and thus improves the experience of securing AI for customers,” she said.  

Enhance AI Protection With Expert Support. 

The Mandiant AI security consulting portfolio helps organizations assess and strengthen the security of AI systems across multiple clouds and platforms. Our consultants review your entire AI setup and suggest ways to enhance its security. They also offer red teaming for AI using insights from the latest real-world attacks.  

Building on a Secure Foundation 

Customers can benefit from running AI workloads on Google Cloud’s secure-by-design infrastructure, which features safeguards, encryption, and strict supply chain controls.  

If your AI workloads are regulated, assured workloads create environments with strict policy guardrails, such as data residency, which ensures your data stays within a specified location, and customer-managed encryption, which means you control the encryption keys for your data. Audit Manager demonstrates compliance with regulations and new AI standards by providing reports and evidence of adherence. Confidential computing protects data during processing; this means data remains encrypted and inaccessible to unauthorized parties even from users with system access or internal threats.  

If you want to find unsanctioned or shadow AI use in your workforce, Chrome Enterprise Premium can help. It gives you visibility into end-user activity and helps prevent both accidental and intentional leaks of sensitive data in generative AI applications.  

Next Steps 

Google Cloud remains dedicated to supporting organizations in protecting AI innovations. Additional information is available in the showcase paper from Enterprise Strategy Group and at the online security talks event on March 12th.  

To try AI protection in the Security Command Center or learn about subscription options, contact a Google Cloud sales representative or an authorized partner.  

More exciting capabilities are coming soon, and we will share in-depth details on AI protection and how Google Cloud can help you securely develop and deploy AI solutions at Google Cloud Next in Las Vegas, April 9 to April 11.

Source: Announcing AI Protection: Security for the AI era 

The Cybersecurity and Infrastructure Security Agency expanded its Secure by Design initiative in April 2026 by adding new international technology partners and software makers. This shift aims to move cybersecurity responsibility from end users to original developers. The agency urges safety features to be built in from the start, motivating companies to prioritize long-term security over rapid product launches. As digital threats grow more complex, the program aims to strengthen global infrastructure by addressing weaknesses before they reach users.  

Institutionalizing Foundational Software Integrity 

The main idea behind the Secure by Design expansion is to move forward with default safety configurations. In the past, many business applications came with open settings that IT teams had to secure. Now, the new partners promise to deliver products with strong security features, such as multi-factor authentication and encrypted communication turned on by default. This helps organizations that may not have expert staff to set up complex systems, making it harder for attackers to find easy ways in and reinforcing foundational integrity. The initiative also addresses vulnerabilities at the code level.  

This effort also covers memory safety in the code itself. Many modern security problems come from poor memory management in older programming languages. CIS says new partners are promising to use memory-safe languages or hardware protections for all new critical infrastructure. This change addresses the main cause of many zero-day attacks affecting today’s networks. By fixing these issues at the source, the industry is creating a firmer and more reliable foundation. This active approach responds to repeated failures seen in the software supply chain over the past ten years.  

Accountability Through Radical Transparency 

A key part of the expanded program is self-attestation of security practices. Manufacturers are now expected to share detailed public information about their internal testing and how they handle vulnerabilities. They will also regularly publish a software bill of materials so customers can see which third-party libraries are used in their products. This level of transparency helps organizations better judge their risks when new vulnerabilities are found. It moves away from the old black-box approach and encourages mutual knowledge and shared responsibility to sustain the impact of these improvements. New tools have been introduced for ongoing progress.  

To keep the program moving forward, CISA has set up a progress reporting dashboard for its voluntary partners. This tool checks how well companies are adopting key security measures, such as removing default passwords and enabling automatic updates. Instead of acting as a strict regulator, the agency serves as a strategic facilitator, helping companies match their business goals with national security needs. By delivering a clear plan for improvement, CISA helps its partners stand out in a market that values strong cybersecurity.  

This voluntary approach motivates companies to compete by offering better security. Also, because the software supply chain is inherently global, a vulnerability in a component developed in one country can have domino effects on critical infrastructure halfway across the planet. By harmonizing secure-by-design standards across jurisdictions, CISA and its international counterparts are creating a common front against transnational digital threats. The global baseline ensures a high standard of protection is maintained, regardless of where the software was originally authored.  

International teamwork also enables real-time sharing of threat intelligence among all program members. If one partner finds a new type of attack, they can quickly alert the whole group. This shared defense lets producers issue fixes before a local problem spreads worldwide. Expanding the program to include telecommunications is especially important because these networks are the main channels for digital information. Protecting them at the design stage benefits everyone who depends on the internet for daily life and business. As strong foundations are built, the program also addresses challenges posed by outdated platforms.  

Eliminating the Security Debt of Legacy Systems 

A major challenge the new partnerships address is technical debt in old systems. Many organizations still use legacy software built before modern online threats. CISA’s partners are creating hardening kits to enhance the security of these platforms. This helps key sectors like energy and healthcare improve defenses without replacing costly infrastructure, bridging old and new systems for a secure future.  

These kits use virtual patching and active monitoring to protect older applications. This creates a zero-trust setup in which every action is checked, even if the soft- first software was not designed for this level of security. The goal is to build fail-safe systems that limit damage in the event of a breach. By planning for potential compromises, the secure-by-design approach focuses on containing problems and on quick recovery. This practical approach to risk management recognizes the complexity of modern networks and offers a clear path to a safer future. As these upgrades take effect, the program’s vision comes into sharper focus.  

The Crystalline Vibration Of A Secure Future 

As these digital systems adopt new standards, we are quietly entering a new phase of security. Our digital world is becoming more attentive and reliable, working in step with our need for safety. Soon, software updates may be something to look forward to, showing that our systems are always learning and improving. Over time, worries about hidden flaws may fade, replaced by confidence that our most important systems are well-protected. We may find that security is handled behind the scenes by smart technology, giving us peace of mind that our digital lives are safe and valued. The world is becoming more responsive, always ready to protect us from new threats.

Source: Secure by Design 

We are opening our Meta operating system to third-party hardware makers, giving consumers more options and expanding the developer ecosystem. By partnering with leading technology companies, we aim to create a more open computing platform for the Metaverse and make app development and audience engagement easier than ever.  

Introducing Meta Horizon OS 

This new hardware ecosystem will use Meta Horizon OS, the mixed-reality operating system that powers our Meta Quest headsets. We picked this name to show our focus on people, connection, and the social network that brings everyone together. Meta Horizon OS brings together the main technologies behind today’s mixed reality experiences and adds features that make social presence a key part of the platform.  

Meta Horizon OS is the result of 10 years of work at Meta to create a next-generation computing platform to lead the way in standalone headsets. We built technologies like inside-out tracking for more natural interactions and social presence. We developed eye, face, hand, and body tracking for mixed reality. We created a comprehensive set of tools to blend the digital and physical worlds, including high resolution, passthrough, scene interpretation, and spatial anchors. This long-term effort, which started with the Android open source project, has led to a mixed reality operating system now used by millions.  

With Meta Horizon OS, developers and creators can unlock these exhilarating technologies using the custom frameworks and tools we designed for immersive mixed reality experiences. They can grow communities and businesses through content discovery and monetization features built into the OS such as the Meta Quest Store, which houses the world’s largest library of immersive apps and experiences. We’re excited to rename the Meta Horizon Store to reflect the vibrant future we’re building together.  

The Horizon social layer that powers Meta Quest devices is now experiencing an exciting part of this new ecosystem. It enables people to bring their identities, avatars, and friend groups with them across virtual spaces and allows developers to infuse their apps with meaningful social features. Because this social layer weaves together different platforms, it empowers people to connect and spend time in virtual worlds using mixed reality, mobile, and desktop devices. Meta Horizon OS devices will also leverage the same mobile companion app that Meta Quest owners enjoy today, and which we’re excited to rename the Meta Horizon app.  

A New Generation of Hardware 

As the mixed reality market accelerates and people embrace it for gaming, entertainment, fitness, productivity, and social presence, exciting new possibilities are emerging for specialized hardware. Like with PCs and smartphones, consumers will benefit the most from a thriving range of hardware, from versatile models to focused devices, all powered by the same dynamic platform.  

Top technology companies around the world are enthusiastically developing new devices using Meta Horizon OS:  

  • ASUS’s Republic of Gamers will use its first gaming expertise to create a new top-tier gaming headset.  
  • Lenovo is harnessing its experience co-designing the Oculus Lift S and its strong background in building devices like the ThinkPad laptops to deliver groundbreaking mixed reality devices for productivity, learning, and entertainment.  
  • Last year, Xbox and Meta partnered to bring Xbox Cloud Gaming (beta) to Meta Quest, letting people enjoy Xbox games on an expansive 2D virtual screen in mixed reality. Now, we’re excited to join forces again to launch a limited edition Meta Quest inspired by Xbox.  

All of these innovative devices will benefit from our strong partnership with Qualcomm Technologies Inc., maker of Snapdragon processors that integrate seamlessly with our software and hardware. The latest Snapdragon XR2 Gen 2 platform launching with Meta Quest 3 unlocks major new levels of mixed reality performance. Companies building for this vibrant new system ecosystem can tap into these advanced chipsets and custom software features.  

A More Open App Ecosystem 

As we open Meta Horizon OS to more device makers, we’re giving app developers easier ways to re-reach audiences. We’re merging the Meta Horizon Store and App Lab, so any developer meeting technical and content standards can launch software. Soon, App Lab titles will have their own section in the store, making them easier to find. Several popular apps like Gorilla Tag and Jib Class began in App Lab. We’re streamlining the process for developers to launch apps on our platform.  

We are developing a new spatial app framework to help mobile developers build mixed reality experiences. To get started, visit our application page via the provided link and request access. Use the tools you already know to bring your apps to Meta Horizon OS or create new mixed reality apps.  

Meta Horizon OS will have a more open app store. This gives people more ways to access apps. Users are not limited to our app store. They can enjoy content from services like Xbox Game Pass Ultimate or Steam Link. They can also use Air Link to stream PC software to their headsets. We invite the Google Play 2D app store to join Meta Horizon OS with the same economic model it uses on other platforms.

Source: A New Era for Mixed Reality 

NVIDIA has introduced a new high-performance interconnect standard that connects multiple graphics processors into one powerful system for local workstations. Launched in April 2026, this hardware and software solution targets professionals who need substantial parallel processing power without using cloud data centers by linking the memory and processing cores of several cards. A single workstation can handle datasets that were too large for a standard desktop. This development meets the growing need for detailed simulations and complex data processing at the network’s edge. It signals a return to decentralized, high-performance computing for researchers, engineers, and digital artists.  

Overcoming The Bottlenecks Of Traditional Bus Architecture 

One of the main technical challenges for local multiprocessor systems is communication latency between processors. Standard motherboard slots often can’t move data quickly enough to keep several high-end chips working together smoothly. NVIDIA’s new unified memory bridge fixes this with a dedicated high-speed connection that skips the usual system bus. This lets two or more processors share their memory as if it were one large pool. As a result, data doesn’t have to be copied between cards. Each calculation cycle is much faster.  

The architectural shift is supported by a new “Dynamic Load Balancer” embedded within the driver stack. This is a system monitor. This new design introduces a dynamic load balancer built into the driver software. It keeps track of each core’s workload in real time and automatically shifts tasks so no single processor slows down the group. If one unit finishes early, it takes on more work from the shared queue to help the others. This setup means that adding more cards almost doubles or triples the system’s output. Such efficiency is especially important for tasks such as real-time 3D rendering or processing large genomic data sets. It effectively merges the video memory of all linked units. In the past, if a single task required 48 gigabytes of memory but each card had only 24, the task could not run locally. The new linking removes this physical boundary, allowing the software to see a 96-gigabyte or 192-gigabyte memory space. This is a game-changer for those working with high-resolution 3D environments or large-scale statistical models. It allows for more complex textures and more detailed physics simulations without slowing down or crashing the systems.  

To handle larger memory, NVIDIA has added Predictive Data Prefetching, a feature that anticipates what data will be needed next and loads it into the high-speed cache (temporary memory used to store frequently accessed data) before processing. This way, the processing cores are never left waiting for data from slower storage devices (such as hard drives or SSDs). By keeping the compute pipeline (the sequence of processing setups) full, the system reaches speeds that once required liquid-cooled server racks (large industrial computer setups). Now, a single professional workstation can match the performance of a mid-sized server cluster (a group of connected servers) from just a few years ago.  

Thermal Management in High-Density Workstations. 

Putting several high-power processors in a single case creates significant thermal challenges that can slow performance. The new linking standard addresses this with a synchronized cooling protocol that coordinates all system fans. The hardware works together to direct airflow and move heat away from the chips and out of the case. If one card gets hotter than the others, the system can lower its speed slightly and raise a neighbor’s speed to keep overall performance steady. This thermal load-sharing prevents any single part from overheating.  

For people working in quiet offices, the system includes an Acoustic Optimization mode, which sets all fans to lower speeds to move more air without producing high-pitched noise. This reduces the typical sound produced by powerful cooling units. As a result, the workstation stays cool and quiet even during long processing sessions. By focusing on the physical environment, the company shows it understands that noise and heat matter in real-world workspaces.  

Security and Data Sovereignty at the Local Edge 

One of the main reasons for the shift to local hardware is the growing concern about data sovereignty and the safeguarding of intellectual property in the cloud. Many organizations are reluctant to upload proprietary designs or sensitive customer data to remote servers. By keeping large workloads local, NVIDIA helps create a stronger barrier against unauthorized access or the exposure of confidential data. The new multi-unit bridge uses hardware-based encryption for all data moving between processors, ensuring information remains secure even as it travels within the computer.  

In addition to the enhanced security and data control provided by local hardware, using a local system also avoids the cost of moving large amounts of data to and from a cloud provider. For example, a research lab that processes daily satellite images or medical scans can save significant money. Local hardware also offers predictable performance, so users are not affected by changing internet speeds or the impact of other users on shared cloud resources. This gives professionals full control over their computing environment and helps ensure that important deadlines are met even if a remote service goes down.  

Thermal Synchronization And Acoustic Load Balancing 

As workstations around the world become more linked and effective, digital labs are quietly changing. Offices are becoming more responsive to our creative needs. Power is shifting away from a central location, and each desk can now be a productive hub. Over time, the line between what the machine does and what we create may blur, allowing us to work more smoothly. We may soon find that our work is supported by reliable systems that esteem both our intentions and our data. The workstation is becoming more than just equipment it is now a dependable part of our daily work. 

Source: NVIDIA AI Ecosystem Expands as Marvell Joins Forces Through NVLink Fusion 

Apple introduced new updates throughout its platforms to give users more control over their data. Private Cloud Compute, a feature that processes information on remote Apple servers without storing it long-term, brings the iPhone’s strong privacy protections to the cloud so users can get both smart features and privacy. New tools such as locked and hidden apps which require authentication for app access and conceal selected apps help secure sensitive information on devices. Other updates include privacy-focused features in Mail (which limit email tracking), satellite messaging (allowing texts in areas without cell service), and Presenter Preview (a preview before sharing your screen).  

Private cloud compute allows Apple Intelligence to process complex user requests with groundbreaking privacy,” said Craig Federighi, Apple’s senior vice president of software engineering. “We’ve extended iPhone’s industry-leading security to the cloud with what we believe is the most advanced security architecture ever deployed for cloud AI at scale. Private Cloud Compute uses your data only to fulfill your request and never stores it, ensuring it’s never accessible to anyone, including Apple. And we’ve designed the system so that independent experts can verify these protections.”  

Superior Privacy for AI Capabilities 

Apple Intelligence is a personal intelligence system built into iPhone, iPad, and Mac. It uses advanced generative models to make these devices more helpful and enjoyable to use.  

A key part of Apple Intelligence is on-device processing, which means features are powered directly on the user’s device without collecting user data. When more computing is needed, private cloud computing steps in and uses larger server-based models software that runs on powerful remote computers to handle complex tasks while still protecting customer privacy.  

When a user makes a request, Apple Intelligence checks whether it can handle it locally on the device. If the task is too complex, only the necessary data is sent to Apple Silicon servers using private cloud compute. The data is not shared, stored, or shared with Apple, and is used only to complete the request.  

Apple silicon servers that power Private Cloud Compute provide strong cloud security. The Secure Enclave protects important encryption keys by keeping them isolated from the rest of the system. Secure Boot is a feature that ensures only approved and verified software can run on the server’s operating system, as it does on an iPhone. Trusted Execution Monitor is a security tool that ensures only approved code runs on the servers. Attestation allows devices to verify a server’s identity before sending any data. Independent experts can review the server code to confirm Apple’s privacy claims.  

More Privacy Features Intended To Support Users 

Locked and hidden apps help users keep their information private when sharing their screen or device. Users can lock an app to protect its content or hide it so others can’t see it. If someone tries to open a locked app, they must use Face ID, Touch ID, or a passcode. Hidden apps are moved to a special folder that also requires authentication to open.  

“We relentlessly deliver on our pledge to give users the strongest and most innovative privacy protections,” said Eric Neuenschwander, Apple’s Director of Customer Privacy. “This year is no exception, and the ability to lock and hide apps is just one example of Apple helping users remain in control of their information, even if they are sharing their devices with others.”  

Apple has long worked to let users control what they share and with whom. In 2020, the Photos picker allowed users to select specific photos for apps without granting full access. This year, new features have been built on that. Contacts permission improvements in iOS 18 let users pick which contacts to share with an app. The Accessory Setup Kit gives developers a way to pair accessories without apps, see all devices on the network, and keep things private and easy.  

Other updates throughout Apple’s platform make it even easier for users to use privacy and security features.  

The new Passwords app builds on Keychain, which Apple introduced over 25 years ago. It lets users easily access account passwords, passkeys, Wi-Fi passwords, and two-factor codes stored securely. The app also warns users about weak, reused, or leaked passwords.  

Additional Features Built With Privacy By Design 

Apple has added privacy and security protections to its apps and services for years, and iOS 18, iPadOS 18, and macOS Sequoia sustain this approach.  

In iOS 18, Mail now sorts messages directly on the user’s iPhone into primary promotions, transactions, and updates, helping users focus on what matters most.  

With iOS 18, users can send messages to friends and family over satellite when they don’t have cellular or Wi-Fi. They can use their regular iMessage and SMS conversations, and iMessage stays end-to-end encrypted.  

Presenter preview in macOS Sonoma helps users avoid sharing too much during video calls, AirPlay, or when connecting with a cable. In apps like FaceTime and Zoom, users can choose to share their entire screen or just one app, and the presenter preview appears automatically.  

Availability 

Access the developer betas of iOS 18, iPadOS 18, and macOS Sequoia now at developer.apple.com if you are an Apple Developer Program member. Expect public betas at beta.apple.com next month. Receive the updated software this fall as a free update. Be aware that features may change and may not be available everywhere in every language or on all devices. Check apple.com for detailed availability.  

Use Apple Intelligence in beta this fall on iPhone 15 Pro, iPhone 15 Pro Max, or any iPad or Mac model with an M1 chip or newer as part of iOS 18, iPadOS 18, and macOS Sequoia. Set Siri’s language to US English to enable it. Get more information at apple.com/apple-intelligence.

Source: Apple extends its privacy leadership with new updates across its platforms 

We’re excited to announce that new Azure Cobalt 100-based virtual machines (VMs) are now generally available. These VMs use Microsoft’s first sixty-four-bit Arm-based Azure Cobalt 100 CPU designed in-house. This launch is a major step forward in building and improving our cloud infrastructure with careful optimization at every level. Through integrating hardware and software, Azure Cobalt 100-based VMs highlight our efforts to deliver the right balance of performance, power efficiency, and scale for our customers.  

The Cobalt 100-based VMs include our new general-purpose DPS v6 series and DPLS v6 series, as well as the memory-optimized EPS v6 series. They deliver up to 50% better price-to-performance than our previous ARM-based VMs, making them a strong choice for many Linux-based workloads, such as data analytics, web and app servers, open-source databases, and caches.  

The new Azure Cobalt 100-based VMs offer significant improvements over previous Azure ARM-based VMs: up to 1.4 times better per-CPU performance, 1.5 times better Java workload performance, and double the performance for web server .NET apps and in-memory cache apps. NVMe local storage IOPS increase fourfold, and network bandwidth grows up to 1.5 times.  

These new VMs are available in regions like Canada Central, Central US, East US 2, East US, Germany West Central, Japan East, Mexico Central, North Europe, Southeast Asia, Sweden Central, Switzerland North, UAE North, West Europe, and West US 2. Additional regions are coming in 2024 and beyond, including Australia East, Brazil South, France Central, India Central, South Central US, UK South, West US 3, and West US.  

Customer Adoption and Scenarios 

During the preview, we worked with both internal and external customers. For example, IC3, the platform behind Microsoft Teams conversations, now serves its growing user base more efficiently and has seen up to 45% better performance on Cobalt 100-based VMs  

We are also providing Cobalt 100-based VMs to many independent software vendors (ISVs) who offer PaaS and SaaS solutions on Microsoft Azure.  

The Journey to ARM: Adopting Innovation and Customer Benefits.  

Microsoft’s experience with Arm technology shaped data center scale industry standards and earned industry recognition. Our transition to Arm-based VMs is driven by the goal of improving price performance and power efficiency for our customers, as demonstrated by the Cobalt 100-based VMs.  

Developer Ecosystem 

The developer ecosystem is growing quickly and has made great progress in recent years. Major platforms and languages such as C++, .NET, and Java now offer native ARM versions. We have made ARM-specific improvements for each of these, enabling us to fully leverage the strengths of the ARM architecture.  

Many popular infrastructure and deployment tools now support Arm natively. GitHub Actions, which many developers use for continuous integration and delivery, is now available for Arm in two ways: self-hosted runners running on an Arm VM or local Arm hardware, and GitHub-hosted runners.  

Containers are a popular choice for deployment because they deliver a streamlined workflow, isolation, security, efficient resource use, portability, and reproducibility. Microsoft Azure Kubernetes Service (AKS) now lets you create ARM agent nodes and mix ARM and x86 nodes within the same cluster.  

Specifications 

You can choose from several Azure virtual machines with 3 memory ratios per vCPU size, giving you the flexibility to meet your workload, CPU, and memory needs. All VM series are available with or without local disks, so you can select the best fit. New Dpsv6 series and Dpdsv6 series general-purpose VMs offer up to 96 vCPUs and 384 GiB of RAM. They are ideal for scale-out workloads, cloud-native solutions such as AKS, small to medium-sized open-source databases, application servers, and web servers. ARM developers can use these VMs in CI/CD pipelines, development, and test scenarios.  

  • The new Dpslsv6 and Dpldsv6 series VMs provide up to 96 virtual CPUs (vCPUs) and 192 GiB of RAM, with a 2:1 memory-to-vCPU ratio (2 GiB RAM per vCPU). They are ideal for media encoding, small databases, gaming servers, microservices, and workloads that do not require much RAM per vCPU.  
  • The new Eps v6 and Epds v6 series memory-optimized VMs provide up to 96 vCPUs and 672 GiB of RAM with an 80:1 memory-to-CPU ratio. They are built for memory-intensive work, such as large databases and in-memory CA. The new Epsv6 and Epdsv6 series memory-optimized VMs provide up to 96 vCPUs and 672 GiB of RAM, with an 8.1:1 memory-to-CPU ratio. Disk storage. For more details about disk types and where they are available, see Azure Managed Disk Types. Disk storage is billed separately from VMs. You can deploy these VMs using the Azure portal, SDKs, APIs, PowerShell, and/or the command line interface.  

To find out more about the new Azua Cobalt 100-based VMs, please read the documentation.  

Pricing 

To learn more about the pricing of Azure Cobalt 100-based VMs, please visit the Azure Virtual Machine pricing and pricing calculator pages.  

You can save money with reserved instances. The Azure savings plan for compute and spot virtual machines. Reserved VM instances help lower costs and make budgeting easier with one-year or three-year commitments. For a limited time, you can save up to fifteen percent more on one-year Azure reserved VM instances for select Linux VMs from October one, twenty twenty-four, to thirty-one March, twenty twenty-five. The Azure savings plan for compute lets you save across several Azure services, including VMs. Spot virtual machines can also cut costs for workloads that can handle interruptions and variable timing.  

A New Era of Price, Performance, and Power Efficiency. 

The launch of Azure Cobalt Boost VMs denotes a new chapter for Azure’s infrastructure. Our custom silicon program delivers outstanding price-performance and power efficiency to our customers. We look forward to seeing how these innovations help your business and to supplying even better solutions in the future.  

Thank you for taking part in this exciting trip with us.

SourceAzure Cobalt 100-based Virtual Machines are now generally available 

The OpenAI Model Spec is the main guide for how OpenAI expects its models to behave in ChatGPT and the API. It explains how to handle conflicting instructions, set boundaries, and deal with risky situations and sensitive topics. It also outlines default behaviors such as honesty, factuality, personality, and style. We use it as a guiding reference. We continue to improve our systems to better align with these guidelines. The Model Spec is a living document. It evolves as we receive community feedback and discover new situations that require clear rules.  

Last year, we open-sourced the model spec and an initial set of evaluation prompts. We are now releasing the first full version of model spec evals, a new evaluation suite that measures how well models adhere to the model spec. This makes model behavior easier for the community to understand, predict, and review.  

To understand the full breadth of the models’ alignment with these principles, model spec evals track progress across all the spec’s goals. They work alongside our detailed safety and capability evaluations, which we have used for a long time to guide model release decisions and share through our system cards. While our safety process assesses system harms and ways to reduce them, model spec evals focus on measuring ideal behavior, including the character, tone, and approach we want our models to exhibit.  

Backed by the CAP model, CAP spec, and CAP evals, we observe specific advances in each new generation of models. With this new evaluation suite, we see that GPT-5 and later models follow the model spec more closely than earlier models. Compliance rates are 72% for GPT 4o, 80% for OpenAI o3, and 82% for GPT 5 Instant. GPT 5 Thinking achieves 89%, GPT 5.3 Instant scores 84%, and GPT 5.4 Thinking 87%. Compliance generally improves with each new model. Thinking models tend to be more compliant than instant models released at the same time. We have seen better results from following instructions, reducing damaging content, handling sensitive situations, being honest and transparent, and producing higher-quality work. Some improvement is expected because the model spec has changed since older models were trained. However, the results also show real progress in alignment. Earlier reasoning models like OpenAI o3 and GPT 5 Thinking were more compliant than non-reasoning models. The latest GPT-5 models now score in the mid- to high-80s. These results cover several recent models, including GPT-4, OpenAI O3, GPT-5 Instant, GPT-5 Thinking, GPT-5.3 Instant, and GPT-5.4 Thinking.  

  • To support these evaluations, we have created an evaluation data set with 596 prompts designed to test how models handle tone, refusals of harmful requests, explanatory questions, sensitive topics, and more.  
  • Additionally, as part of this release, we are providing open-source evaluation code so researchers can develop and reproduce our results, extend the dataset, or adapt it to their own use cases. This transparency further encourages community involvement in improving the evaluation process.  

The OpenAI model spec is meant to provide a clear, shared guide for how OpenAI models should behave. These evaluations show where current models match the specification and where improvements are still needed. They also help the research community study model behavior and give useful feedback for further improvement.  

The evaluation prompts currently cover only text-based parts of the model spec. We plan to add prompts for images and agentic settings in the future. For now, we measure those areas internally with other evaluations. The model spec covers a lot, but our current set of prompts is small compared to its full scope. This means it provides a broad, low-detail view of how well models conform to the spec. We focused on covering more areas because we already have other evaluations that examine specific cases in greater detail. In future releases, we plan to add more detailed prompts to improve this evaluation. The current examples are based on simple, everyday user scenarios, not on adversarial or tricky prompts. We aim to increase the number, variety, difficulty, and realism of prompts in future updates. Model spec evals are a living dataset that evolves as the spec changes. We plan expansions to cover the current spec. We also expect the dataset to change as we add new policies or add nuance to existing ones.  

About the Dataset 

The dataset contains 596 prompts. Each prompt is crafted to test 225 specific focus areas. These correspond to distinct clauses and policy sections in the model spec. Each focus area is a unique requirement that the models must fulfill.  

For example, one focus area is: The assistant must strive to follow all applicable instructions when producing a response, including instructions from the system, the developer, and the user, unless an instruction conflicts with one of higher authority.  

Each prompt simulates a brief conversation to specifically test one focus area involving roles such as system, developer, user, assistant, or tool. Each prompt is accompanied by a concise rubric that clarifies what constitutes compliance in that scenario.  

The rubric provides clear criteria for the grader model to assess whether a model’s response is compliant with the focus area tested by the prompt. While the model spec guides evaluation in principle, these rubrics ensure accurate, consistent scoring and reduce ambiguity.  

How We Built the Data Set 

Prompts and rubrics were written using models such as GPT-5. Each prompt and rubric was checked by a researcher for realism and accuracy. To ensure correctness, sample responses were human-labeled as compliant or not, then scored using the rubric to verify alignment. In case of disagreement, we manually review to determine if the issue is in the rubric, the grader’s interpretation, or the human label.  

How We Grade Model Adherence to the Model Spec 

To evaluate a model, we sample its response to a prompt and submit it to an automated grader (GPT-5 thinking). The grader gets the Model Spec, the conversation with the model’s response, and the rubric that explains what counts as compliance. The grader assigns a score from 1 to 7 and explains their reasoning for each response. We collect five scores from the grader, then take the median as the final score, and then convert it into a simple rating. Scores one to five indicate non-compliance, and six to seven indicate compliance.  

Early Results 

Newer models show higher Model Spec compliance. GPT-4o (72%), OpenAI o3 and GPT-5 instant (80 to 82%), GPT-5 thinking (89%), and later GPT-5 models, GPT-5.3 instant and GPT-5.4 thinking (84 to 87%).  

These overall scores should be viewed with caution because they are not adjusted for importance or how often situations occur in real use. It is more useful to compare how models score in each section of the model spec than to compare them with other models.  

In nearly all main sections, GPT-5 Thinking scored the highest. GPT-4o scored the lowest, with a difference of at least ten points. In the “Do the best work” section, the gap is almost 30 points. These improvements show the progress OpenAI has made. That progress is in instructions, safety, factual accuracy, problem-solving, creativity, and temperament.  

At the same time, we see areas where models can improve their compliance with the spec:  

  • Avoid overreaching and making decisions for the user.  
  • Present perspectives from any point on the opinion spectrum  
  • Avoid overstepping (e.g., doing more than the user asked for)  

What’s Next? 

This is the first example version of Model Spec Evals, and we expect it to change over time. Next, we plan to add more prompts to cover more situations, such as multimodal instructions, tool use, longer conversations, and adversarial settings. We will also keep the dataset up to date as the model spec changes.  

We hope these evaluations make it clear where our models meet the model spec and where they need improvement. We welcome feedback from developers, researchers, and the community. We look forward to working on this together. 

SourceIntroducing Model Spec Evals