AWS has introduced the Kiro framework, which lets AI agents run complex multi-step tasks for days using Long-R durable functions. These agents, now in preview, maintain context across sessions and can perform tasks such as code maintenance, bug triage, and automated testing without ongoing human input.  

Key Features of the Kiro Framework and Its Agents 

  • Kiro uses reliable Lambda functions to run longer workflows, avoiding common timeout problems in serverless computing.  
  • Kiro agents remember information across sessions and improve over time by learning from previous pull requests and user feedback.  
  • These agents autonomously tackle tasks for extended periods, following instructions, making plans, writing code, and running tests with minimal or no human input.  
  • Kiro uses a method called Spec Mode, which turns prompts into user stories, technical documents, and clear tasks for agents to follow.  
  • Kiro offers agent hooks for file-based triggers, MCP servers that provide specialized knowledge and tools to help agents follow coding standards, and tools that support agents in doing so.  
  • Deployment & access: the Kiro Framework preview is available to subscribers of Kiro Pro, Pro Plus, and Power Plans. Kiro is typically deployed as part of an AI-focused integrated deployment environment based on Code OSS, allowing eligible users to integrate and use the framework within their existing workflows. It is designed to reduce the need for constant supervision in AI-assisted development. This lets developers spend more time on important work while the agent manages complex asynchronous tasks.  

On Tuesday, Amazon Web Services introduced three new agents called Frontier agents. One of them is designed to learn your work preferences and then operate independently for several days.  

The Frontier agents each serve a unique function: one focuses on writing and maintaining code, another on reviewing security, and a third automates DevOps tasks to prevent issues when new code is deployed. Preview versions of all three are currently available.  

AWS claims its Kiro autonomous agent can operate independently for days, maintaining context and performing work without close supervision.  

Kiro is a coding agent built on AWS’s earlier AI tool of the same name, which launched in July. The first tool was meant for prototyping, but could also create code ready to go live. To maintain reliability, the AI adheres to the company’s coding standards through Specification-Driven Development.  

While Kiro writes code, people guide, confirm, or correct it, helping to make clear instructions. The Kiro autonomous agent learns by watching how the team uses tools and by reviewing existing code. After learning, AWS says it can work on its own.  

You simply assign a task from the backlog, and it independently figures out how to get that work done, AWS CEO Matt Garman promised during his keynote at AWS re:Invent on Tuesday.  

It actually learns how you like to work and continues to deepen its understanding of your code, your products, and the standards your team follows over time, he said.  

According to Amazon, Kiro maintains persistent context, so it does not lose track of tasks, enabling independent operation over hours or days with little human help.  

Garman gave an example of updating important code used in 15 different company programs. Instead of dealing with each update one by one, Kiro can fix all 15 with a single prompt.  

To further automate coding, AWS also created the security agent. The agent works independently to spot security issues as code is written. It tests them afterward and suggests fixes. The DevOps agent completes the set. It automatically tests new code for performance and checks if it works with other software, hardware, or cloud setups.  

Amazon is not the first to offer agents that can work for long periods. For example, last month, OpenAI said its GPT-5.1 Codex Max coding model is also built for long runs up to 24 hours.  

It is not certain that the main challenge in using these agents is the context window or their ability to run continuously. Large language models still struggle with accuracy and sometimes make mistakes, so coders often need to closely supervise them. Developers usually prefer to give short tasks and check the results quickly.  

However, for agents to truly work like co-workers, their context windows need to get larger. Amazon’s new technology is an important step toward that goal.

Source: Amazon previews 3 AI agents, including ‘Kiro’ that can code on its own for days 

The next wave of AI-powered robots, such as humanoids and self-driving vehicles, needs high-quality physics-based training data. If their datasets lack diversity and realism, these systems may not train well and could struggle with unexpected situations. Gathering large real-world datasets is costly, time-consuming, and often constrained by pragmatic constraints.  

NVIDIA Cosmos addresses this problem by accelerating the development of world-class models (WCMs). Cosmos WFM enables faster synthetic data generation and provides a foundation for training specialized physical AI models. In this post, we’ll look at the newest Cosmos WFM’s, their main features to advance physical AI, and how you can use them.  

Cosmos World Foundation Model Updates 

NVIDIA Cosmos world-based models are improving rapidly, making it easier for users to access high-quality synthetic data and accelerated physical AI development. After just one year, recent updates ensure users benefit from faster, more flexible, and realistic data generation processes.  

  • Cosmos Transfer 2.5: Delivers Faster, More Scalable Data Augmentation. The process of creating varied data by altering existing data from simulations and 3D spatial inputs provides greater variety within environments, lighting, and scene setups.  
  • Cosmos predict 2.5: improves generation of rare scenarios for sequences up to 30 seconds, attaining up to 10 times higher accuracy when post-trained on custom or sector-specific data. It also supports multi-view outputs, custom camera setups, and various policy outputs, such as action and simulation.  
  • Cosmos Reason 2: offers advanced physical AI reasoning with better spatio-temporal understanding (the ability to interpret spatial and temporal relationships) and more precise timestamps. It adds: Object Detection, 2D and 3D point localization (Finding locations in flat and 3D spaces), Bounding box coordinates (Boxes that identify the positions of objects), Reasoning explanations, and labels. It now supports Long Context Improved inputs up to 256,000 tokens (a token is a unit of text, like a word or character).  

Cosmos Transfer Creates Photorealistic Videos That Adhere To Real-World Physics 

Cosmos Transfer creates detailed word sense from structural inputs, ensuring accurate spatial alignment and composition.  

Cosmos Transfer uses the controlnet architecture to retain pre-trained knowledge, resulting in structured, consistent outputs. It uses spatial-temporal control maps to match artificial and real-world scenes, giving detailed control over:  

  • scene layout  
  • object placement and movement  
  • eye points  
  • lidar scans  
  • trajectories  
  • HD maps  
  • 3D bounding boxes  

Ground Truth Annotations: High Fidelity References for Exact Alignment  

Output: photorealistic video sequences with controlled layout, object placement, and motion.  

Key Capabilities 

  • Generate scalable, photorealistic, synthetic data that aligns with real-world physics, allowing users to train more reliable AI and robotics models.s.  
  • Control object interactions and scene composition with structured multi-modal input, giving users precise customization and more relevant training data for their specific use cases.s.  

Using Cosmos Transfer for Controllable Synthetic Data 

With Generative AI APIs and SDKs, NVIDIA Omniverse enables users to create accurate 3D simulations for real-world training and testing. These experiments provide ground-truth video inputs for Cosmos Transfer, improving photorealism and diversifying datasets to fit user-specific conditions, ensuring your AI agents are better prepared for real-world deployment.  

This process speeds up the generation of high-quality data, enabling users’ AI agents to learn more efficiently from simulation to real-world applications, reducing development cycles and boosting performance in practical tasks.  

As a result, Cosmos Transfer helps users train robots and AI for diverse environments and conditions by adding realistic lighting and textures. This improves model robustness and makes it easier for users to transition from simulation to real-world use, especially for robotics platforms like GR00T-N1.1.  

Cosmos Predict for Generating Future World States 

Cosmos Predict WFM enables users to generate predictive video sequences for future scenarios using varied inputs such as text, video, and image sequences. Its smooth, accurate video generation helps users test and refine how AI systems might respond in real-world situations.  

The following key capabilities were developed in our Cosmos Credit functions. It creates realistic video scenes directly from text prompts.  

  • Predicts subsequent events in a video by generating missing frames or continuing motion  
  • Generates multiple frames (intermediate images) between a starting and ending image to create a smooth, complete video sequence.  

Cosmos Predict WFM is a solid starting point for training world models, AI systems that simulate environments used in robotics and self-driving vehicles. After initial training, you can teach these models to generate actions rather than videos for policy modelling and AI decision-making, or adapt them for visual language tasks to build custom AI perception models (systems that understand visual information).  

Cosmos Reason: Designed to Perceive Reason and Respond Intelligently 

Cosmos Reason is a flexible AI model designed to understand motion, how objects interact, and relationships over time and space. It uses chain-of-thought reasoning to examine visual input, predict outcomes from prompts, and choose the best actions. Unlike text-only models, it bases its reasoning on actual physics and provides clear natural-language context for its answers.  

Video: observations along with a text question for instruction (prompt).  

Output: a text response created using long-term chain of thought reasoning (step-by-step analysis over time).  

  • Understands how objects move, interact, and change  
  • Predicts and selects optimal next actions based on observations.  
  • Continuously refines its decision-making ability over time.  
  • It is designed for further training to help build perception AI and embodied AI models.  

Let’s Get Started 

Explore our Cosmos Cookbook for user-focused step-by-step guidance, technical tips, and examples that help you streamline and accelerate your Cosmos WFM projects.s.  

Access open Cosmos models and datasets on Hugging Face and GitHub to quickly enhance your projects or evaluate models, making experimentation and implementation faster and easier for users.  

Join our Cosmos Discord community now—connect with peers, get real-time support, and share unique experiences. Become part of our vibrant network today!  

Be inspired: Watch the GTC Keynote from NVIDIA founder and CEO Jensen Huang. Then explore Cosmos sessions and kick-start your own breakthrough projects with insights at https://www.nvidia.com/gtc/sessions/physical-AI-days/. Start your journey today! 

Source: Scale Synthetic Data and Physical AI Reasoning with NVIDIA Cosmos World Foundation Models 

Today’s semiconductor market is highly competitive. Mobile application processors (APs) must continue to improve performance, even as devices get slimmer and more powerful. As on-device AI becomes more common, power use in smaller spaces increases, leading to higher power density and more heat. People still want longer battery life and thinner, lighter phones. Because of this, mobile APs need more than small performance upgrades; they need new designs that use space efficiently.  

To address these issues, mobile AP packaging is evolving. It now does more than just protect the chip; it also manages heat and maximizes space. By improving package architecture and thermal dissipation, packaging maintains performance and reliability. It enables slimmer designs and larger batteries, making packaging increasingly critical for mobile APs.  

When APs use more power to boost performance, temperature rises. Cooling them requires reducing power, which stops the chip from reaching its full potential. Now, controlling thermal resistance is key to a stable mobile AP design. Traditional solutions use heat-conductive materials or a thicker silicon die. As devices shrink, these methods alone are not enough to solve the heat problem.  

The Shortcomings of Conventional PoP Designs 

For high-end APs and SoCs, Package-on-Package (PoP) design is common to improve performance. In this design, DRAM sits on top of the AP chip. As mobile devices get thinner, so does the package, including the AP die. This reduces the path for heat to escape. Heat builds up faster, and the chip reaches its thermal limit sooner, which limits sustained peak performance.  

When the AP is in operation, the heat generated by the silicon die inside the AP package must be quickly dissipated to reduce chip temperature. Lower thermal resistance improves heat dissipation efficiency, helping preserve stable performance even under high watt loads. To achieve this, device makers use heat-dissipation components, such as heat spreaders and vapor chambers, to transfer the heat generated by the AP to external cooling structures. However, in conventional POP structures, the DRAM package is positioned above the AP chip, limiting direct heat transfer between the AP and the heat dissipation components. This structural characteristic reduces heat transfer efficiency and acts as a basic limitation on performance improvements at both the package and system levels.  

Samsung’s HPB addresses the need for steady performance improvements among mobile APs. 

Samsung has improved thermal management by placing the heat path block (HPB) on top of the AP chip. This is the first time HPB has been used with Fan-Out Wafer-Level Packaging in the industry. It reduces thermal resistance within the package and maintains stable performance under heavy use.  

This method creates a new package that moves heat from the AP die to the phone’s cooling parts more effectively. The DRAM package is now away from the main heat source, unlike in previous PoP designs. The HBM is placed directly above the heat source, helping heat escape more quickly and efficiently.  

The HPB’s Core Benefits 

Samsung created the HPB to efficiently remove heat from the AP die, maintain stable performance, and retain structural strength. They also introduced a new thermal interface material (TIM) with high thermal conductivity and strong bonding. This combination improves both heat dissipation and package reliability.  

In a POP (package-on-package) structure, heat from the bottom AP (application processor) die must be transferred upward through the intermediate DRAM (dynamic random access memory) package. The heat transfer path passes sequentially through:  

  • the DRAM packages  
  • the bottom solder balls  
  • the substrate  
  • the DRAM die  
  • the EMC (epoxy molding compound)  

Along this path, the solder balls, which have relatively high thermal conductivity, are distributed only in a limited region. The substrate dielectric layer, D-AF, used for die-stacking and EMC is composed entirely of low-thermal-conductivity materials. DRAM packages are inherently inefficient, limiting effective heat transfer to mobile device cooling components such as vapor chambers.  

Unlike these materials, the HPB used in the Exynos 2600 is made of copper, which has a thermal conductivity of about 400 W/m·K. This is 500 to 1000 times better at transferring heat than the polymer materials used in substrates, DAF, or EMC. As a result, heat from the AP die quickly dissipates from the package, keeping temperatures lower at the source and helping maintain strong performance. This performance improvement is shaped not only by the materials but also by the evolution of the package’s design and development.  

From Challenges to Breakthroughs: The Development Journey Shaping the Future of Samsung’s Mobile Packaging 

To improve heat transfer from the AP die to the HPB, the new package design cut the DRAM package size by about half. It also adjusted both the overall package height and AP package thickness. These changes are intended to improve the thermal path without significantly increasing the package size.  

Because of the asymmetric placement of the DRAM, Samsung reconfigured the AP DRAM interface. The overall chip and package architectures were also redesigned for performance and reliability. They pre-validated HPB thermal reduction through multi-perspective simulations. Target performance was achieved and sustained using root cause analysis and progressive optimization across materials, processes, and product stages. This progress was supported by close teamwork among departments.  

Mobile processors must perform better while staying small. As a result, designing packages to manage heat will become even more important for stable AP performance. The HPB-based package shows how changing the heat transfer path can solve these problems.  

By developing the HPB, Samsung Electronics has gained important technical know-how, testing methods, and a solid approach to team collaboration. With this foundation, the company plans to keep improving AP Packaging to deliver better performance, thermal stability, and spatial optimization in future mobile devices.

Source: Introducing a New Package Architecture for Improved Thermal Efficiency in Mobile Application Processors 

The Buzz 

  • Microsoft has introduced Azure Local Disconnected Operations, Microsoft 365 Local, and Foundry Local—allowing large AI models and productivity tools to run fully offline, as detailed on Microsoft’s official blog.  
  • Organizations can now run multi-modal AI models on NVIDIA hardware within their own secure environments, ensuring full compliance and security without any cloud connection.  
  • The Microsoft 365 productivity suite, including Exchange, SharePoint, and Skype for Business, can now run completely offline through at least 2035.  
  • Defense, government, and other regulated sectors that were previously restricted by compliance rules can now access enterprise AI infrastructure, improving accessibility and enabling innovation.  

Microsoft’s three new sovereign cloud updates let enterprises run large AI models, productivity software, and cloud infrastructure fully offline. This enables high-tech privacy and total data control for regulated sectors. Azure Local, Microsoft 365 Local, and Foundry Local with NVIDIA GPU support empower organizations to securely deploy advanced AI on-premises.  

Microsoft is transforming secure Enterprise AI by allowing large language models and productivity suites to run offline, ensuring strong data privacy and operational control without a cloud connection.  

Azure Local Disconnected Operations are now available, allowing organizations to set up critical infrastructure using Azure’s management tools without any external connections. All management, policy enforcement, and workload execution occur within the customer’s environment. This is a significant change for defense contractors handling classified work or for financial institutions operating in jurisdictions with strict data residency laws.  

The availability of Azure Local Disconnected Operations represents a breakthrough for organizations that need control over their data without sacrificing the power of the Microsoft Cloud. Gerard Hoffman, CEO of Proximus Luxembourg, told Microsoft in a statement, “For Luxembourg, where digitalized sovereignty is not simply a principle but a key necessity, this model offers the strength, autonomy, and trust our market expects.”  

Productivity is just as important as infrastructure. Microsoft 365 Local Disconnected now provides Exchange, SharePoint, and Skype for Business servers fully within customers’ own environment, with promised support through at least 2035. Teams can collaborate, share files, and communicate without any data leaving their network. Customers have full control over access, compliance, and data protection.  

The biggest news is that Foundry Local can now run large-scale AI models on-site. Foundry Local is a set of AI tools that operate entirely within an organization’s own network. Microsoft is adding NVIDIA GPUs so organizations can run computationally intensive AI tasks without connecting to external networks. This update means highly secure organizations can have local, advanced AI capabilities while ensuring strict privacy and compliance.  

The technical architecture is straightforward: In connected mode, a central management component, the control plane, runs in a Microsoft Cloud region and sends configuration and monitoring commands to local, customer-owned servers. In disconnected mode, the control plane itself runs as a virtual machine on the customer’s infrastructure, directly managing Foundry Local, Microsoft 365 Local, and Azure Local without any data or communication reaching Microsoft’s external clouds. The user experience for configuring, monitoring, and updating these services stays the same, whether systems are online, offline, or fully air-gapped.  

Azure Local and Microsoft 365 Local are now globally available in disconnected mode, with Foundry Local’s large AI models offered to qualified compliance-driven customers.  

Digital Sovereignty Roles are patterning worldwide. Microsoft is designed to meet real customer needs, be independent of external connections, and ensure guaranteed continuity.  

Foundry Local is built to handle large models and GPU needs. Microsoft provides support for setup, updates, and work while customers retain full control over their data and hardware.  

The competitive landscape is shifting. While Amazon Web Services offers Hot Posts and Google Cloud provides distributed cloud, neither lets customers run large AI models in completely disconnected, secure environments as Microsoft does. Microsoft leverages its experience with on-premise products like Exchange and SharePoint to give organizations an offline operations and scalable AI advantage.  

In industries such as defense, intelligence, health care, and critical infrastructure, and in certain jurisdictions, this opens up AI capabilities that were previously off-limits. A defense contractor can now run the same multimodal models used in commercial settings, just air-gapped inside a classified facility. A European bank can deploy large language models for internal tools without data crossing borders or touching external networks.  

This setup also addresses operational complexity. Organizations no longer need separate management systems, different governance rules, or split architectures for online and offline workloads, and can achieve consistent management. Whether systems are online, sometimes connected, or always offline, simplifying operations and enhancing efficiency  

Douglas Phillips, President and CTO of Microsoft Specialized Clouds, leads the engineering effort behind these capabilities. His team is responsible for bringing Azure, Microsoft’s adaptive cloud portfolio, and the Microsoft 365 Collaboration Suite to customers with sovereignty, security, edge, and compliance requirements that standard cloud offerings can’t address.  

These changes affect more than just Microsoft customers. This level of secure AI sets a new standard for what businesses can expect from cloud providers. It shows that advanced AI can be used without giving up data control, regulatory compliance, or operational independence.  

Microsoft’s sovereign cloud expansion fundamentally changes what’s possible for enterprises operating under strict compliance regimes. By enabling large AI model deployment, full productivity sockets, and cloud infrastructure to run completely disconnected from external networks, the company is opening up AI capabilities to sectors that were previously locked out by regulatory limitations. The question now isn’t whether sovereign AI is technically feasible. Microsoft just proved it is. The question is how quickly competitors respond and how fast regulated industries adapt these capabilities to close the AI gap within their commercial counterparts.

Source: Microsoft Sovereign Cloud Goes Fully Offline With AI Support 

Key Details 

  • Boston Dynamics launches immediate production of the humanoid Atlas robot.  
  • In 2026, the rover will be deployed at Hyundai and Google DeepMind, with additional customers anticipated the following year.  
  • Atlas will be trained with new AI-based models to handle many industrial tasks, starting with the automotive industry.  

Boston Dynamics, a leader in mobile robotics, introduced the product version of its new Atlas robot at the Consumer Electronics Show in Las Vegas. The fully electric humanoid was shown during Hyundai’s CES Media Day, which also included a live demo. It is the latest Atlas prototype and a lively dance performance by the well-known Spot Robots.  

Production of the new Atlas robots will begin immediately at the company’s Boston headquarters. All units for 2026 are already spoken for, with fleets set to ship to Hyundai’s Robotics Metaplant Applications Center (RMAC) and Google DeepMind soon. More customers will be added in early 2027.  

For more than 30 years, Boston Dynamics has been building some of the world’s most advanced robots, said Robert Playter, the company’s CEO. This is the best tool we have ever built. Atlas is going to change the way the industry works and make its mark. It will be the initial step toward a long-term goal we have dreamed about since we were children. Useful robots can walk into our homes and help make our lives safer, more productive, and more fulfilling.  

Atlas is an enterprise-grade humanoid robot capable of handling many tasks. From moving materials to filling orders, it learns from new tasks quickly, adapts to evolving environments, lifts heavy loads, and works independently with little supervision. It keeps working at a steady, reliable pace and does not need to stop when its battery runs low. Instead, it will find a charging station, swap its own batteries, and return to work.  

The robot connects easily to manufacturing systems such as MES (Manufacturing Execution System) and WMS (Warehouse Management System), as well as other industrial software, via Boston Dynamics’ Orbit software. Once one Atlas robot learns a new task, that skill can be shared instantly with the whole fleet.  

Atlas operates autonomously via remote or with a tablet. It has 56 degrees of freedom. A 2.3-meter reach lifts up to 50 kg, is water-resistant, and works from -22°C to 40°C.   

Safety features include human detection and fenceless guarding (protecting people without physical barriers), with workflow integration via barcode or RFID (radio-frequency identification).  

“Our new Atlas is the most production-friendly robot we’ve got,” said Zach Jackowski, GM of Atlas at Boston Dynamics. This generation of Atlas uses fewer unique parts, and every component is made to fit with automotive supply chains. With support from Hyundai Motor Group, we will reach the highest reliability and economies of scale in the industry.  

Along with launching Atlas at CES, Boston Dynamics announced a new partnership with Google DeepMind. They plan to use Google DeepMind’s advanced base models to improve Atlas’s cognitive abilities. The company also shared that Hyundai Mobis will supply Atlas actuators. Both organizations will work together to build an efficient supply chain and speed up actuator development and production.  

Hyundai Motor Group holds a majority stake in Boston Dynamics. The company is preparing to deploy tens of thousands of Boston Dynamics robots in its manufacturing facilities. Hyundai also announced a $26 billion investment in its U.S. operations. This includes plans for a new robotics facility with an annual capacity of 30,000 robots.  

To learn more, visit www.bostondynamics.com.  

About Boston Dynamics. 

Boston Dynamics leads the world in developing and deploying highly mobile robots that handle tough industrial and safety challenges. Our robots have advanced mobility, dexterity, and intelligence, enabling automation in hard-to-reach or unsafe environments such as factories, power plants, construction sites, warehouses, and distribution centers.  
 
Our portfolio includes three robots:  

  • Spot, a four-legged robot for industrial inspections and public safety  
  • Stretch, a robot that moves boxes for logistics and retail  
  • Atlas, our electric humanoid platform, is now in development.

Source: Boston Dynamics Unveils New Atlas Robot to Revolutionize Industry

Apple has introduced MacBook Neo, a new laptop designed to bring the Mac experience to more people at an affordable price. MacBook Neo features a durable aluminum body and comes in four colors: Blush, Indigo, Silver, and Citrus. Its 13-inch Liquid Retina display offers sharp images and supports 1 billion colors. Powered by the A18 Pro chip, MacBook Neo manages everyday tasks such as web browsing, streaming, photo editing, creative projects, and AI features with ease. It is up to 50% faster for daily tasks and up to 3 times faster for on-device AI tasks compared to the best-selling PC with the latest Intel Core Ultra 5, with up to 16 hours of battery life. Users can work or play all day on a single charge. The 1080p FaceTime HD camera and dual microphones help users look and sound their best, while side-firing speakers with Spatial Audio provide clear, immersive sound. The Magic Keyboard and large multi-touch trackpad make typing and navigation comfortable and precise. MacBook Neo runs macOS Tahoe and includes built-in apps such as Messages, Pages, Calendar, and Safari. Smooth integration with iPhone, Apple intelligence, and support for third-party apps. Starting at $599 or $499 for education, MacBook Neo is Apple’s most affordable laptop yet. Pre-orders begin today, and it will be available starting Wednesday, March 11.  

“We are incredibly excited to introduce MacBook Neo, which delivers the magic of the Mac at a breakthrough price,” said John Ternus, Apple’s Senior Vice President of Hardware Engineering. “Built from the ground up to be more affordable for even more people, MacBook Neo is a laptop only Apple could create. It features a durable aluminum design in four beautiful colors, a brilliant Liquid Retina display, Apple Silicon-powered performance, all-day battery life, a high-quality camera and audio, and the intuitive power features of macOS. There is simply no other laptop like it.”  

Beautiful And Durable Aluminum Design 

The MacBook Neo features a carefully crafted aluminum design for durability. Its soft, rounded corners give it an elegant look and a comfortable feel. Weighing only 2.7 lb, it is easy to carry in a backpack or bag. MacBook Neo adds personality and style to daily use. It comes with four colors:  

  • Blush  
  • Indigo  
  • Silver  
  • Citrus  

These colors also appear on the Magic Keyboard in lighter shades and in new wallpapers, creating a unified and colorful look.  

Stunning 13-inch Liquid Retina Display 

The 13-inch Liquid Retina Display provides a sharp 2400 x 1600 resolution, 500 nits of brightness, and support for one billion colors, surpassing the brightness and sharpness of most PC laptops in this price range. The anti-reflective coating helps maintain clarity and comfort in various lighting conditions, whether you are watching movies, editing photos, or in a video call.  

Apple Silicon Powered Performance 

The MacBook Neo runs on the A18 Pro Chip, enabling everyday tasks like browsing, writing, streaming, and photo editing to run fast and smoothly. You can easily switch between apps such as Messages, WhatsApp, Canva, Excel, and Safari on the top-selling PCs with the latest Intel Core Ultra 5, compared to the top-selling PCs with the latest Intel Core Ultra 5. MacBook Neo is up to 50% faster for daily use and for more demanding tasks. It’s up to three times faster for on-device AI and twice as fast for photo editing. The five-core GPU delivers great graphics. For games and creative projects, the 16-core neural engine powers Apple Intelligence features and AI tasks like summarizing notes or cleaning up photos, while keeping your data safe. Plus, MacBook Neo is fanless, so it stays completely silent.  

All-day Battery Life 

Thanks to Apple Silicon, the MacBook Pro delivers up to 16 hours of battery life on a single charge. This reliability makes it well-suited for work or play, whether in class, at a coffee shop, or on the move.  

Magic Keyboard and New Multi-touch Trackpad 

The MacBook Neo comes with Apple’s Magic Keyboard for comfortable, precise typing. The large, multi-touch trackpad lets you click, scroll, swipe, and pinch anywhere on its surface. If you choose the model with Touch ID, you can log in quickly and securely and easily approve purchases with Apple Pay.  

1080P Camera, Dual Speakers, and Mics 

The MacBook Neo’s 1080p FaceTime HD camera uses advanced image processing for sharp, vibrant video calls. Dual microphones block background noise for a clear voice during meetings. Side-firing speakers with spatial audio and Dolby Atmos deliver immersive sound whether watching movies, listening to music, or working in GarageBand.  

Essential Connectivity 

MacBook Neo features two USB-C ports for connecting accessories or an external display. Both ports can be used for charging. MacBook Neo also includes a headphone jack for audio. Wi-Fi 6E delivers fast wireless connectivity, and Bluetooth 6 enables and ensures reliable connectivity for peripherals and accessories.  

Powerful Productivity with macOS 

MacOS is Apple’s easy-to-use and powerful operating system for Mac. With built-in apps like Safari, Photos, Messages, and FaceTime, you can get started right away. Apple Intelligence features, such as writing tools and live translation, are built into macOS to make everyday tasks smarter and easier. You also get strong privacy and security, including top-level encryption, virus protection, and free automatic security updates.  

Flawless integration with iPhone 

If you use an iPhone, you can take advantage of continuity features in MacOS to make switching between your iPhone and MacBook Neo easy.  

  • Handoff lets you start a task on your MacBook Neo and finish it on your phone.  
  • Universal clipboard lets you copy and paste between devices.  
  • With iPhone mirroring, you can see and use your iPhone right on your MacBook Neo.  
  • If you are new to a Mac, you can use your iPhone to quickly and securely transfer your settings, files, photos, passwords, and more.  

Built With The Environment In Mind 

MacBook Neo is Apple’s lowest-carbon MacBook yet, helping the company move closer to its goal of being carbon-neutral by 2030. It uses 60% recycled materials, the highest of any Apple product. This includes 90% recycled aluminum and a battery made with 100% recycled cobalt. The enclosure is made with a process that uses half as much aluminum as standard methods. MacBook Neo is built using 45% renewable electricity, such as wind and solar, throughout the supply chain. It also meets Apple’s strict standards for energy efficiency and safe materials. The paper packaging is made entirely from fiber and is easy to recycle.

Source: Say hello to MacBook Neo 

Xcode 26 introduces powerful local AI capabilities by leveraging base models on the Neural Engine for secure on-device AI processing. This reduces reliance on cloud services. Developers can now benefit from inline code suggestions, automated test and documentation generation, and integration with third-party models.  

Below are the main AI features and improvements introduced in Xcode 26, setting the stage for enhanced development workflows. 

  • On-device AI power: base models, which are foundational artificial intelligence algorithms, now run locally, enabling fast, secure processing on Apple Silicon chips.  
  • Intelligent Coding Tools: Xcode 26 offers in-line code generation and debugging tools that automatically generate and test code as you work, improving developer efficiency.  
  • Model flexibility: developers can use local models (AI systems processed on their computer) or connect to third-party providers such as ChatGPT and Claude, which are external AI services, directly within the editor.  
  • Model training: fine-tune on-device models with local data—meaning training the AI using information on your device—to enable apps to learn specialized tasks and improve intelligence.  
  • Performance optimization: algorithms such as Lexicographical_compare, a tool for sorting data by character order, now execute faster, and vector computation (calculations on lists of numbers) has improved.  
  • Enhanced tools: this update brings improved localization catalogs and new resources for developing AI models.  

With these updates, you can build apps that are faster, smarter, and more private, unlocking the full potential of the Apple ecosystem.  

Xcode 26 comes with Swift 6.2 and SDKs for:  

  • iOS 26  
  • iPadOS 26  
  • tvOS 26  
  • WatchOS 26  
  • MacOS Tahoe 26  
  • VisonOS 26  

You can debug on devices running:  

  • iOS 15 or later  
  • tvOS 15 or later  
  • watchOS 8 or later  
  • visionOS  

To use Xcode 26, your Mac needs to run macOS Sequoia 15.6 or newer.  

Enhance your workflow with a Coding Intelligence tool to write code, create tests and documentation, fix errors, refactor, and navigate projects. Xcode now supports ChatGPT, Claude, and API keys for providers using the Chat Completions API or a local model on Apple Silicon Macs.  

  • You can use natural language instructions to work with code in the coding assistant. The assistant gathers information relevant to your current code, remembers past conversations, and lets you attach files for more context.  
  • Coding Tools deliver actions to generate documentation, explain code, create previews and playgrounds, and edit inline.  
  • Predictive code completion, a feature that suggests how to finish writing code based on context, is faster and uses more code contexts, all locally on your Mac.  

Also in Xcode 26: 

  • The ‘#’ playground macro is a command that lets you debug and experiment with code in real time in the preview panel, which visually displays code output as you write.  
  • One Composer makes it easier to create icons from one design. You can adjust depth, add dynamic lighting, and customize icons for default dark and mono modes.  
  • Tabs have been redesigned to make navigation easier. You can now use tab navigation and pin files to keep them in view.  
  • Compilation caching stores data from previous builds, so build times are faster, especially when switching between code branches (different versions of your project) or performing clean builds, which means compiling everything from scratch.  
  • New Instruments helps you analyze your app’s:  
  • Performance  
  • Processor  
  • Trace records capturing every function call made by the app.  
  • Swift UI Profiles help you monitor Swift UI Views. Power Profiler measures how your app uses Battery and creates Heat. CPU Counters help you find and fix slow parts of your code.  
  • Swift Concurrency Debugging now monitors execution across asynchronous (async) functions, which are tasks that run at the same time, and threads (sequences of tasks handled separately by the processor). It allows clear types of concurrency ways in which multiple tasks operate at once and helps you see the properties and relationships for each task in your code.  
  • String catalogs help organize and manage localization, which is the translation of your app into different languages. They use type-safe Swift symbols special labels that prevent errors so you can reference strings directly in code, support auto-complete for string lookup, and give AI-generated comments using on-device processing.  
  • Voice control now lets you dictate Swift code using syntax-aware recognition, which is an input system that understands the structure of the language and automatically formats your code as you speak.  

General. 

New features. 

  • Hang and Launch Diagnostics now include trending insights. These highlight issues that have become more common across the last four adversions and provide further context on their impact. Look for the flame icon in the source list to spot this data. See when an issue started. Emphasize Performance Fixes in New App Versions (135376723). There is a new setting for how function names appear in the C++ frames plugin.cplusplus.display.function-name-format. By default, this displays the entire function name but can be customized to drop various parts of a function signature (e.g., return type, scope qualifiers, etc.). see FUNCTION FUNCTION-NAME-FORMATS-FOR-NO-DETAILS)  
  • LDB now marks the version base name by default when showing C++ frames—a backtrace is a report showing the call sequence of functions leading to a certain point in code.  

Xcode 26 

Turn your ideas into reality using Generative Intelligence powered by your preferred large language model. The coding assistant lets you interact with your code using natural language. With coding tools, you can quickly write documentation, fix issues, and make changes directly in your code. Use the playground macro to preview your known UI code. The redesigned tab experience makes it easier to move through your files. Plus, improved eye localization catalogs help you reach more users worldwide. Building on these powerful tools, Xcode introduces additional ways to optimize your workflow and app performance.  

Instruments. 

Optimize your app for Apple Silicon using two new hardware-assisted tools, Instruments:  

  • Processor Trace  
  • and CPU Counter  

Use the new SwiftUI instrument to observe how changes in your app’s data affect SwiftUI. View updates to these performance insights. Complement the enhanced automation capabilities found in XCUI automation tests.  

XCUI Automation Tests 

Now you can record, run, and manage XCUI automation tests directly in Xcode. Test plan configurations let you replay your XC test UI tests across many locales, device types, and system conditions. Review your results in the Xcode test report and download screenshots and videos from test runs as you refine your testing and development. Xcode’s new design resources simplify design asset management, starting with Icon Composer.  

Icon Composer 

Icon Composer helps you create layered icons using Liquid Glass from a single design for iPhone, iPad, Mac, and Apple Watch. The new multi-layer icon format lets you adjust Liquid Glass properties, preview dynamic lighting effects, and add annotations for different appearance modes. Icon Composer works smoothly with X-Core and lets you export a flattened icon for marketing or communication.

Source: Xcode 26 Release Notes 

Agentic AI systems must use deep models capable of independently solving complex technical problems.  

Multi-agent systems can produce up to 15 times more tokens than standard chats since they keep sending history, tool outputs, and reasoning steps at each step during long tasks. This context explosion can lead to the world rift, where agents slowly lose track of the main objective due to the need to use large reasoning models for every subtask, known as the thinking tax. Also, these applications are too costly and slow in real-world use.  

Today, we are announcing Nemotron 3 Super to solve these problems. The new Super model has 120 billion total parameters, with 12 billion active at a time. It is designed for maximum effectiveness and precision in complex multi-agent tasks such as software development and security triage. This release follows our introduction of Nemotron 3 Nano in December.  

Number 23 Super solves the thinking tax problem with its hybrid mixture-of-experts (MoE) design. It offers more than five times the throughput of the previous Nemotron Super. The model also handles context explosion with a built-in 1-million-token context window, providing agents with long-term memory for accurate reasoning. It is fully open, with open weights, datasets, and recipes, so developers can easily customize, optimize, and deploy it on their own systems.  

WhatsApp’s Nemotron 3 Super Apart 

Nemotron 3 Super introduces design features to reduce trade-offs between effectiveness and correctness in large reasoning models.  

  • Latent MOE (a type of mixture-of-experts architecture that compresses hidden data representations) uses token compression to activate more experts (specialized submodels) per inference at the same computational cost.  
  • Multi-token prediction accelerates long-sequence generation by enabling single-step future-token prediction, thereby enabling speculative decoding.  
  • A Hybrid Mamba Transformer backbone means the model has two main types of layers. Mamba layers process long sequences efficiently, while Transformer layers are specialized for exact reasoning. This combination increases the model’s speed and makes it four times more memory- and compute-efficient.  
  • Native NVFP4, used for pre-training, is a special low-memory format built for NVIDIA Blackwell chips. It reduces memory usage and speeds inference (model output) by 4x on NVIDIA B200 compared to FP8 on NVIDIA H100, while maintaining high accuracy.  
  • After initial training, the model uses reinforcement learning (AI learns by trial and error) across 21 environments, running on NVIDIA Nemotron Gem and Nemotron RL and accumulating 1.2 million simulated experiences (environment rollouts).  

These advantages combine to create a model well-suited for long-running autonomous agents on Pinchbench, an innovative benchmark for evaluating how well LLMs perform as the brain of an open-cloud agent. Nemotron 3 Super scores 85.6% across the full test suite, making it the best open model in its class.  

Delving Deeply Into The Architecture 

Hybrid Mamba Transformer MOE Backbone 

Super uses the same blended approach as nano, but it operates on a much larger scale to understand how its architecture supports this. Consider the way its backbone combines three types of layers.  

Mamba-2 layers handle most sequence processing. These State-Space Models operate in linear time with respect to sequence length, enabling the practical use of a 1M Token Context Window. Mamba layers efficiently manage memory when processing large code lengths, extended chat histories, or many documents.  

Transformer attention layers are inserted at key points because SSMs may struggle to locate specific facts in a known context. These attention layers preserve the ability to retrieve targeted information within large inputs.  

MOE layers increase the number of effective parameters without needing heavy computation. Only some experts are used for each token, which keeps latency low and throughput high. This is important when many agents run concurrently in a shared system.  

Latent MOE 

In a typical MOE setup, tokens are routed from the full hidden dimension (all the internal data space) to the different experts (specialized sub-networks). This process can slow down computation, increase costs, and limit the number of experts the model can use.  

Supersuper uses latent MOE (working in a reduced dimension before routing). Here, token embeddings (numerical summaries of tokens) are compressed into a simpler, smaller space. The experts perform their tasks in this compressed space, and then their results are expanded to match the full model size. This approach has practical effects:  

More experts can be used at the same cost per computer (i.e., the same amount of computer resources). Compression enables the model designer to support four times as many experts without requiring more computation.  

Finer-grained specialization, when more experts are available, the model can afford highly specialized routing, for example, activating distinct experts for Python syntax versus SQL logic only when strictly necessary. This granularity is especially valuable in agentic settings where a single conversation may span tool calls, code generation, data analysis, and dialogic reasoning within a few terms.  

Multi-Token Prediction (MTP) 

Standard language models are trained to predict one token at a time. A fundamentally myopic objective, Super is trained with MTP, where specialized prediction heads simultaneously forecast seven future tokens at each position.  

This has two concrete benefits:  

This leads to better reasoning during training. Predicting multiple future tokens helps the model learn longer patterns and logical connections, rather than just guessing the next word. The model is trained to predict entire sequences, which improves performance on tasks that require step-by-step logic.  

MTP also speeds up inference by predicting multiple future tokens in a single pass. This reduces the time required to generate long outputs. The MTP heads make draft predictions that can be checked in parallel, enabling up to 3x faster generation for tasks such as code and tool calls without requiring a separate draft. Both benefits come from a single design. Shared weights across all MTP heads limit the number of extra parameters and stabilize training by aligning heads, keeping draft predictions more consistent across longer sequences.  

Native NVFP4 Pre-Training 

Most quantized models are first trained in full precision and then compressed, which usually results in some loss of accuracy. Super does things differently. Most floating-point operations during pre-training use NVFP4 and NVIDIA’s 4-bit floating-point format. This format, optimized for Blackwell, greatly reduces memory usage and speeds up inference compared to FP8 while maintaining high accuracy.  

By training natively in reduced precision (only using 4-bit math from the beginning), the model learns to be accurate even with small numbers, meaning it stays stable and effective while using much less memory.  

Training Super: A Three-Stage Process  

  1. Supervised fine-tuning is a phase in which the model is further trained on specific examples relevant to the tasks it will face, so it learns to act appropriately for those tasks.  
  1. Reinforcement learning is used to further improve model behavior by letting it learn from actual outcomes in test scenarios.  

Pre-training: Super is pre-trained on 25 trillion tokens using the NVFP4 format, learning to be accurate with four-bit math throughout pre-training. The data includes 10 trillion unique curated tokens focused on reasoning and coding.  

Supervised Fine-Tuning: Before Reinforcement Training, Super is being fine-tuned with about 7 million supervised samples. These come from a larger set of 40 million samples that include reasoning, following instructions, coding safety, and multi-step agent tasks. This stage lays the foundation for the behavior RL will employ. The model teams learn to give correct responses across different tasks, so RL starts from a stable base rather than a raw, pre-trained model.  

Multi-environment reinforcement learning makes the model more agent-like by training it across many environments (such as Nemotron, GitHub, NMEDIAS, and the Open RL Library). These scenarios test tasks such as tool use, coding, and complex planning, creating the main dataset for reinforcement training.   

This type of reinforcement learning step helps the model work reliably in multi-step workflows, reduces reasoning errors, and manages the organized tasks often found in agent pipelines.  

The Super + Nano Deployment Pattern 

Nemotron 3 Nano works well for processing specific targeted steps in an agent-like workflow; however, as multi-agent applications become more complex and involve several steps, a more powerful model is needed for better planning and reasoning. For example, imagine an agent that needs to choose among tools to create a presentation with 10 high-quality slides.  

Nemotron 3 Super is a great fit for these situations in software development. For example, Nemotron 3 Nano can handle single merge requests, while Nemotron 3 Super can handle more complex coding tasks that require a deeper understanding of the codebase. For expert-level coding, proprietary models are best.  

Building With Super’s Open Resources 

Nemotron 3 Super is fully open source, including model weights,datasets, and architectural recipes, enabling developers to tailor, refine, and deploy the model for privacy and security.  

Get Started 

With these capabilities in mind, getting started is straightforward. Start using Nemotron 3 Super today. Deploy it on your preferred platform, whether on a workstation or in the cloud. To experience its capabilities, sign up for a pro subscription on Perplexity. Access it directly by your API, use OpenRouter, or visit build.NVIDIA.COM. Take the next step. Explore Nemotron 3 Super Now! 

Source: Introducing Nemotron 3 Super: An Open Hybrid Mamba-Transformer MoE for Agentic Reasoning 

At Universe 2025, GitHub unveiled Agent HQ, a platform that serves as machine control for AI coding assistance from vendors like OpenAI, Anthropic, Google, and xAI within the GitHub ecosystem.  

This initiative, often referred to as an AI Army command center, aims to transition developers from juggling separate tools to coordinating a team of agents that seamlessly write, test, and debug code together.  

Main details of GitHub Agent HQ (2025-2026) 

  • Centralized Control Plane: Agent HQ provides a single interface in GitHub, VS Code, and the command line for assigning, tracking, and managing AI tasks in real time.  
  • Third-party integration: column developers can now use models from Anthropic Cloud 3.7 Sonnet, Google Labs Jules, and xAI in their GitHub workflow, not just GitHub by Copilot.  
  • Agents can now perform tasks in sequence independently, starting with picking up issues.  
  • creating branches  
  • committing code  
  • opening pull requests  
  • Human developers then review the agents’ work and provide feedback as needed.  
  • Enterprise Governance: The platform provides advanced code review by agents, a control panel to manage agent actions, and a dashboard to track AI performance.  
  • Following the October 2025 announcement, third-party agents gradually became available to GitHub Copilot subscribers. Over the next few months, advanced features will be offered through Copilot Pro + or to enterprise clients.  

Shift To Agentic Development 

GitHub COO Kyle Daigle stated that the objective is to bring order to the condition caused by rapid AI growth. Agent HQ enables developers to move beyond basic chat-based assistants by leveraging agents for more structured, step-by-step programming assignments.  

Related security concerns (EchoLeak)  

In early 2025, researchers identified the first zero-click AI vulnerability in Microsoft’s broader Copilot ecosystem, though not an Agent HQ-specific one. It highlighted the risks of AI agents accessing sensitive data. Microsoft responded with stronger security and auditing in its Frontier Suite (Microsoft 365 E7), launched in early 2026.  

If you maintain open source projects or work on an enterprise team, seeing automated documentation fixes, new unit tests, or refactoring suggestions can be a real eye-opener. Still, automation raises a key question: how do you set limits on agents that can access your repository and the internet? You could worry about an agent using information from unreliable websites, or accidentally exposing an API token, or maybe it could start posting unnecessary comments on every open issue. For its automation to be truly valuable, it needs to be predictable.  

What is the safest way to add agents to existing automations like CI/CD? Agents are unpredictable and handle untrusted inputs. Examine your repository’s state and make decisions as they run. Along with agents in CI/CD with constant oversight, you can scale your engineering, but it requires safeguards to address security risks.  

GitHub agentic workflows are built on GitHub Actions. Normally, everything in an action shares the same level of trust. This means an unauthorized agent could interfere with MCP servers, access authentication secrets, or send network requests to any destination. If an agent has bugs, is manipulated by prompts, and has no restrictions, it could behave in unforeseen and unsafe ways.  

This is why security is a core part of Agentic Hub workflows. We see agent execution as an extension of the CI/CD model, not as something separate. We keep the creative part of building workflows apart from the control part of running them. Then, we turn workflow into a GitHub action with clear limits on permissions, inputs, audit records, and network access.  

In this post, we will explain how we designed Agentic workflows to be secure from the start, starting with the threat model and needed security architecture.  

Threat Model 

Two key features of agentic workflows affect the threat model for automation.  

Agents can understand repository state and act independently. While useful, they should not be trusted by default, especially with untrusted inputs.  

Second, GitHub Actions offer a very open execution environment. Sharing a trust domain helps with automation, broad access, and good performance; however, if untrusted agents are involved, a single trust domain can lead to extensive problems if something fails.  

With this model, we assume agents may access or modify unauthorized data, use or misuse channels, or perform actions beyond their permissions through deferred GitHub agentic workflows. Use strict security settings based on this threat model, adhering to four security principles:  

  1. Defense In-Depth  
  1. Not Trusting Agents With Secrets  
  1. Reviewing all writes  
  1. Comprehensive Logging.  

Defend in Depth 

GitHub Agentic workflows use a layered security system with state configuration and planning layers. Each layer helps limit the impact of failures in the layers above by enforcing its own security rules.  

The Substrate Layer is built on a GitHub Access Runner, running on a virtual machine, with several trusted containers that control which resources an agent can use. This layer keeps components separate, manages privileged operations and system calls, and enforces communication boundaries at the kernel level. These predictions remain valid even if an untrusted component is compromised and runs code within its container.  

On top of the substrate layer is the configuration layer. This layer uses declarative artifacts and toolchains to set up a secure system and its connections. It decides which components are loaded, how they connect, which communication channels are allowed, and what privileges each has. External tokens, such as agent API keys and GitHub access tokens, are important inputs. The configuration controls which tokens are placed in which containers.  

The last layer of defense is the planning layer. While the configuration layer decides which components exist and how they connect, it does not control when they are active. The planning layer’s main job is to set up a staged workflow with clear data exchanges between components. The Safe Outputs subsystem, explained later, is the main example of secure planning.  

Don’t Trust Agents Bearing Secrets 

From the start, we aimed for workflow agents to have no access to secrets and to maintain strict trust boundaries. Agentic workflows run as GitHub actions, with all components sharing a single trust domain on the runner VM. In this setup, sensitive items such as agent authentication tokens and MCP server API keys are stored in environment variables and configuration files that all processes in the VM can access. No extra measures are required to prevent agents from breaching these trust boundaries.  

This is risky because agents can fall victim to prompt injection. Attackers might cause harmful impacts, such as web page or repository issues, that trick agents into revealing sensitive information. For example, an agent affected by prompt injection and with access to shell commands could read configuration files, SSH keys, LNS/PROC state, and workflow logs to find credentials and other secrets. It could then upload these secrets online or hide them in public GitHub objects, such as issues, pull requests, and comments.  

Our first step to reduce risk was to put the agent in its own container and to implement strict controls on what it can access. This includes:  

  • Firewall internet access  
  • MCP access only through a trusted gateway  
  • NLM API calls are routed through an API proxy to limit internet access  

Agentic workflows set up a private network between the agent and the firewall. The MCP gateway runs in a separate trusted container, starts MCP servers, and is the only one with access to MCP authentication material.  

Agents like Cloud, Codex, and Copilot need to talk to an LLM over a secure channel, but we do not give these tokens directly to the agent’s container. Instead, we keep LLM auth tokens in a separate API proxy and set up agents to send modern traffic through that proxy.  

Zero-Secret Agents need a balance between security and usefulness. Programming tasks often need access to compilers/interpreters/scripts/repository data. However, increasing the container setup would duplicate existing provisioning steps and add more network destinations to the five-where rules.  

Instead, we use container volume mounts to give the agent access to needed host files and programs, and we run it in a chroot jail. First, we mount the whole VM. The host system has a read-only /host. Then we cover certain paths with empty tmpfs layers and start the agent in a chroot jail at /host. This way, the host setup stays unchanged, and the agent can only read and write what it needs for its work.  

Stage and Vet all Writes 

Even without access to secrets, prompt-injected agents can still cause problems. For example, an agent’s interest might flood a repository with unnecessary issues or pull requests to overwhelm maintenance, or add unwanted URLs and other content to repository objects.  

To prevent this kind of behavior, the Agentic Workflows Compiler decomposes every workflow into clear, explicit stages. It acts as a control point, defining for each stage.  

  • The Active Components and Permissions (read vs write)  
  • The data artifacts emitted by that stage  
  • The admissible downstream consumers of those artifacts  

While the agent runs, it can read GitHub state through the GitHub MCP server and can only prepare its updates through the safe outputs MCP server. After the agent finishes the safe outputs, the MCP server processes any buffered write operations using a set of safe output checks. It includes operations that an agent can perform. Authors can choose which GitHub update types are available, such as:  

  1. Creating issues, comments, or pull requests  
  1. Safe outputs limit the number of updates allowed, such as restricting an agent to creating at most three pull requests per run.  
  1. Safe outputs analyze and update content to remove unwanted patterns, such as sanitizing URLs  

Only artifacts that pass through the entire safe outputs pipeline can be passed on, making sure that each stage’s side effects are explicit and vetted.  

Log Everything. 

Even with no secrets and checked rights, an agent can still change repository data, use tools in ways we did not expect, or try to get around the limits we set. Agents will try many tricks to complete their tasks. If something goes wrong, we need to see the full execution path to understand what happened.  

Agentic workflows make observability a first-class property of the architecture by logging extensively at each trust boundary. Network and destination-level activity is recorded at the five one-layer model request/response metadata, and authenticated requests are captured by the API proxy. All invocations are logged by the MCP gateway and MCP servers. We also have an internal implementation in the agent container to audit potentially sensitive actions such as access to environment variables. Together, these logs support end-to-end forensic reconstruction, policy validation, and rapid detection of anomalous agent behavior.  

Extensive logging also sets the stage for future information flow controls. Anyway, we observe communication; we can control it. Agentic workflows already support GitHub MCP servers’ lockdown mode. In the coming months, we will add more safety controls that enforce policies across MCP servers based on whether something is public or private and who created a repository object.  

What’s Next? 

Join the discussion in our community or on the #GitHubNext Discord. We look forward to seeing what you build with GitHub Agentic Workflows. Stay tuned for more updates.

Source: Under the hood: Security architecture of GitHub Agentic Workflows 

Tesla has advanced general-purpose robotics. The latest Optimus AI update adds vision-language navigation, enabling the robot to reason. By merging language understanding with spatial cognition, Tesla addresses the main challenge of deploying humanoid robots: executing complex real-world instructions.  

For robotics engineers and AI researchers, this update marks a shift from traditional SLAM methods to a more comprehensive, embodied AI approach for humanoid robots. Now the focus is not just on avoiding obstacles but on helping the robot comprehend its environment using human language.  

The shift to vision-language navigation (VLN) for robotics engineers and AI researchers. This update marks a shift from traditional SLAM methods to a more comprehensive embodied AI approach for humanoid robots. Now the focus is not just on avoiding obstacles but on helping the robot comprehend its environment using human language.  

The Shift to Vision-Language Navigation (VLN) 

Historically, autonomous navigation was a geometric problem. Robots use LiDAR-based vision to create a voxel map of the world and navigate to specific coordinates. However, coordinates are not how humans communicate. We do not tell a co-worker to move to 45.2-12.8 in. We say, “Take the red folder from the messy desk and bring it to the lounge near the coffee machine.”  

With Vision-Language Navigation (VLN), Optimus can now understand these kinds of instructions. The new AI uses a transformer model (a type of neural network, especially good at understanding language and images) that processes video from the robot’s eight cameras along with language input. This lets the robot find objects or rooms it hasn’t seen before by matching what it sees to the words it hears.  

Embodied AI: The Fusion of Logic and Limbs 

This update focuses on embodied AI, meaning the robot’s intelligence is integrated with its physical form. Unlike pure text-based models, a humanoid robot must interact with the physical world and obey its laws. Tesla has redesigned its FSD for robots to enable detailed step-by-step reasoning about space and time, allowing Optimus to plan and act within its environment.  

When Optimus receives a command, the vision language model first breaks the task into sub-goals. If the goal is to clean up the spill in the lab, the robot must identify it using its vision system. Understand that cleanup requires a tool, such as a mop or paper towels. Use language/logic) and then navigate to where those items are typically stored (using memory and spatial reasoning, or the ability to recall and understand places). By running this logic locally on Tesla’s D1 chip, the robot achieves sub-millisecond latency to adjust its balance and gait while simultaneously processing high-level cognitive tasks.  

Mastering Active Environments with World Models 

A major challenge for humanoid robots is that human spaces change constantly. Factories, homes, and offices are never static. The new AI stack uses Neural World Models to help Optimus predict possible changes based on past data.  

If a human walks across the robot’s path, Optimus does not simply stop. It predicts the person’s path and adjusts their velocity and path in real time. This is where the vision-language component becomes critical for safety and social etiquette. The robot can distinguish between a stationary object, such as a box, and a temporary obstruction, such as a person, and chooses a wider berth to ensure people’s comfort. This subtle behavior is a direct result of training the navigation stack on millions of hours of human-human interaction data, allowing the robot to emulate natural spatial social norms.  

The Role of End-to-End Neural Networks 

Tesla is committed to an end-to-end approach. While others use separate modules for vision, planning, and movement, Optimus depends on a single large neural network. The Vision Language Navigation update feeds raw data, images, and text directly into this network, which controls the robot’s actions.  

This approach provides for emergent problem-solving. During recent internal testing, an Optimus unit was tasked with moving a crate that was blocked by a rolling chair. Rather than failing or waiting for the path to clear, the robot used its vision-language understanding to recognize the chair as a movable object, pushed it out of the way, and proceeded to its goal. This type of reasoning, identifying affordances in the environment and seeing what actions an object allows (such as a chair’s mobility), is the hallmark of true humanoid autonomy.  

Scaling Through The Dojo Training Fabric 

Tesla’s Dojo supercomputer drives Optimus’s advanced embodied AI. To train vision-language navigation, Tesla uses a special auto-labeling system. Thousands of Optimus robots in factories collect data. When a robot encounters a new situation or tricky instructions, it sends the data to Dojo.  

There is a larger teacher model that analyzes the Dojo video. A bigger teacher model reviews the video and results, then labels the data for the student model on the robot. This cycle makes the navigation system stronger every day. In 2026, Tesla began using generative world simulations in which Dojo creates millions of challenging scenarios, such as a robot in a dark room with mirrors or a busy hospital hallway, to test the VLN system before it’s used in real robots. The technical ability to move forward with vision-language navigation is an economic strategy that makes the robot easier to perform via voice or text.  

Tesla is reducing the barrier to entry for small-scale manufacturing and elder care facilities. You no longer need a staff of robotics engineers to define waypoints or no-go zones. A floor manager can simply walk the robot through a facility, giving verbal indications such as “this is the shipping dock” and “don’t enter this area during shift changes”. The robot’s VLN stack will build a semantic map that adheres to those rules.  

Tesla believes accessible robotics will help Optimus reach millions of users. When using a robot is as simple as conversation, it becomes an everyday workplace tool, not a luxury.  

The Road Ahead: General Purpose Intelligence 

Adding Vision Language Navigation to the Optimus AI stack is a step toward Tesla’s goal of Artificial General Intelligence. While a chatbot explains a recipe, Optimus is getting closer to seeing the ingredients, understanding the recipe, and completing the task.  

Looking to 2026, integrating vision and language will drive social robotics. Optimus will move through our world and communicate, saying things like, “Excuse me, I need to reach that shelf,” or, “I have completed the inventory check.” This will ease collaboration between people and robots.  

Final Thoughts: The Humanoid Constitution 

Tesla’s vision for Optimus has always been bold, but Tesla’s big plans for Optimus are now becoming real with the latest AI update, which adds vision-language navigation. Tesla is reaching, teaching the robots to see and hear the world as we do. This marks the start of the Autonomous Digital Coworker, a machine that understands not just what to do, but also how and why. The general-purpose humanoid is no longer simply an idea; it’s already working on factory floors.

SourceAI & Robotics