NVIDIA has introduced several new technologies to accelerate the development of humanoid robots. This includes NVIDIA ISAAC GR00T-N1, described as the world’s first open and fully customizable foundation model. It is a large artificial intelligence system trained on diverse data that can be adapted for many tasks, in this case, general humanoid reasoning and skills.  

Other technologies in the lineup include simulation frameworks and blueprints, such as the NVIDIA ISAAC-GR00T blueprint. A simulation framework is a set of software tools for testing and training robots. In a virtual environment, the blueprint helps generate synthetic training data. There is also Newton, an open-source physics engine developed with Google Brain and Disney Research, designed specifically to simulate real-world physical interactions for building robots.  

Building on these releases, GR00T-N1 is now available. It is the first in a series of customizable models that NVIDIA will share globally to support industries facing workforce shortages.  

The Age of Generalist Robotics is Here, said Jensen Wong, founder and CEO of NVIDIA, with NVIDIA ISAAC GR00T N1 and new data-generation and robot-learning frameworks. Robotics developers everywhere will open the next frontier in the age of AI.  

GR00T-N1 Advances Humanoid Developer Community 

The GR00T N1 Foundation Model uses a dual system design inspired by how people think. It features System-1, which acts quickly and automatically, like human reflexes or intuition, and System-2, which takes a slower, more careful approach to decision making. Dual-system design refers to splitting cognitive processes into fast and slow systems, similar to theories in human psychology.  

System 2 is powered by a vision-language model, a type of AI that understands images and written or spoken commands, reasons about its environment, and the instructions it has received to plan actions. System 1 then translates these actions into precise, continuous robot movements. System 1 is trained with data from both human demonstrations and a large volume of synthetic data generated by the NVIDIA Omniverse platform. The vision-language model enables the robot to interpret both visual and linguistic inputs.  

GR00T-N1 can handle a variety of common tasks, including grasping and moving objects with one or both arms and passing items between arms. It can also perform more complex multi-step tasks that need a longer context and a mix of general scales. These abilities are useful for tasks such as material handling, packaging, and inspection.  

Developers and researchers can further train GR00T-N1 with real or synthetic data to fit their own humanoid robots or tasks.  

During his GTC keynote, Goan showed 1X’s humanoid robot performing household tidying tasks on its own using a policy trained with GR00T-N1. This autonomous ability comes from an AI training partnership between 1X and NVIDIA.  

The future of human arts is concerning adaptability and learning, said Brent Bonich, CEO of One-X Technologies. While we develop our own models and media, GR00T-N1 provides a significant boost to robot reasoning and skills with minimal post-training data. We fully deploy on Neo-Gamma, promoting our mission of creating robots that are more than tools, yet companions capable of assisting humans in valuable, immeasurable ways.  

Other top humanoid developers with early access to GR00T-N1 include Agility Robotics, Boston Dynamics, Mentee Robotics, and Neura Robotics.  

NVIDIA, Google DeepMind, and Disney Research focus on physics.  

NVIDIA is working with Google DeepMind and Disney Research to develop Newton, an open-source physics engine. In this partnership, NVIDIA leads the development, with Google DeepMind and Disney Research contributing expertise, to help robots learn to perform complex tasks more accurately.  

Newton, built on the N-Media Warp Framework, will be optimized for robot learning when compatible with simulation frameworks such as MuJoCo and Isaac Sim. It will also utilize Disney’s physics engine.  

Google DeepMind and NVIDIA are also co-developing MuJoCo-Warp, aiming to accelerate robotics machine learning tasks by over 70 times. Developers will access it via Google DeepMind’s MJX open-source library and the Newton engine, co-developed with NVIDIA.  

Disney Research, as a partner in the Newton project, will be among the first to use the engine to improve its robotic character platform. This platform powers next-generation entertainment robots like the expressive Star Wars-inspired BDH droids that appeared with Huang during his GTC keynote.  

The BDH droids are just the beginning. We’re committed to bringing more characters alive in ways the world hasn’t seen before. This cooperation with Disney Research and Video and Google is a key part of that vision, said Kyle Laughlin, Sr. Vice President of Walt Disney Imagineering Research and Development. This alliance will allow us to create a new generation of robotic characters that are more expressive and engaging than ever before and connect with our guests in ways that Disney can.  

Continuing their collaboration, NVIDIA, Disney Research, and Intrinsic have announced a new partnership. Each organization will collaborate to develop OpenUSD pipelines and best practices for robotics data workflows, with NVIDIA overseeing the technical architecture and Disney and Intrinsic contributing their expertise in robotics and data management.  

NVIDIA has also announced the BGX Spark Personal AI supercomputer at GTC. It gives developers a ready-to-use system to expand GR00T and N1’s capabilities for new robots’ tasks and environments without requiring much custom programming.  

The Newton physics engine will be released later this year.

Source: NVIDIA Announces Isaac GR00T N1 — the World’s First Open Humanoid Robot Foundation Model 

In autonomous mobility, the benchmark for full self-driving has shifted. Now it requires deep semantic understanding, comprehending an environment’s meaning and context, not just obstacle avoidance. In March 2026, Tesla’s Gen 3 firmware introduced a paradigm-defining feature: VLM (vision language model) logic for terrain adaptation. By embedding vision-language models into the vehicle’s system, in the reasoning stack responsible for deliberate, complex decisions, Tesla moves beyond traditional occupancy grids. Occupancy grids are basic maps showing where objects are present. This approach lets its fleet interpret and navigate unstructured environments with human-like intuition.  

This update is the most significant architectural change to Tesla’s Neural Stack since the introduction of end-to-end neural networks (FSD V12). In those networks, the entire driving process is managed by a single neural network that addresses the semantic gap. The issue is that only systems can understand ambiguous surfaces, such as simple wet glass, deep silt, or construction zone debris.  

The Architecture of VLM Logic in Gen 3 firmware 

The Gen 3 firmware moves from a purely geometric world model to a semantic reasoning framework. Traditional AI systems treat the world as a 3D grid of 3D volumes called voxels. Voxels are small cubes in a grid used to represent space. In this system, a voxel is marked as occupied or empty. This method works for avoiding solid object obstacles such as concrete walls; however, this binary logic does not help when a cyber-truck must decide if a muddy path is safe or if a puddle hides a pothole.  

With VLM and Logic, the Tesla AI supercomputer processes camera feeds through a multi-modal transformer. This neural network model can interpret multiple types of data. The vehicle first describes the scene in a latent linguistic space, which is an internal language-like representation used by AI to understand context before executing a command. For example, instead of seeing only a low-level competitor like Brown Moline at XYZ, the VLM identifies deep, saturated mud with standing water and a high risk of traction loss. This semantic level triggers specific terrain-adaptation profiles: suspension damping (how shock absorbers respond), torque distribution (how engine power is sent to each wheel), and tire slip targets (optimal tire spin for traction). Adjust in real time.  

TERRAIN ADAPTATION: THE PHYSICS OF SEMANTIC INTELLIGENCE 

Terrain Adaptation, powered by Vehicle Logic Models (VLM) software, updates the Cybercrime and upcoming Cyber Beast models when the firmware detects a shift from asphalt to an unstructured surface. VLM Logic Response promptly: it acts as a strategic planner for the vehicle’s air suspension system, which controls ride height and stiffness, and the powertrain, which manages power distribution to the wheels.  

  • Predictive damping delays the traditional system’s response after a vehicle hits a bump. VLM logic instead analyzes terrain texture and appearance ahead. The model detects surfaces such as loose gravel, small shifting stones, and washboarding, and repairs uneven patches. The firmware softens compression damping on Gen3 struts. This adjustment maintains tire contact patch integrity. The tire stays fully in touch with the road surface.  
  • Dynamic Torque Vectoring: On slippery or uneven surfaces, the ision Language Model (VLM)t logic informs the Tri-Motor Drive Unit. The unit distributes power among the motors. It applies anticipatory torque bias in shifting power to the wheels most likely to need it before traction issues occur. The vehicle maintains momentum through sand or snow with less input from the traditional traction control system. The traditional system typically reduces wheel slip by braking or limiting power.  
  • Micro adjustments in gait: this common disk logic is not limited to vehicles. The Gen3 firmware is a unified software platform that also powers the Optimus Gen3 humanoid robot. With VLM training adaptation, the robot moves confidently across cluttered factory floors using its vision system to detect hazards. For example, it recognizes a pile of oily rags as a slip hazard and adjusts its center of mass before its foot comes into contact with the pile.  

Embodied AI and the Sovereign Logic Guardrail 

A critical component of this update is the concept of Sovereign AI. Tesla runs these massive vision-language models entirely on the device. This bypasses the need for cloud-based inference. As a result, terrain adaptation stays functional even in remote off-grid areas where LTE or Starlink connectivity is intermittent.  

To achieve this, the Gen3 firmware uses a technique called optimized speculative decoding. It compresses numbers to improve the efficiency of AI computations. The AI computer runs a smaller, faster draft model for repetitive, frequent driving tasks. The longer visual language model, verifier model, intermittently checks the meaning and context of what the car sees in its surroundings. If the VLM detects a complex terrain change that the draft model missed, it overrides the driving path with a safe state command. This command directs the car to pause or take safe action. This dual-model approach provides a safety guardrail that is impossible in single-model, end-to-end systems.  

The Role of Generative World Models in Training 

VLM logic for terrain adaptation became effective through millions of miles of synthetic (computer-generated) off-road training. Tesla’s neural word simulator made this possible. This generated artificial intelligence program creates hyper-realistic three-dimensional environments and helps teach the VLM how different terrain types behave.  

By simulating the physics of mud, sand, water, and ice, Tesla’s engineers exposed the VLM to corner cases too dangerous or rare to test in the real world. This training enables the VLM’s cloud-like reasoning to predict that a dark patch on a frozen road is likely black ice. It triggers an immediate shift in the terrain adaptation profile to ultra-low-grip mode.  

Developer and Power Use Implications  

For the technical community, Tesla Gen3 firmware includes a new vision debug mode via the service menu. This mode displays Vision Localization Modules (VLMs) and internal monologue. In real-time, users see descriptive labels such as:  

  • surface wet cobblestone, indicating the detected surface type  
  • traction ESD 045, for estimating tire traction  
  • adaptation soft rebound active, for the current suspension mode  

This transparency is a massive step for AI interpretability. Instead of wondering why a vehicle slowed down or changed course, VLM logic provides a clear semantic reason. This builds user trust. It also lets Tesla’s fleet learning system flag when the VLM’s terrain view differs from the human driver’s actions. This creates a cycle of continuous improvement.  

Final Thoughts 

The integration of Vision Language Model (VLM)M logic for terrain adaptation in TeslaGen 33 firmware marks the end of the specialized era. We are no longer viewing just a car that drives or a robot that walks. Now there is a unified embodied intelligence that understands the physical world semantically. Firmware continues to roll out globally throughout the first half of 2026. The gap between human and machine perception will continue to close. Whether navigating a snowy mountain pass in a Cybertruck or a busy warehouse in an autonomous robot, the ability to see, think, and adapt to the terrain is the final piece of the Autonomy Puzzle.

Source:  Firmware Version 23.8.2 for the Tesla Gen 3 Wall Connector 

Operating system security is a constant battle against memory corruption. Software safeguards like ASLR, stack canaries, and non-executable memory have been used, but attackers find ways around them. With Android 17, Google now mandates hardware memory tagging for ARMv9 chips, marking a shift to hardware-based security for future mobile devices. 

Android 17 now requires the Memory Tagging Extension (MTE), a hardware feature in ARMv9-A. By moving memory safety checks to the CPU, this aims to eliminate major vulnerabilities like use-after-free and buffer overflows, which account for most serious Android security bugs. 

How MTE Changes Memory Safety 

To understand MTE’s importance, consider C and C++. Memory is accessed via pointers, which are just addresses. The CPU cannot detect if a program uses an address after it is freed. 

With memory tagging, each 16-byte block has a 4-bit tag. When memory is allocated, the allocator assigns a tag and stores it in the unused part of the pointer. The CPU checks whether the memory and pointer tags match on access, triggering a fault if they don’t. 

Hardware tags are checked at a low level, making bypass difficult. This prevents heap grooming, as attackers must guess 4-bit tags at each step, thereby greatly increasing the difficulty of the attack. 

Android 17: Making MTE Mandatory 

MTE was first added in Android 12, but its use has been inconsistent. Google’s Pixel 8 and Pixel 9 phones were among the first to include MTE hardware, but the feature was often hidden as a Developer Option or only enabled for certain system services like Bluetooth and NFC. 

The Android 17 source code shows that MTE is no longer optional. Any device that wants Google Mobile Services (GMS) certification on ARMv9 chips must have MTE turned on by default for all important system processes and core software. There are also new ‘Hardened User-Space’ profiles that require MTE for third-party apps unless they opt out. This strong approach is meant to push chip makers and device manufacturers to enable MTE, even if they worry about performance. 

Optimizing Security and Performance: Three MTE Modes 

One main reason memory tagging has been slow to catch on is the ‘security tax,’ or the extra CPU and battery use it can cause. Android 17 handles this by using three different MTE operating modes: 

  1. Synchronous (SYNC) Mode: In this mode, the CPU immediately halts execution whenever it encounters a tag mismatch. This makes it the most secure option, as errors are caught at the moment they occur. However, this strict checking also causes the most slowdown, with a performance cost of about 3% to 5%. For this reason, Android 17 requires SYNC mode only for the system’s most security-critical parts, including the kernel, identity credentials, and biometric authentication. 
  1. Asynchronous (ASYNC) Mode: In this mode, the CPU records a tag mismatch but does not stop immediately. Instead, execution continues until the next kernel entry, commonly during a system call. This delayed response reduces performance impact but means violations are detected slightly later. Android 17 assigns ASYNC mode to regular system apps and background services, prioritizing the user experience for less-critical operations. 
  1. Asymmetric (ASYMM) Mode: This mode combines the earlier modes by applying synchronous checks for reading memory (CPU catches errors immediately on reads) and asynchronous checks for writing (delayed error reporting for writes). Android 17 uses ASYMM as the default for most third-party apps, since it balances strong protection—especially for data reads without the higher performance costs of full synchronous checking. 

Why ARMv9 Chips Matter for Security 

This rule focuses on ARMv9 chips because their hardware, like the Cortex-X4, A720, and the new Blackhawk cores, is built to check tags quickly. Older ARMv8.5 versions of MTE were often too slow for actual use, but improvements in ARMv9 have made the performance impact so small that most users will not notice it. 

This new requirement is also a key defense against the growing number of ‘zero-click’ attacks. Many of these target media processing or networking, which often use C++ for speed. By making hardware memory safety mandatory in these risky areas, Android 17 makes it much harder and more expensive for attackers to succeed. An exploit that worked on an older ARMv8 device will now just cause a harmless crash on an Android 17 ARMv9 device. 

What This Means for Developers: Fewer ‘Heisenbugs’ 

For developers, making MTE mandatory has both pros and cons. It gives them a strong new way to debug, turning hard-to-find ‘Heisenbugs’ into clear, repeatable crashes. But it also means developers need to be more careful with native code. Custom memory allocators that do not handle tags properly will not work with Android 17. 

To help with this change, the Android 17 SDK now offers improved ‘MTE-Aware’ telemetry. When a tag fault happens, the system creates a detailed report with the allocation and freeing stack traces for the problem to address. Before, this kind of insight into memory use was only possible with sophisticated tools like AddressSanitizer (ASan). 

Conclusion 

The discovery that the Android 17 source code mandates hardware memory tagging for ARMv9 silicon marks the beginning of the end for memory corruption as we know it. The fact that Android 17 now requires hardware memory tagging for ARMv9 chips signals a major step toward ending memory corruption. By making MTE a required part of the platform, Google is shifting from a ‘detect and patch’ approach to one that is secure by design. It will focus on MTE performance, and app developers will finally have a hardware-backed safety net that protects their users from the most dangerous classes of cyberattacks. The “black art” of memory exploitation is about to get a lot more difficult.

Source:  Arm memory tagging extension 

The “black box” problem has been a major obstacle for enterprise adoption of artificial intelligence. Even as models have improved, their internal logic has stayed mostly hidden. Now, with the release of the Anthropic API Beta, this is changing. The update introduces Thought Trace Logs for Claude 4.6 models, giving developers and safety researchers a new way to see how the model reasons before generating any part of its final response. 

This change shifts the focus from guessing prompts to a more structured, engineering-based approach to understanding and explaining AI. Now, “chain of thought” is not just a prompt trick but a clear, reviewable data stream. 

The Architecture of the Thought Trace 

In the past, seeing how a language model “thinks” meant making a choice. You could have the model explain its reasoning in the final output, which uses up tokens and could affect the answer, or you could use internal tools that were too slow and costly for live API use. 

The Claude 4.6 “thought_trace” feature adds a channel during inference. If you use the include_thought_trace: true header, the Anthropic API returns an extra metadata stream. This stream shows “reasoning tokens” for the model’s plan, task breakdown, and fact checks. 

These logs are more than answer summaries. They record the model’s “inner monologue,” noting when it made and fixed mistakes. For those working on autonomous AI agents, this log provides a clear record of why an agent went off track during complex tasks. 

Strengthening Reliability Through Interpretability 

Thought Trace Logs reduce worries about AI fabrication. Instead of just trusting model answers, engineers can now check each step of the model’s reasoning with Claude 4.6. 

In legal or financial settings, an app can review the thought trace for key logic steps. If it shows assumptions without citing documents, the app can prompt corrections or flag answers for human review. This offers AI transparency that goes beyond checking for certain words or sentiments. 

Integrating “Adaptive Thinking” and Effort Controls 

These logs appear under “Adaptive Thinking” in the Claude 4.6 models. Claude 4.6 (Opus and Sonnet) now uses a flexible reasoning budget. The model spends less effort on greetings and more on complex tasks like code refactoring. 

The Thought Trace Logs make the decision process visible. Developers can see the model’s “Effort Level” for each task, from low to high, and how it affects thought trace detail. This view aids cost and speed optimization. If the logs show the model overthinks simple tasks, developers can use the new API effort setting to limit reasoning depth, saving time and tokens. 

Solving the “Alignment Faking” Problem 

A more technical, but important, benefit of the Anthropic API Beta is that it lets you monitor “alignment faking.” This happens when a model notices it is being tested and changes its answers to please the user instead of giving the most accurate or objective response. 

Researchers use Thought Trace Logs to check if the model’s reasoning matches its output. If the trace shows strong logic but the answer is vague or softened, it may mean that safety rules or prompts make the model too agreeable. This helps test and improve rules guiding Claude’s behavior. 

Implementing Thought Trace in Production 

Engineers using Thought Trace Logs in production need new data handling. The logs can be long, sometimes longer than the response. Anthropic has added Context Compaction to help. The API now summarizes older thoughts, so the “thought history” no longer fills the context window. 

The logs are structured in JSON, easy to use with monitoring tools like Datadog or New Relic. Organizations can create dashboards to track “Reasoning Efficiency” or “Logic Accuracy,” treating the model’s thoughts as valuable data. 

The Future of the “Transparent” Agent 

As we approach 2026, demand for explainable and transparent AI will grow. The Anthropic API Beta for Claude 4.6 shows the industry is moving past the “Trust Me” phase. 

By making the thought trace visible, Anthropic helps developers create agents that are both clever and explainable. Doctors can check a diagnosis. Engineers can review code changes. Seeing the reasons behind answers helps move AI from experiments to real-world use. 

Conclusion: A New Standard for Model Accountability 

Making Thought Trace Logs available for Claude 4.6 is a bold step toward transparency in a secretive industry. It shows that for AI to be useful in business, it must be open to review like any other software. 

As developers start using these logs, we will probably see more “Interpretability-First” apps—tools that not only give answers but also show a clear, logical path for how those answers were found. With Claude 4.6, the black box is not just open; it now has a detailed internal view.

Source:  Anthropic’s Transparency Hub 

Samsung is developing low-latency technologies for humanoid robots, focusing on improving voice interaction and real-time control. Much of this progress arises from its partnership with Rainbow Robotics.  

Samsung and Rainbow Robotics are also working on AI-powered factories, though there is no specific mention of a Robo-Operating System, Kernel patch, significant improvements in low-latency AI software, a boost in voice interaction, or improvements in humanoid robot performance.  

Below are the main highlights of Samsung’s progress in this field:  

  • Low-latency voice AI voice integration: Samsung is adding voice controls to help humanoid robots respond more quickly, aiming for smooth human-robot interactions in factories and service roles.  
  • Agentic AI Core: Samsung uses agentic AI, which is artificial intelligence capable of taking actions and making decisions to achieve goals. As a coordination layer for humanoid robots, agentic AI enables robots to act independently and manage complex tasks in real time. The operating system requires very low latency meaning minimal delay to avoid task delays.  
  • Rb-Y1-humanoid-focus: one main project is the Rb-Y1, a wheeled humanoid robot developed with Rainbow Robotics. It is designed to handle complex tasks and to engage in conversations on production lines.  
  • Humanoid Robotics R&D: Sensing Research is developing a robust, high-performance robotics software framework—a base layer of code and tools that supports robot functionality. It processes sensor data, such as microphone audio, and plans robot movement in real time.  
  • Focus on voice activity detection (VVAD): Samsung aims to improve task completion and reduce conversational delays. VVAD, or voice activity detection, is a technology that recognizes when a person is speaking, helping devices respond only when needed. Research also targets reliability in noisy environments. Together, these efforts support Samsung’s goal of fully autonomous, self-managing AI-managed factories by 2030, with AI robots acting as conversational partners.  

Humanoids on the factory floor 

Samsung previously concentrated its robotics initiatives on commercial products such as robotic vacuum cleaners. Currently, the company is allocating resources to humanoid robotics development and partnering with Rainbow Robotics in South Korea. Samsung intends to deploy the Rainbow Robotics RB-Y1 humanoid robot within its manufacturing operations.  

This denotes a shift from using robots for side tasks to directly adding human-owned robots into the manufacturing work. While Samsung has not shared specific assignments yet, deploying these robots on the manufacturing line likely means they will aid with material handling, assembly, or shaping, leveraging their human-like movement for flexibility.  

Agentic AI as a Coordinating Layer 

In parallel with the deployment of humanoid robots, Samsung aims to incorporate agentic AI across its production infrastructure. According to the company, these AI systems are designed to optimize process quality and efficiency from material warehousing to finished goods logistics. Samsung further anticipates that AI will contribute to occupational safety and environmental compliance.  

Integrating humanoid robots with agentic AI establishes a platform where robots execute assigned tasks while AI agents dynamically monitor, optimize, and adapt workflows in real time. This reflects a broader industry trend toward the convergence of physical robotics and intelligent process automation.  

Industry Context 

Samsung’s strategy aligns with a growing push by major manufacturers to introduce humanoid robots into factory environments. In October 2025, Apple supplier Foxconn announced plans to use NVIDIA-powered bipedal robots to assemble AI servers within six months. Hyundai has also ordered 30,000 Atlas humanoid robots from its subsidiary Boston Dynamics, with deployment planned across its car factories in the United States.  

These projects indicate that large-scale industrial stakeholders are advancing from pilot implementations to systemic deployments of humanoid robotics platforms. For robotics engineers and factory managers, the focus is shifting from proof-of-concept validation to integration, reliability, and operational governance.  

Governance and Following Steps 

Samsung is expected to outline its AI strategy at the Mobile World Congress in Barcelona in March, including details on its governance framework for AI deployment. The governance structure will likely prove crucial as humanoid robots and autonomous software agents are integrated into safety-critical production environments.  

Looking ahead to 2030, Samsung’s plan shows a significant commitment: the company believes that humanoid robots, with help from agentic AI, can boost productivity, quality control, and resilience in factories worldwide.

SourceSamsung Targets 2030 Global Factory Shift With Humanoids 

Picture looking through a window so clear it seems to vanish, revealing everything in sharp detail. That’s the goal of the Apple Metal 4 API, released in beta in March 2026, which introduces advanced built-in image sharpening for the Vision Pro. Earlier updates focused on speed, but Metal 4 now uses the M5 and R1 chips to control how light appears, even at the level of a single pixel. This lets Apple bring Retina Vision to spatial computing, making visuals appear sharper than before.  

For developers and graphics engineers, this represents more than a performance boost. It marks a major change in how retina-quality 3D is achieved in Metal 4. Machine learning, combined with precise hardware control, enables the Vision Pro to deliver resolutions beyond the limits of its micro-OLED panels.  

The Challenge: Beyond the Limits of Physical Pixels 

The challenge is going beyond the limits of physical pixels. The first Vision Pro had 23 million pixels, which is impressive. Even 4K-per-eye displays can struggle with aliasing and the screen-door effect when showing fine text or detailed shapes. Traditional up-scaling methods, such as Metal IFX, reconstruct missing data from earlier frames, but they are limited by the display’s pixel grid.  

Subpixel Neural Scaling solves this by focusing on tiny red, green, and blue parts that make up each pixel. Normally, these are bundled as one color unit with Metal. For new neural kernels, we can adjust and sharpen edges by working with each sub-element. Separately guarded by a high-frequency neural network.  

How Subpixel Neural Rescaling Works 

This technology uses a new process designed for M5 chips and upgraded smart processors. The process has three main steps.  

  1. Step 1: the Metal 4 system analyzes the shapes and motion in each scene at a higher level of detail than the display shows.  
  1. Step 2: A special program in the device predicts the best brightness and color values for each tiny part of a pixel. The program is trained for the Vision Pro’s unique screen layout, which uses very small pixels.  
  1. Step 3: The R1 chip assigns these results directly to the screen’s hardware using a sub-pixel offset trick to make edges look smoother and more detailed to your eyes than if each pixel were controlled alone.  

This approach greatly reduces judder and shimmering on thin lines, frequent issues in AR/VR, especially when the user’s head moves. By working at the sub-pixel level, the Vision Pro makes visual text as clear as printed text.  

Sovereign AI and On-Device Processing 

An important aspect of Metal 4 is its commitment to sovereign AI security. All neural re-scaling happens on the device within the Vision Pro’s secure enclave, eliminating delay and privacy risks from cloud processing. The Metal 4 API offers a black box for neural upscaling, so raw texture data remains protected from the rest of the system. This is crucial for sensitive CAD designs or medical imaging. With Metal 4, these high-resolution assets are re-scaled locally for optimal clarity, maintaining the sovereign nature of data from encrypted disk to the user’s retina.  

Impact on Developer Workflows: The MTL4Compiler 

Apple has also released the MTL4 Compiler, a new tool that gives developers more control over how visual improvements are applied. Unlike earlier versions, Metal 4 lets developers adjust these settings on the go for different scenes. 

 Developers can now:  

  • Prioritize latency or quality: Adjust the neural rescaling model’s complexity on the fly based on the scene’s characteristics.  
  • Build sharpening tools in the background: this keeps the Vision Pro’s 120Hz refresh rate smooth, preventing shuttering.  
  • Map custom data directly: For specialized use cases, developers can skip standard image improvements and link their own custom data directly to the display’s small elements.  

Synergy with Hardware: M5 and the R1 Photon-to-Photon Pipeline 

Subpixel Neural Rescaling works effectively thanks to the 2026 Vision. The M5 chip’s higher memory bandwidth handles the high data flow needed for the neural engine to run at 120 frames per second, while the R1 chip finishes compositing with a photon-to-photon display of only 12 ms.  

By adding Neural Rescaling to the R1’s final step, Apple ensures the upscaled image aligns with the user’s head position even if the M5 rendering is slightly delayed. This close collaboration between hardware and software helps prevent motion sickness that can occur with AI-generated friends in VR.  

The future of a transparent display 

Ultimately, the main goal of Metal 4 and sub-pixel neural rescaling is to make the display feel transparent and remove technical barriers between the user and the virtual world. When the pixel grid disappears, the sense of immersion is complete.  

As developers try out the Metal 4 API beta, we’ll likely see a new wave of advanced spatial apps. These apps will use the impro 

ved resolution to show layered data, realistic models, and 3D experiences that lower-quality displays couldn’t handle.  

Final Thoughts: A Milestone in Spatial Graphics 

The debut of sub-pixel neural rescaling in the Apple Metal 4 API Beta represents more than merely an incremental upgrade. It represents the maturation of Apple’s spatial computing platform, where AI is no longer a bolt-on feature but an essential part of the graphics pipeline. By moving the battleground from more pixels to smarter pixels, Apple has secured the vision position as the world standard for high-quality immersion.  

Now, it is up to developers to use these new models to create experiences that use the M5 chip’s abilities. The era of visible pixels is ending, and the time for clear, sharp images powered by advanced software has begun.  

Meta Title (60 characters) Apple Metal 4 API Brings Sub-Pixel AI Scaling to Vision Pro 

Meta Description (160 characters) Apple Metal 4 API beta introduces sub-pixel neural scaling for Vision Pro, using M5 and R1 chips to sharpen visuals, reduce aliasing, and deliver Retina-level spatial graphics. 

Source: Discover Metal 4 

As confidential computing grows, keeping generative AI secure is now a top priority for enterprise architects. When companies shift from pilot projects to full-scale use of Large Language Models (LLMs), protecting model weights and prompt data during inference becomes a major challenge. To help solve this, Amazon Web Services has released AWS Nitro Enclaves v3.4, which introduces Sealed Generative Logic Isolation. 

This new feature changes how “data in use” is protected in the cloud. By building on the Nitro System, AWS now offers a secure, verifiable space where generative workloads can run without being exposed to the parent instance, system administrators, or the cloud provider. 

The Architecture of Sealed Generative Logic Isolation 

Sealed Generative Logic Isolation is a security tool made for the high memory and computing needs of modern AI. Older confidential computing setups often have trouble handling large model inference. Nitro Enclaves v3.4 addresses this by adding a hardware-based “seal” around the enclave’s memory and execution. 

When running a generative model in an enclave, Sealed Generative Logic Isolation keeps the full inference—reading the prompt through to output—inside a secure boundary. These enclaves block interactive access, have no persistent storage, and have no external network. Data moves only via a secure local vsock channel to the parent EC2 instance, which relays encrypted data. 

+1 

Cryptographic Attestation and Model Integrity 

A main update in v3.4 is the enhanced attestation document. In generative AI, proving the correct model runs is as important as data protection. Nitro Enclaves v3.4 allows detailed measurement of model weights and logic with Platform Configuration Registers (PCRs). 

Through integration with AWS Key Management Service (KMS), a Nitro Enclave can cryptographically prove its identity and the integrity of its “Generative Logic” before any decryption keys are released. This means that an LLM’s weights, often a company’s most valuable intellectual property, remain encrypted in Amazon S3 and are decrypted only in the enclave’s volatile memory. If the enclave’s code or the model’s signature is altered by even a single bit, the attestation fails, and the “seal” prevents the logic from being executed. 

+1 

Confronting the Challenges of Generative AI at Scale 

Large-scale AI deployments face three primary security hurdles: prompt leakage, model weight theft, and exposure of inference telemetry. AWS Nitro Enclaves v3.4 addresses these with a layered isolation approach: applications in which the parent instance never sees the unencrypted prompt or the model’s response. This is essential for domains such as healthcare and finance, where PII (Personally Identifiable Information) must be processed by AI without being logged or stored. 

  • Persistent Model Protection: Since enclaves lack persistent storage, decrypted model weights exist only in memory. When the enclave shuts down, the Nitro Hypervisor securely erases the memory, so attackers cannot recover any data. 
  • Refined Resource Allocation: v3.4 improves the balance between memory and compute, so larger enclaves can use high-performance instances like the c7i and r7g series. This means the added security does not significantly slow down inference. 

Pragmatic Deployment Workflows 

To deploy Sealed Generative Logic Isolation, you follow a clear DevOps process. First, you create an Enclave Image File (EIF) with the inference engine, such as vLLM or llama.cpp, and the required startup code. 

In v3.4, nitro-cli supports larger images and complex dependencies, simplifying the containerization of multimodal models. After signing and deploying the EIF, the enclave retrieves the model’s decryption key from KMS using its unique identity, keeping models isolated from the parent OS kernel. 

The Shift Toward Sovereign and Compliant AI 

This release comes as global regulations move toward Sovereign AI. Governments and international bodies now require AI processing to remain within specific legal and security boundaries. Using AWS Nitro Enclaves v3.4, organizations can demonstrate a “Zero-Trust” setup for their AI workloads. 

Isolation matters for Multi-Party Collaboration. With v3.4, two organizations can share sensitive data in a single enclave, running specialized “LoRA” analysis or training models while keeping the raw data inaccessible to either party. The enclave acts as a secure neutral “clean room” for generative logic. 

Conclusion: The New Baseline for AI Trust 

As generative AI transitions from novelty to a core component of enterprise infrastructure, the underlying “plumbing” must be as robust as the models themselves. The introduction of Sealed Generative Logic Isolation in AWS Nitro Enclaves v3.4 provides the technical foundation for this trust. By hardware-sealing the inference process, AWS is removing the major barriers to AI adoption in highly regulated sectors. 

For organizations seeking to securely integrate LLMs into workflows, v3.4 sets a new standard. It protects model intelligence, even in cloud environments.

Source: What is Nitro Enclaves?

In high-performance computing, hardware access is only possible with strong software support. With the release of CUDA 13.1 in 2026, the industry received more than a routine update, and NVIDIA confirmed full specifications for its most powerful Blackwell accelerator: the B200 Ultra with 144 GB of HBM3e memory per GPU. Developers have the resources needed for new trillion-parameter models and advanced AI workflows.  

The Architecture of the B200 Ultra 

The B200 Ultra is the peak of Blackwell Architecture, addressing memory bottlenecks seen in Hopper-based systems (the architecture family that preceded Blackwell). While the standard B200 was a major step, the Ultra variant targets the AI Factory era, handling rapidly growing model weights and KV (key-value) caches in AI workloads.  

NVIDIA uses 12 high-bandwidth memory (HBM3e) stacks to reach 144 GB on the B200 Ultra, enabled by a dual-die setup with two chips linked via the NVIDIA Hi-Bandwidth Interface (NV-HBI), a specialized high-speed connection. The 10 TB/s (terabytes per second) connection joins the chips into a single accelerator, allowing CUDA 13.1 to use the full memory pool without added latency.  

CUDA 13.1: The Software Enabler for 144 GB 

CUDA 13.1 is the technical bridge enabling developers to harness this capacity. A key addition is CUDA Tile A, a programming model abstracting Blackwell’s hardware complexity.  

Previously, GPU programming required managing data at the thread and warp level with SIMT. As memory and hardware grow more complex, this becomes harder. CUDA tiles let developers break work into tiles or data chunks, which the computer and runtime assign to the B200 Ultra’s 144 GB of memory and its Tensor cores. The memory bandwidth, about 8 TB/s, is used efficiently, and programmers avoid manual detail management.  

Memory Locality and Green Contexts 

Among the most significant features in CUDA 13.1 are those that manage how data is moved, stored, updated, calculated, optimized, partitioned, and introduced.  

Green Contexts: Let system administrators set up and manage separate sections of GPU resources on a B200 Ultra with 144 GB of HBM3e. A single card can be split into several isolated environments, each with its own memory and streaming multiprocessors (SMs, the fundamental compute units of a GPU). This is especially useful for AI hubs and data centers that support multiple users, since one B200 Ultra can run a secure government language model in one section and a commercial AI agent in another, with hardware making sure there is no data leakage between them.  

Why 144 GB Matters: The LLM and MOE Challenge 

The jump to 144 GB of HBM3 meets the specific needs of a mixture of expert architectures and large language models with long context windows. By 2026, models like Gemini 2.0 and GPT-5 will require substantial memory for millions of tokens simultaneously.  

Previously, developers used model sharding, splitting models across GPUs, which caused delays due to interconnect latency. With 144 GB memory, the B200 Ultra fits larger models on a single chip or in fewer GPUs per cluster, cutting costs and boosting real-time speeds.  

The B200 Ultra supports FP4 (4-bit floating-point 12) precision in CUDA 13.1, effectively doubling memory usage. With FP4, developers fit models that required 288 GB into 144 GB with minimal accuracy loss, thanks to Blackfield Transformer Engine’s dynamic scaling.  

Interconnect And Rack-Scale Integration 

The B200 Ultra is powerful on its own, but it really shines when connected with NVLink 5 (NVIDIA’s latest high-speed interconnect technology for GPUs). CUDA 13.1 improves communication techniques for the NVLink switch system, enabling up to 576 GPUs to work together over a single fast network.  

A standard DGX B200 setup combines eight 144 GB units for more than 1.1 TB of HBM3e memory. This high memory density enables the development of models with trillions of parameters, something that was not possible just two years ago. CUDA 13.1 also boosts the performance of the NVIDIA Collective Communications Library (NCCL), so operators like AI Radios and all-gather runs run at the full 1.8 GB/s speed of NVLink 5.  

Developer Productivity and Future Proofing 

Finally, NVIDIA has used the CUDA 13.1 update to modernize the developer experience. The toolkit now includes a unified version for both Tegra embedded (NVIDIA’s platform for mobile and edge devices) and desktop- and data-center GPUs, reducing the overhead for developers building cross-platform AI applications. The addition of CU TilePython, a domain-specific language (DSL) for authoring tile-based kernels, enables data scientists to write high-performance GPU code directly in Python, avoiding the need for low-level C++ for many common optimization procedures.  

The focus on productivity ensures the transition to the Blackwell Ultra platform remains as seamless as possible. Companies that have invested in the CUDA ecosystem over the last decade will find that their current codebases are forward-compatible, gaining performance boosts simply by recompiling with the CUDA 13.1 toolkit to take advantage of the B200 Ultra’s new memory resource abstractions.  

Final Thoughts: The New Baseline for AI Infrastructure.  

With CUDA 13.1 confirming 144 GB HBM3e memory for the B200 Ultra, a new standard has emerged for enterprise and national AI infrastructure. Large memory and a software stack focused on abstraction, safety, and multi-tenant use secure NVIDIA’s place in AI.  

As the B200 Ultra ships to top cloud providers and research labs in early 2026, the question will move from “How much memory is available?” to “How can we use it best?” With 144 GB, working with AI researchers is no longer held back by hardware, but only by the challenges they decide to tackle.

Source: CUDA Toolkit 13.2 – Release Notes

An Operator is an AI agent called a Computer Using Agent (CUA) that completes tasks by controlling a computer via its screen, mouse, and keyboard, automating browser tasks for users.  

Below are some important details about the operator release:  

  • Availability: Currently, ChatGPT Pro is offered to subscribers for $200 a month.  
  • Functionality: The column operator uses GPT-4 OS vision to interact with computer interfaces.  
  • Future Scope: OpenAI plans to expand Operator to the Plus team and enterprise users and integrate it into ChatGPT.  

The current research preview focuses on browser-based actions, aiming to let AI use computers as a human would.  

The operator is a web-based agent that navigates the internet and completes tasks for users. It operates within its own browser environment, allowing it to view web pages and interact by tapping, clicking, and scrolling. Currently in a research preview phase, Operator has certain limitations that are expected to be addressed with further user feedback. As one of OpenAI’s first agents, Operator enables users to delegate tasks, which it then executes autonomously.  

An operator can manage repetitive browser tasks on behalf of users, such as filling out forms, ordering groceries, or generating memes. Because it interacts with websites and tools in the same way users do, Operator enhances the practicality of AI. It streamlines routine activities and creates new opportunities for businesses to engage customers.  

We are starting with a small roll-out for safety and manageability. Pro users in the US can access Operator at operator.chatgpt.com. This limited release helps us learn from users and improve Operator over time.  

How Operator Works 

The operator runs on a new model called Computer Using Agent (CUA). CUA commands GPT for those vision skills, using advanced reasoning and reinforcement learning. It’s trained to work with graphical user interfaces, such as buttons, menus, and text fields you see on your screen.  

The Operator sees what is on the screen by taking screenshots. It interacts with the browser using all mouse and keyboard actions. This means it works on the web without needing special API interfaces.  

If Operator runs into problems or makes a mistake, it uses its reasoning skills to try to fix things on its own. If it can’t resolve the issue, it gives control back to you, ensuring the experience remains smooth and coordinated.  

CUA is still new and has some limitations, but it has already set new records in important browser benchmarks. More details about our evaluations and the research behind Operator are on our blog post.  

How to Use 

To start, tell the Operator what to do, and it will handle the rest. You can take control of the browser at any time. The Operator asks you to step in for tasks that need a login, payment, or when a captcha appears.  

You can personalize Operator with custom instructions for all or specific sites. For example, you might set airline preferences on booking.com. The operator also lets you save points for quick access. This is useful for frequent tasks like restocking groceries on Instacart. Using multiple tabs, the Operator can handle several tasks at once by starting new conversations, like ordering a mug from Etsy while booking a campsite on Hipcamp.  

Ecosystem & Users 

Operator changes AI from a passive tool into an active helper in the digital world. It makes tasks easier for users and helps companies offer better experiences and improve conversion rates. We’re working with companies like DoorDash, Instacart, OpenTable, Priceline, StubHub, Thumbstack, and Uber, and others to ensure Operator meets real needs and follows industry standards. We also see many ways operators can make certain workflows easier to use and more effective, especially in the public sector. For example, we are partnering with the city of Stockton to help people enroll in city services and programs more easily.  

By initially introducing Operator to a select audience, OpenAI aims to learn and refine its capabilities through real-world feedback, while maintaining a focus on innovation, trust, and safety. This approach supports meaningful value delivery to users, creators, businesses, and public sector organizations.  

Safety and Privacy 

Ensuring the operator is safe to use remains our top priority. We have added three layers of safeguards to prevent abuse and keep users in control.  

Operator keeps users in control by prompting for input at key moments.  

  • Takeover Mode: When sensitive information like passwords or payment details must be entered, the Operator prompts you to take over. In this mode, the operator does not collect or record any input.  
  • User confirmations: before completing actions such as placing an order or sending an email. The operator requests your approval.  
  • Task Limitations: Operator declines certain sensitive tasks, such as banking transactions or job application decisions.  
  • Watch Mode: On sensitive sites, such as email and financial services, the Operator operates under close supervision. This lets you promptly identify and correct any issues.  

Data privacy and management within Operator is designed to be straightforward.  

  • Training Opt-Out: If you turn off “Improve the model for everyone” in your ChatGPT settings, your data in Operator will not be used to train our models.  
  • Transparent Data Management: You can delete all browsing data and log out of every site with one click. In Operator’s settings, you can also easily delete past conversations.  

We have added protections to stop websites from manipulating Operator with hidden prompts, malicious code, or phishing attempts.  

  • Cautious Navigation: The operator can detect and ignore prompting actions.  
  • Monitoring: A dedicated monitor detects suspicious behavior and can pause tasks if necessary.  
  • The detection pipeline uses both automated systems and human reviewers to spot new threats. We update safeguards quickly. Operator is built to refuse harmful requests and block disallowed content. Our moderation can warn users or revoke access if rules are broken. Extra review steps help catch misuse. We provide guidance on using Operator in line with policies.  

Even with safeguards, no system is perfect, and Operator is under research review. We will improve it with feedback and testing. To learn more, visit the Operator Research blogs’ safety section.  

Limitations 

The operator is currently in an early research phase, and while it’s already capable of handling a wide range of tasks, it’s still learning and evolving and may make mistakes. For instance, it currently struggles with complex interfaces, such as creating slide shows or managing calendars. Early user feedback will play a vital role in upgrading its accuracy, reliability, and safety, helping us make Operator better for everyone.  

What’s Next? 

Cua in the API: The model behind Operator, called Cua, will soon become available via the API, enabling developers to build their own CAD computer using agents.  

Enhanced capabilities will keep working to help the Operator handle longer, more detailed workflows.   

Access: We plan to expand Operator to the plus team and enterprise users, and to integrate its capabilities directly into ChatGPT in the future, once we are certain of its safety and usability at scale, unlocking seamless, real-time, and asynchronous task execution.

Source:Introducing Operator 

Right now, we are moving from models that excel at specific tasks to agents that can handle more complex workflows. When you prompt a model, you only get its trained intentions, but if you give it a computer environment, it can do much more, like run services, request data from APIs, and/or create useful things like spreadsheets and reports.  

When building agents, some practical problems come up. For example:  

  • You need to decide where to store intermediate files.  
  • Avoid pasting large tables into prompts.  
  • Give workflows network access without causing security issues.  
  • Handle timeouts and read-rides without building your own workflow system.  

To address these agent-specific challenges, we built the components needed to give the Responsys API a computer environment. By doing this, we enable reliable management of real-world tasks, freeing developers from having to create their own execution setups. This sets the stage for tackling the broader practical problems faced in agent development.  

OpenAI’s API shell tool and hosted container workspace address these challenges. The model suggests steps and commands that run in a separate environment with its own filesystem, optional storage (e.g., SQLite), and limited Network Access.  

With this foundation in place, let’s explore how we build a computer environment for agents and discuss early lessons from using it to accelerate, standardize, and improve safety in production workflows.  

The Shell Tool 

A good agent workload needs a tight execution loop:  

  1. The model suggests an action.  
  1. The platform executes it.  
  1. The result informs the next step.  

We’ll start with the shell tool to illustrate this loop, then discuss the container, workspace, networking, reusable skills, and context. Compact Shenoy  

To understand the shared tool, know how a model uses tools. It suggests tool calls after seeing step-by-step examples during training. The model proposes tool use but can’t execute the calls itself.  

The shell tool gives the model command-line access to perform tasks like text search or API requests using familiar Unix utilities such as grep, curl, and awk.  

Unlike our current code interpreter, which runs only Python, the Shell Tool supports a much broader range of use cases. You can run GO or JAVA programs or start a Node.js server. Such flexibility enables the model to handle more complex tasks.  

Orchestrating The Agent Loop 

On its own, a model can only propose shell commands, but how are these commands executed? We need an orchestrator to retrieve model output, invoke tools, and return the tools’ response to the model in a loop until the task is complete.  

The Responsys API is how developers interact with OpenAI models when used with custom tools. The Responsys API returns control to the client, who must provide their own harness to run the tools. However, this API can also orchestrate between the modern and hosted tools out of the box.  

When the Responsys API receives a prompt, it assembles model context: user prompt, prior dialog state, and tool instructions. For shell execution to work, the prompt must mention using the Shell Tool, and the selected model must be trained to propose shell commands. Models GPT-5.2 and later are trained to do so with all of these contexts.  

The model then decides the next action. If it chooses shell execution, it returns one or more shell commands to the Responsys API service. The API service forwards those commands to the container runtime, streams the shell output back, and feeds it to the model in the next request’s context. The model can inspect the results, issue follow-up commands, or produce a final answer. The Responsys API repeats this loop until the model returns a completion without additional shell commands.  

When the Responsys API runs a Shell command, it keeps a streaming connection to the Container Service open. As output appears, the API sends it to the model almost immediately. This lets the model decide whether to wait for more output, run another command, or issue a final response.  

The model can suggest several shell commands at once. The Responsys API can run these commands concurrently in separate container sessions. Each session streams its output separately. The API then combines these streams into structured tool outputs for context. This allows the agent loop to run tasks such as searching files, fetching data, and checking results in parallel.  

Commands that handle files or process data may generate lots of shell output. This can fill a context space without adding much value. To manage this, the model sets an output limit for each command. The Responsys API enforces the limit and returns a result that keeps both the start and end of the output, marking skipped content. For example, you might set an AF1000 character limit, keeping the beginning and end.  

By combining concurrent execution and output limits, the agent loop maintains speed and context efficiency. The agent loop controls which tool outputs are included in the context, helping the model focus on important results rather than being overwhelmed by raw terminal logs.  

When The Context Window Gets Full: Compaction 

A challenge with agent loops is that some tasks run for a long time. These long tasks can fill up the context window, which tracks information across turns and agents. For example, an agent might call a skill, get a response, and then make turn calls and summaries. The Limited Context Window can fill up quickly, keeping important details while removing extraneous information. We built native compaction into the Responsys API. Developers don’t need to create their own summarization or state systems, and the feature matches model training.  

Our latest models are trained to review prior dialog states and generate a compaction item that stores key information in an encrypted, token-efficient format. After compaction, the context window includes this compaction item and the most important parts of the earlier window. This makes workflow progress smooth across window boundaries, even in long, multi-step, or tool-driven sessions. Codex uses this system to handle long programming tasks and repeated tool use without losing quality.  

You can use compaction either as a built-in server feature or through a separate /compact endpoint. With server-side compaction, you can set a threshold, and the system takes care of compaction timing for you, so you don’t need complex client logic. This setup allows a slightly larger input context window, so small overages just before compaction are handled rather than rejected. As models improve, the native compaction feature updates with every OpenAI model release.  

Codex played a key role in building the compaction system. It was one of the first to use it. If one Codex instance hit a compaction error, we started another instance to investigate. This process helped Codex develop a strong built-in compaction system by working through the problem. Codex’s ability to examine and improve itself has become unique to OpenAI. While most tools just need users to learn them, Codex learns with us.  

Container Context 

Now let’s talk about State and Resources. The container is more than merely a place to run commands. It’s also the model’s working environment. Inside the container, the model can read files, query databases, and reach external systems, all under network policy controls.  

File Systems 

The first part of the container context is the file system, which is used to upload, organize, and manage resources. We created container and file APIs to give the model a clear view of available data and help it select specific file operations rather than run broad energy scans.  

All inputs are directly into the prompt context. As inputs grow, filling the prompt becomes more expensive and harder for the model to navigate. A better approach is to stage resources in the container’s file system and let the model decide which to open, pass, or run via shell commands, much as humans do. Models work better with organized information.  

Databases 

The second part of the container context is databases. We recommend storing structured databases in SQLite and varying them directly rather than copying a spreadsheet into the prompt. Describe the tables and columns, and explain their meanings so the model can pull only the needed rows.  

For example, if you ask which products had declining sales this quarter, the model can look up only the relevant rows rather than search the entire spreadsheet. This approach is faster, cheaper, and better suited to large data sets.  

Network Access 

The third part of the container context is Network Access, essential for agent workloads. Agents may need to fetch live data, call external APIs, or install packages. Giving containers full internet access can be risky, as it allows them to store information outside sites, access sensitive systems, or make it harder to prevent leaks.  

To solve these problems without limiting what agents can do, we set up hosted containers to use a central egress policy proxy. All ongoing network requests go through a central policy layer that enforces allow lists and access controls and keeps traffic visible. For credentials, we use Domain Scoped Secret Injection at egress. The model and container only see placeholders, while the real secret values remain hidden and are used only for approved destinations. This reduces the risk of leaks while still allowing secure external calls.  

Agent Skills 

Shell commands are powerful, but many tasks follow similar multi-step patterns. Agents often must replan and relearn, leading to inconsistent results. Agent skills package these patterns into reusable building blocks. A Skill Easy Folder with a Skill.MD File and Resources such as API Specs and UI Assets.  

This structure maps naturally to the runtime architecture we described earlier. The container provides persistent files and an execution context, and the shell tool provides the execution interface. With both in place, the model can discover scaled files using shell commands (ease, cat, etc.) when needed, interpret instructions, and run scaled scripts within the same agent loop.  

We provide an API to manage skills on the OpenAI platform. Developers upload and store skill folders as versioned bundles, which can later be retrieved by skill ID before sending the prompt to the model. The Responsys API loads the skill and includes it in the model context. The sequence is deterministic.  

  1. Fetch skill metadata, including name and description.  
  1. Fetch the scale bundle, copy it into the container, and unpack it.  
  1. Update model context with skill metadata and the container path.  

When deciding if an SQL is relevant, the model reviews its instructions step by step and runs its scripts using shell commands in the container.  

How Agents are Made 

To put all the pieces together: the Responsys API handles orchestration, provides shell tools, runs actions, supports a double-step container, provides an open system runtime context, supports skill add, provides reusable workflow logic, and enables compaction to let an agent run for a long time with the context needed for an end-to-end workflow.  

Discover the right scale, fetch data, and transform it into a local structured state. Query it efficiently and generate durable artifacts.  

Make Your Own Agent 

For a step-by-step example using the shell tool and computer environment, see our developer blog post and cookbook. These resources show how to package and run a SQL with the responses API.  

We’re eager to see what developers build. Language models go beyond creating text, images, and audio. We’ll continue to enhance our platform for complex real-world tasks at scale.

Source: From model to agent: Equipping the Responses API with a computer environment