NVIDIA Expands Agent Safety: Runtime Permissions Do Not Guarantee Correct Work
TL;DR
NVIDIA launches its Open Agent Safety Platform, with OpenShell 0.1.0 enforcing file, network and tool permissions outside the model. That can contain mistakes, but more than 100 participating organizations do not establish deployment results or correct work.
NVIDIA launched its Open Agent Safety Platform on 2026-09-28, placing enforcement of agent permissions outside the model. Whether this supports longer autonomous work can be tested through attempted violations: when an agent switches tools or starts a subprocess, are prohibited operations still blocked? A promise in a conversation to follow the rules cannot answer that question.Announcement Technical explanation
The event and this article both fall on September 28 in Taipei. This Reuters edition appeared at 05:01 EDT, or 17:01 in Taipei, before the reporting cutoff of 23:56. The official announcement gives only a date; the report’s timestamp is not the software’s activation time. This is a same-day announcement, not a previous-day fallback.Publication time
OpenShell 0.1.0 is available as open source, with external controls over files, networks, processes and tool calls. An application-layer proxy can inspect requested operations and assign different permissions to reading and writing within the same service. Credentials can remain outside the agent workload. That offers a way to block the operation itself, without depending on the model to keep obeying its instructions after reading malicious material.Runtime mechanism Source code and architecture
The platform also includes the NVIDIA Sentry reference design. Infrastructure such as BlueField-4 separates continuous monitoring from the resources running the agent. This is an architecture for vendors to integrate, not a product already installed at every customer, nor proof that any agent will operate without mistakes.Monitoring architecture
NVIDIA says more than 100 organizations are working with its platform technologies. It has also integrated OpenShell with Slack so teams can inspect activity and approve or reject requests for additional permissions. The organization count does not establish the number of production deployments. Reuters reports NVIDIA’s claim that these tools could have prevented the earlier Hugging Face incident. That is a vendor’s counterfactual assessment, not an independently reproduced test.Participation and integration Limits of the claim
HP announced the same day that it will build AI infrastructure components using the platform’s technologies. This is a future product plan, without delivery dates, prices or comparative customer outcomes in the announcement. It should be distinguished from the existing Slack integration; a partner list does not mean every feature is generally available.HP’s plans
Let agents propose changes without handing over authorization
I favor separating operational permissions from the agent. Suppose a patching agent encounters malicious text asking it to upload credentials. If an external rule prohibits that destination, transmission should remain blocked even if the model accepts the instruction. This benefit does not require proving that the model always recognizes attacks. It does require requests to pass through the enforcement layer and the agent to remain unable to change its own rules.
That arrangement can also change how a product offers autonomous work. A team could allow continuous modifications on a reversible branch while retaining release authority with a person or another system. The useful commitment to users is which actions can proceed independently and which must stop, rather than a general promise about how long an agent can keep working. This is a product choice inferred from the control mechanism, not a feature already present in every integration.
An agent with legitimate permissions can still produce the wrong patch. Suppose it edits a permitted file but breaks a feature. The permission system can work perfectly while the result remains unacceptable. Authorizing an operation and accepting an outcome therefore need different evidence. The absence of a permission violation is no reason to omit tests and review.
Overly broad permissions reduce protection; overly narrow ones can repeatedly interrupt legitimate work. I would make the reason for an interruption visible to users rather than let an agent expand its own authority to finish. Useful subsequent measurements are the proportion of legitimate tasks interrupted by permissions and the proportion of delivered results that pass acceptance. This release supplies enforcement tools, but not independent data sufficient to answer both questions.
The cover reuses an archival photograph of NVIDIA premises from this site, not an image of the new platform. Downloading the event image failed because DNS resolution was unavailable.
Sources:
Related Articles
OpenAI Launches Dots: Ongoing Work Should Not Become a Growing Review Queue
OpenAI is rolling out dots to eligible paid users for work that continues between conversations. Read-only background research and limited memory controls make missed commitments and accumulated review work more useful tests than activity alone.
Copilot Unites Office and Persistent Agents: Can Work Continue Without Reconciliation?
Microsoft announces Home, Code and Autopilot, connecting shared Office files with Copilot. Integration could reduce version reconciliation, while persistent agents must recognize obsolete goals. Access remains phased rather than universal.