Artificial Intelligence & Machine Learning

NVIDIA Drives Next Wave of AI-Powered Creativity and Physical Intelligence at SIGGRAPH 2026

The SIGGRAPH conference, a premier event for computer graphics and interactive techniques, is currently underway in Los Angeles through Thursday, July 23. This year’s iteration is highlighting how advancements in graphics research, neural rendering, simulation, and artificial intelligence are fundamentally reshaping the creation and comprehension of digital and physical worlds for both humans and machines. A pivotal moment at the conference was the NVIDIA keynote address on July 20, where AI research and engineering leaders Neil Ashton, Edward Liu, and Ming-Yu Liu unveiled a suite of innovations poised to democratize advanced creative tools and push the boundaries of physical AI.

The keynote focused on several key areas: the expansion of AI agents into mainstream creative applications, the development of more sophisticated world models for AI, and the acceleration of real-time simulation. These developments signal a significant shift from merely accelerating existing workflows to creating entirely new paradigms for how content is produced and how intelligent systems interact with the physical environment.

AI Agents Expand Creative Tools to Millions

A central theme at SIGGRAPH 2026 is the pervasive integration of AI agents into professional creative software, empowered by the Model Context Protocol (MCP). This initiative aims to bring advanced AI capabilities directly into the hands of millions of creators by allowing AI agents to operate within established workflows for scene creation, shot composition, timeline editing, asset management, and more. Critically, these integrations are designed to keep creators firmly in control of the creative process.

For over two decades, NVIDIA technologies have been instrumental in accelerating the Digital Content Creation (DCC) pipeline. From GPU-accelerated viewports and CUDA-powered effects to NVIDIA RTX PRO ray tracing, AI denoising, neural rendering, and real-time simulation, these innovations have consistently enhanced the tools used by artists, studios, and developers to build the world’s games, films, television shows, and advertising. MCP represents the next evolutionary step, transforming applications not just to be faster, but "agent-ready."

From Acceleration to Action:

The introduction of MCP-connected tools signifies a move from passive acceleration to active assistance. Artists and technical directors can now leverage AI agents to perform tasks such as inspecting scenes for missing textures, ensuring color management consistency, generating various export formats, creating playblasts for daily reviews, or validating shots against pipeline rules. These operations can be executed while the creative decision-making remains with the human user.

The same NVIDIA platform that has powered accelerated viewports, rendering, simulation, and AI effects is now being extended to support local agents, model inference, and multi-application workflows on systems specifically designed for professional creators. NVIDIA RTX PRO workstations, alongside DGX Spark and DGX Station systems, are being positioned to bring accelerated AI performance directly to the desks of artists and developers, streamlining studio pipelines. Running models and agents locally offers significant advantages, including improved responsiveness, reduced reliance on external cloud services, and enhanced security for sensitive creative data within controlled environments.

The NVIDIA Agent Toolkit is also being updated to support MCP integration, featuring an MCP client for connecting to remote servers and an MCP server for publishing tools to any MCP client. This interoperability is crucial for building a cohesive and scalable agent-driven ecosystem.

The Creative Ecosystem Goes Agent-Ready:

Across the creative industry, leading applications and platforms are embracing MCP, providing AI agents with more contextually relevant access to production environments.

At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
  • Adobe is enhancing its creative agent across Firefly, Express, and Creative Cloud, powering AI Assistant experiences. These assistants enable creators to describe desired outcomes, with the AI orchestrating multi-step workflows. Adobe is also extending its professional creative tools to third-party AI platforms via the Adobe Connector, broadening accessibility. For developers, the Adobe Express Developer MCP Server facilitates the creation of Adobe Express add-ons using official documentation and APIs.

  • Affinity by Canva has introduced an AI Connector for Claude, leveraging MCP to integrate natural-language automation directly into Affinity applications. Designers can now instruct Claude to handle repetitive production tasks like renaming layers, resizing assets for multiple channels, applying bulk edits, optimizing vector paths, and preparing files for delivery. Claude can also assist in building reusable scripts and custom features, reducing production overhead and freeing up creative professionals to focus on design.

  • Blender, through Blender Lab, offers a lightweight MCP server, providing a natural-language interface to its extensive Python API, documentation, and complex setups. This initiative demonstrates how open-source creative tools can become accessible to AI agents without compromising their core functionality.

  • Boris FX Silhouette has integrated an MCP server, allowing AI assistants to work directly within projects. By utilizing Silhouette’s FX Scripting API as first-class MCP tools, assistants can inspect projects, construct node trees, edit shapes and keyframes, and render frames. A new preferences panel simplifies setup, including MCP package installation, client configuration generation, and connection testing. Both interactive online and headless offline modes are supported for automation and batch processing.

  • Foundry Griptape natively supports MCP, offering AI orchestration specifically tailored for professional VFX pipelines. This integration allows studios to securely manage multiple AI models and agents while maintaining traceability and creative control. Griptape automates repetitive tasks like cleanup, matte painting, and quality control in conjunction with tools like Blender and Foundry Nuke, ensuring artists remain in command.

  • SideFX is incorporating MCP support into Houdini 22 via its new APEX Script workflow. AI assistants will have access to a curated set of APEX Script syntax, functions, documentation, and examples, aiding artists in generating and refining code for procedural character rigs. While the initial focus is on APEX Script and character rigging, community-developed MCP servers promise broader agent interaction with Houdini.

  • Unreal Engine has announced MCP connectivity for Unreal Editor, enabling AI workflows that can interact with editor capabilities through a standardized protocol. This development opens new possibilities for game developers, virtual production teams, and real-time artists, allowing AI assistants to analyze scenes, assets, and project states.

These advancements underscore a commitment to making powerful AI tools more accessible and integrated into the daily workflows of creative professionals, accelerating production cycles and unlocking new levels of artistic expression.

NVIDIA AI for Media Helps Newsrooms Detect Synthetic Video

In an era where video is the primary medium for understanding global events, maintaining public trust in visual media is paramount. NVIDIA announced at SIGGRAPH 2026 the Synthetic Video Detector NVIDIA NIM microservice, a crucial addition to the NVIDIA AI for Media platform. This tool aims to integrate an AI-assisted detection signal directly into editorial and media workflows, helping news organizations combat the growing threat of synthetic or manipulated video content.

The NIM microservice operates by analyzing video frame by frame, producing a classifier score indicating the likelihood of synthetic content. Editorial teams can utilize this score to efficiently prioritize clips for review, flag potentially problematic footage, or escalate it for further in-depth analysis. Rather than replacing established verification practices, the microservice acts as an additional layer of intelligence for time-sensitive decision-making, enabling newsrooms to maintain high editorial standards and preserve public trust.

A critical aspect of the Synthetic Video Detector’s effectiveness is its resilience to common post-production processes. NVIDIA testing has shown that the model maintains significant accuracy even after video compression, resizing, cropping, and re-encoding – steps frequently applied in newsroom and social media workflows. The detector achieved up to 92% accuracy on uncompressed video, with performance dropping to 87% at 15% compression and 82% at 50% compression. Furthermore, the NIM microservice demonstrates impressive speed, processing 1080p video in as little as 22 milliseconds on NVIDIA RTX systems and approximately 30 milliseconds on NVIDIA L40 GPUs, making it suitable for real-time applications.

At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI

Deploy Detection Where Video Lives:

The flexibility of the NIM microservice allows for deployment closer to the source of video content, whether that be on-premises, at the edge, in hybrid environments, or within approved air-gapped systems. This capability empowers organizations to maintain stringent control over their video data, access, and operations.

Industry adoption is already underway, transforming the microservice from a research capability into deployable media infrastructure. Wowza is integrating the microservice into its Wowza Video Intelligence Framework, bringing real-time synthetic video detection to livestreaming workflows used in over 35,000 deployments globally. This integration is particularly significant for organizations facing high risks from synthetic media, such as broadcasters, government agencies, financial institutions, and critical infrastructure operators, many of whom operate under strict data residency and security requirements. By embedding AI-assisted verification within existing video infrastructure, Wowza enables real-time flagging of questionable video while keeping sensitive footage within secure environments.

This advancement is a critical step in safeguarding the integrity of visual information in the digital age, providing news organizations with a powerful tool to discern authentic content from increasingly sophisticated synthetic media.

Now Openly Available, NVIDIA Cosmos 3 Edge Brings Frontier World Models to Edge GPUs for Local Physical AI

The development of intelligent systems that can perceive, reason about, and predict their physical environment hinges on sophisticated "world models." However, the inherent vastness, unpredictability, and constant flux of the real world present significant challenges for these systems. Whether deployed in robots navigating complex environments or in camera networks monitoring industrial settings, physical AI systems require the ability to understand the present, anticipate the future, and act with sufficient speed to be effective. Historically, delivering such advanced AI capabilities at the edge has often necessitated a compromise between model sophistication and deployment efficiency.

NVIDIA Cosmos 3 Edge aims to eliminate this tradeoff. This 4-billion-parameter omnimodel has been meticulously optimized for memory-efficient deployment and high throughput on a range of NVIDIA hardware, including NVIDIA Jetson, NVIDIA RTX PRO, and NVIDIA DGX systems, as well as GeForce RTX GPUs. Building upon the foundation of NVIDIA Cosmos 3, this compact world foundation model possesses the ability to understand and generate text, images, video, ambient sound, and actions. Its innovative mixture-of-transformers architecture facilitates physically grounded, real-time vision analytics and on-device robot action.

Cosmos 3 Edge represents a significant leap forward in delivering frontier physical AI at the edge. It has achieved the No. 1 ranking on VANTAGE-Bench for vision analytics success within its parameter class and enables state-of-the-art robot learning through post-training capabilities.

On-Device Physical AI Across Robotics, Autonomous Vehicles and Smart Infrastructure:

Developers can leverage NVIDIA DGX Station, a deskside AI supercomputer, to post-train Cosmos 3 Edge on proprietary robot and sensor data. This process allows for the creation of specialized world action models that can then be deployed onto NVIDIA Jetson Thor for real-time robot control policies, including complex manipulation and locomotion tasks. Several leading companies, including Agile Robots, Doosan Robotics, Siemens, and Skild AI, are actively evaluating Cosmos 3 Edge for their robotics workflows.

For autonomous vehicles, Cosmos 3 Edge supports crucial functionalities such as road-scene understanding, traffic reasoning, object-intent prediction, and policy-model distillation, all within resource-constrained hardware environments. The model can serve as a student backbone for automotive policy model distillation, including integration with NVIDIA Alpamayo vision-language action models.

In the realm of smart infrastructure, Cosmos 3 Edge delivers best-in-class throughput and accuracy for real-time inference on Jetson Thor. This enables vision agents to analyze live video streams for applications such as traffic monitoring, public safety, logistics, and industrial inspection. Developers also have the option to run the 2-billion-parameter NVIDIA Nemotron-powered reasoning module independently on NVIDIA Jetson Orin 8GB. Companies like Centific, Vaidio, and YUAN are currently evaluating Cosmos 3 Edge to accelerate vision agents operating at the edge.

At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI

Cosmos Platform Now Openly Available:

Cosmos 3 Edge is an integral part of the broader NVIDIA Cosmos platform, designed for the development of physical AI world models. The platform offers Cosmos 3 in three sizes: Edge (4B parameters), Nano (16B parameters), and Super (64B parameters). This tiered approach allows developers to select the most appropriate model for each stage of development, from edge deployment to high-fidelity generation. Cosmos 3 Edge, Cosmos 3 Nano, and Cosmos 3 Super are now available on Hugging Face, with inference and post-training frameworks and recipes accessible on GitHub.

This initiative democratizes access to cutting-edge world models, empowering a new generation of intelligent systems capable of more sophisticated and nuanced interaction with the physical world.

AI Agents Made Easy: Build and Run Personal AI Agents Locally on DGX Station With NVIDIA Agent Toolkit

The advent of powerful, personalized AI agents is now within reach for individuals and teams, thanks to the NVIDIA Agent Toolkit and the NVIDIA DGX Station. The DGX Station, positioned as the ultimate deskside AI supercomputer, now offers a streamlined path to setting up and running advanced AI agents. The process has been simplified to just three steps, with systems capable of being operational in approximately 30 minutes.

On the DGX Station, the NVIDIA Agent Toolkit consolidates NVIDIA NemoClaw, the NVIDIA Nemotron 3 Ultra open model, NVIDIA Omniverse libraries as agent-accessible tools and skills, and a secure runtime environment into a single, local system that requires no internet connection. As workloads scale, developers can connect multiple DGX Stations to support concurrent users, a greater number of agents, and larger models. This capability empowers creatives and engineers to "own their own intelligence," with a system pre-configured for local operation. The complete stack—model, agent, and tools—provides a robust platform for creating and deploying domain-specific "super agents" customized with users’ own data and knowledge.

The open NVIDIA Agent Toolkit stack on DGX Station includes:

  • NVIDIA NemoClaw: A framework for building AI agents.
  • NVIDIA Nemotron 3 Ultra: An open foundation model for enhanced agent capabilities.
  • NVIDIA Omniverse Libraries: Agent-accessible tools for simulation and world building.
  • Secure Runtime: A local environment for agent execution.

Harness Efficiency at Scale:

For teams deploying agents at scale, the economics fundamentally shift with DGX Station. Nemotron 3 Ultra, tuned for an open harness, delivers leading-edge performance without the per-token costs associated with cloud-based services after the initial hardware investment. This allows users to build and run agents as much as needed.

NVIDIA has also introduced a blueprint for integrating NVIDIA Omniverse libraries into Blender. This integration equips NemoClaw agents with callable RTX sensor simulation and physics tools, essential for preparing 3D scenes for physical AI workflows. On the DGX Station, designers and engineers can run all core components of this workflow—the frontier model, open harness, secure runtime, and 3D tools—in a single, connected system.

Frontier models can orchestrate NemoClaw as a specialized sub-agent, delegating domain-specific tasks to agents running locally on DGX Station. These agents can directly access Omniverse tools and Blender, enabling complex workflows. LangChain has optimized its Deep Agents harness for Nemotron 3 Ultra, providing designers and engineers with a production-ready pathway to achieve benchmark-leading agentic performance at a significantly reduced cost.

Nous Research has fine-tuned Nemotron 3 Ultra for its Hermes Agent harness and adopted it for production workloads, demonstrating the value of owning AI intelligence. Tuning the model for a developer’s specific stack results in agents that are both faster and more capable within particular domains. The Hermes Agent has also expanded its Model Context Protocol catalog to include Blender, allowing teams to activate Blender directly from their agent. This represents a live example of a tool-using NemoClaw agent operational on DGX Station. For teams utilizing OpenClaw, this stack extends capabilities by bringing Nemotron 3 Ultra, Omniverse tools, and local inference on DGX Station into an environment where OpenClaw’s persistent, long-running agents can act continuously.

At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI

Develop and Deploy Quickly With New Playbooks:

Two new playbooks are now available to assist developers in building and running agents out-of-the-box with NemoClaw and dual-node deployments. NVIDIA DGX Stations are available for order from ASUS, Dell Technologies, Exxact, GIGABYTE, HP, MSI, and Supermicro, making this powerful AI development platform accessible to a wide range of organizations.

NVIDIA Brings Graphics Research Breakthroughs to Simulation and Physical AI

At SIGGRAPH 2026, NVIDIA’s research is pushing beyond the creation of visually realistic worlds to focus on worlds that behave realistically and respond in real time. This paradigm shift is evident across NVIDIA’s 21 accepted technical papers, which form the foundation for real-time systems capable of generating virtual worlds and driving machine training in the physical world. Whether the end product is a game, film, robot, or digital twin of a factory, the overarching goal is to expand the creative canvas with AI-generated worlds that are grounded in 3D, governed by physics, and directed by creators.

A prime example of this research is MotionBricks, a real-time motion model trained on over 350,000 motion clips and running at game-engine speeds. MotionBricks allows creators to direct and connect character movements. Remarkably, the same model that drives animated characters on screen can also control a Unitree G1 humanoid robot in a physical setting, utilizing computer graphics and simulation to accelerate physical AI development.

The framework for training generative controllers on large-scale motion datasets, GPC (Generative Pose Controllers), extends this concept. NVIDIA pre-trains a single controller on extensive human motion data, imbuing it with transferable motor skills that can be applied to new tasks. This represents a significant step towards a foundational model for motor control.

To facilitate the creation of virtual worlds for testing these movements, ArtiFixer transforms messy real-world 3D captures into clean, complete virtual scenes. It also introduces a novel method for predicting photorealistic global illumination directly from a scene’s geometry, eliminating the need for ray tracing.

To ensure these virtual worlds behave according to real-world physics, a new solver has been developed for the NVIDIA Newton physics engine, bringing difficult-to-simulate materials like snow, sand, and elastic solids to life. Furthermore, to maintain creator control, the VideoNeuMat pipeline provides reusable, relightable materials extracted from generative video models. The ARDY autoregressive diffusion model allows creators to steer 3D character motion in real time using text prompts.

These research breakthroughs, made openly available with accompanying code and models for download, highlight NVIDIA’s commitment to advancing the fields of computer graphics, simulation, and physical AI, enabling creators and developers to build more immersive, interactive, and intelligent experiences.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button