THE APEX TIMES
NVIDIA Research unveils foundation models aimed at general-purpose grasping, faster driving reasoning, and agent training in virtual worlds
At CVPR, NVIDIA highlighted three research papers that focus on a common idea: train at scale so models can handle new grippers, new driving situations, and new environments with less retraining.
NVIDIA Research used this year’s CVPR conference to push a message it says is central to the next wave of “physical AI”: training at scale should produce systems that generalize instead of behaving like narrow specialists. In a June 3 research update, NVIDIA described three papers and related efforts that target three bottlenecks it sees across robotics and autonomous systems: robot grasping that often requires per-device retraining, autonomous driving reasoning that must run quickly on in-car hardware, and virtual-agent training that depends on exposing models to enough variety before deployment. The company also pointed to conference recognition for two of the workstreams, saying NitroGen and another NVIDIA-authored paper were best-paper finalists at CVPR.
For robotic grasping, NVIDIA said most current AI approaches are tailored to a specific end effector, or gripper. In practice, a vision-language-action policy trained for a two-finger gripper typically only learns to operate with those two fingers, and dexterous grasping policies tend to assume a particular multi-finger embodiment. That specialization forces robotics teams into repeated cycles of new training data, fine-tuning, and validation when they change grippers, NVIDIA said.
To address that “embodiment lock-in,” NVIDIA presented GraspGen-X as what it calls the first foundation model for zero-shot grasping. A foundation model is a general-purpose model trained on large datasets that can often be adapted to new tasks with minimal or no retraining. NVIDIA said GraspGen-X uses geometry and contact understanding to generate grasp pose proposals for a new gripper it has never seen before and an object it has not encountered. The company said the training data was built synthetically because a comparable dataset is “impossible” to collect in the real world at scale, and it reported generating 2 billion simulated grasps across thousands of object shapes and synthetic gripper configurations.
NVIDIA said GraspGen-X can reduce per-gripper training cycles for developers, and that it can be paired out of the box with curoboV2, described as a CUDA-accelerated motion planning library, to turn grasp proposals into motions in unknown environments. NVIDIA also connected this grasp-generation work to a next-stage closed-loop approach, citing another paper, Grasp-MPC, which it said was presented at ICRA 2026. Closed-loop grasp execution means the system uses ongoing feedback during the grasp process, rather than treating grasping as a purely open-loop prediction.
On the autonomous vehicle side, NVIDIA argued that reasoning systems must be fast enough to run on the compute inside a vehicle, not just accurate in the abstract. The company said text-based “chain-of-thought” reasoning generates words token by token, and that token count becomes a practical constraint on response time. NVIDIA’s LCDrive paper, it said, tackles this by replacing words with compressed latent representations, so the model “thinks” in a compact internal space rather than producing human-readable reasoning steps.
NVIDIA said LCDrive alternates between proposing candidate actions and predicting what the world will look like if those actions are taken. It then uses the predicted world state to refine its next step, essentially keeping the same reasoning loop but in a computationally more efficient form than natural-language reasoning. The company reported that LCDrive achieved output trajectory quality comparable to text-based reasoning while using roughly half the tokens. NVIDIA also said LCDrive was built on its Alpamayo model and trained with supervision derived from existing vehicle data.
The third workstream focused on embodied agents trained in virtual worlds. NVIDIA described Isaac GR00T as an open foundation model for humanoid robots, built around the idea that exposing a model to enough diverse situations can improve its ability to generalize to unseen ones. NitroGen, NVIDIA said, extends that principle to virtual environments by using the GR00T architecture to train a generalized gameplay AI foundation model for embodied agents. The company argued that video games supply structured, varied worlds with explicit goals and success conditions, making them well-suited to large-scale training.
NVIDIA said NitroGen was trained across more than 1,000 games and 40,000 hours of interaction, and it reported evaluating the model across categories that include action role-playing games, platformers, roguelikes, and open-world games. The company said the resulting agents learned to demonstrate gameplay behaviors spanning combat, navigation, and exploration. For data efficiency, NVIDIA reported that in low-data conditions, starting with NitroGen improved performance by up to 52% versus previous state of the art. It also said the model is open source and is available on GitHub and Hugging Face.
Taken together, the CVPR package reflects NVIDIA’s broader push to make “physical AI” less dependent on bespoke data collection and narrow training setups. If these approaches hold up beyond demonstrations and paper benchmarks, they could reduce integration friction for robotics and driving teams by offering models that transfer across grippers, reason faster inside constrained compute budgets, and bootstrap agent training with large-scale simulation. Still, NVIDIA’s blog post does not provide details on the hardware configuration used for the token-rate claims, the specific evaluation environments and metrics for every result, or the practical availability timeline for developers seeking to adopt the systems. It also does not clarify what level of fine-tuning, safety testing, or real-world validation will be required for each use case.
Why It Matters
- Reducing per-embodiment training could lower the cost and time it takes to deploy robotic grasping systems across different grippers and robot platforms.
- Token-efficient reasoning may make agentic driving approaches more feasible on embedded in-car compute, where response latency matters.
- Large-scale training in simulation, including video-game worlds, could accelerate development and evaluation of embodied agents without requiring constant real-world data collection.
- By packaging research around reusable foundation models and developer tooling, NVIDIA is indicating an ecosystem strategy that goes beyond GPUs into end-to-end AI development stacks.
Sources
Key Facts
- NVIDIA said three CVPR papers share a theme of training at scale to improve generalization across diverse applications.
- GraspGen-X is presented as a zero-shot grasping foundation model designed to work with new grippers and unknown objects without per-gripper retraining.
- NVIDIA reported training GraspGen-X on 2 billion simulated grasps spanning thousands of object shapes and synthetic gripper configurations.
- LCDrive replaces text-based chain-of-thought with compressed latent reasoning and, NVIDIA said, achieved similar trajectory quality using roughly half the tokens.
- NitroGen is described as a generalized gameplay foundation model trained across more than 1,000 games and 40,000 hours, with open-source availability on GitHub and Hugging Face.
- NVIDIA reported NitroGen and PixelDiT were named best-paper finalists at CVPR, given to 15 of over 4,000 accepted papers.
Technology Related
Anthropic agrees to a $35 billion cloud computing deal tied to Nvidia-backed Lambda, report says
Anthropic PBC is reportedly moving to lock in large-scale compute capacity through a major multi-year arrangement with Lambda, a cloud provider backed by Nvidia. Terms and timelines were not fully disclosed in the report.
AMD has tended to fall in September, but market history is only part of the story
A review of the past decade points to a recurring pattern for AMD in September. The stock has declined in eight of the last 10 Septembers, though broader market seasonality appears to explain only some of the weakness.
Apple escalates claims against OpenAI, alleging evidence destruction in trade-secrets fight
In a new court filing, Apple accused OpenAI of actively destroying evidence tied to a trade-secrets dispute involving a former iPhone engineer. The company also pressed claims tied to alleged downloads of confidential information.
Duolingo shares jump after results point to steady user momentum, according to Yahoo Finance
A Yahoo Finance report highlighted that Duolingo’s second-quarter revenue rose 18% year over year, using the framing of a “Netflix-like comeback” after a period of volatility in the online learning category.
Netflix confirms production of Korean series “Materesa (WT),” led by “Queen of Tears” director and writers behind “The East Palace”
The streamer says its next Korean mystery drama, centered on a cold-blooded criminal psychologist who probes unsolved murders, is in production and has set a cast for “Materesa (WT).”
FTC and 22 States Sue Amazon, Alleging It Secretly Marked Up Ads Shown to Marketplace Sellers
The federal competition regulator and a coalition of states claim Amazon undercut third-party sellers on its platform by allegedly embedding surcharges into advertising terms.
FTC lawsuit by 22 states targets Amazon’s ad auction pricing, putting focus on high-margin advertising
The U.S. Federal Trade Commission says Amazon.com secretly inflated prices in its advertising auctions for more than seven years, while states joined the agency in the legal challenge.
Jensen Huang’s “Buy at a Discount” remark returns to focus as Nvidia shares rise and an AI basket gains
A CEO message to investors in June has been replayed after Nvidia’s stock moved higher over the following months, alongside gains in a broader AI peer group. Analysts caution that short-term trading often reflects many forces beyond a single CEO comment.
AMD says it is expanding its AI infrastructure footprint in Saudi Arabia
The chip designer announced a new platform initiative in Saudi Arabia, while investors appeared focused on how quickly the move could translate into additional AI-related revenue. AMD shares were little changed in Monday premarket trading.
Nvidia shares show a rare trading pattern, underscoring how investors are rethinking semiconductor correlations
A market-linked read of Nvidia’s stock behavior suggests its relationship with broader semiconductor moves has shifted, a change that can affect hedging, positioning, and how traders interpret near-term momentum.