
Physical AI, AD and Robotics
Newsletter | Technical Briefing
Curated insights for AD professionals, Roboticists, Physical-AI engineers, Founders & Tech leaders
π Last 48 Hours
π₯ TOP STORY
π§ NVIDIA Cosmos 3 Can Be Post-Trained for Custom Agent Skills in a Single Day
- βNVIDIA’s technical blog details a workflow that lets developers post-train Cosmos 3 β the company’s physical-world foundation model β on domain-specific agent skills within one day, dramatically compressing the customization cycle.
- βThe approach targets the fine-tuning bottleneck that has made deploying large world models in production robotics pipelines impractical, enabling task-specific adaptation without retraining from scratch.
- βIf the one-day benchmark holds across diverse skill sets, it shifts Cosmos 3 from a research asset to a practical deployment tool β watch for adoption signals from NVIDIA’s robotics ecosystem partners.
- βπ Read More β
- What matters: One-day post-training on Cosmos 3 makes world-model customization a routine engineering task, not a research project.
π§ͺ TECHNOLOGY, RESEARCH & INNOVATION
π§ New VLA Inference Framework Cuts Latency by Skipping Redundant Visual Frames and Diffusion Steps
- βResearchers identify two compounding latency sources in Vision-Language-Action pipelines: redundant re-encoding of near-identical consecutive frames and multi-step iterative diffusion sampling β both avoidable without sacrificing output quality.
- βBy selectively skipping temporal redundancy in visual encoding and reducing diffusion sampling steps, the method targets the real-time deployment gap that has kept VLA models off low-latency robot controllers.
- βEfficient VLA inference is the missing link between strong lab generalization and factory-floor deployment β this line of work directly addresses that constraint.
- βπ Read More β
- What matters: Eliminating temporal redundancy in VLA pipelines is the practical path to real-time robot control without sacrificing generalization.
π€ GaitSpan Teaches Humanoids to Walk, Jog, and Run from a Single Policy β No Gait Schedules Required
- βGaitSpan proposes a single locomotion policy that spans walking through running on humanoid robots, eliminating the need for separate gait schedules, motion-clip imitation, or expert-switching architectures.
- βExisting multi-gait approaches trade flexibility for coverage β GaitSpan’s unified policy avoids mode-switching discontinuities that cause instability at gait transitions, a key failure point in prior work.
- βA single continuous locomotion policy is a prerequisite for humanoids operating in unstructured environments; GaitSpan’s approach sets a cleaner baseline for the field to build on.
- βπ Read More β
- What matters: Spanning the full walk-to-run gait range in one policy removes a fundamental brittleness from humanoid locomotion systems.
π PRODUCT, HARDWARE & MODEL LAUNCHES
π€ Booster T2 Humanoid Ships with Onboard NVIDIA Compute, Targeting Real-World Deployment
- βBooster Robotics has unveiled the T2 humanoid, equipped with powerful onboard NVIDIA compute β moving inference on-device rather than relying on cloud offload for real-time control.
- βOnboard GPU compute at humanoid scale is a hardware inflection point: it enables low-latency policy execution and reduces the network dependency that makes cloud-reliant robots fragile in dynamic environments.
- βThe T2 joins a growing cohort of NVIDIA-compute-equipped humanoids; the differentiator going forward will be which platforms pair that compute with the most capable onboard policies.
- βπ Read More β
- What matters: Onboard NVIDIA compute in the T2 signals that edge inference β not cloud dependency β is becoming the baseline expectation for production humanoids.
π XPENG Launches Camera-Only, HD-Map-Free VLA Robotaxi Testing in Guangzhou
- βXPENG has begun robotaxi testing in Guangzhou using a camera-only Vision-Language-Action system that operates without HD maps β a significant architectural departure from lidar- and map-dependent AV stacks.
- βDropping HD maps removes a major operational scaling constraint: map maintenance costs and coverage gaps have been a persistent barrier to rapid AV geographic expansion, particularly in China’s dense urban grids.
- βCamera-only, map-free VLA robotaxis are now being tested on public roads β the next milestone to watch is safety-disengagement rates compared to sensor-rich competitors.
- βπ Read More β
- What matters: XPENG’s map-free VLA robotaxi is a live test of whether language-conditioned vision alone can replace the expensive sensor and mapping stacks that have defined AV infrastructure.
π° BUSINESS, STARTUPS & INVESTMENT
π Waymo Expands Robotaxi Service to San Diego, Its Largest U.S. Geographic Push Yet
- βWaymo is launching commercial robotaxi service in San Diego, extending its rider-only operations beyond its established San Francisco, Phoenix, and Los Angeles markets.
- βEach new city adds a distinct operational domain β San Diego’s coastal geography, military corridors, and cross-border traffic patterns present edge cases not present in Waymo’s existing fleet data.
- βGeographic expansion pace is now Waymo’s clearest competitive moat signal β watch how quickly San Diego reaches the ride volume that made San Francisco operationally profitable.
- βπ Read More β
- What matters: Waymo’s San Diego launch is less about one city and more about proving that its operational playbook scales to new geographies without regression.
π§ SoftBank and Yaskawa Demonstrate Physical AI Development Powered by GPU Cloud Data Center
- βSoftBank and Yaskawa have jointly demonstrated a physical AI development pipeline that uses SoftBank’s AI data center GPU cloud to train and iterate on robot skills β bringing hyperscale compute directly into industrial robotics development.
- βYaskawa’s industrial robotics installed base combined with SoftBank’s GPU infrastructure creates a vertically integrated loop: real-world robot data feeds cloud training, which pushes updated policies back to the factory floor.
- βThe SoftBank-Yaskawa pairing is a template for how telco-scale compute owners can monetize physical AI β expect similar partnerships between cloud providers and industrial OEMs to accelerate.
- βπ Read More β
- What matters: SoftBank and Yaskawa’s GPU-cloud-to-factory pipeline shows that physical AI’s next infrastructure layer is being built by telcos and industrial OEMs, not just hyperscalers.
π THE BOTTOM LINE
β‘Foundation Model Customization::One-day post-training on Cosmos 3 collapses the gap between world-model research and production robotics deployment.
β‘VLA Efficiency::Temporal redundancy elimination in VLA pipelines is the near-term unlock for real-time robot control at scale.
β‘AV Architecture::XPENG’s camera-only, HD-map-free robotaxi is the most aggressive public test of whether VLA models can replace expensive sensor-and-map stacks in autonomous driving.
β‘Physical AI Infrastructure::The SoftBank-Yaskawa GPU-cloud-to-factory model signals that physical AI’s compute layer is being claimed by telcos and industrial OEMs β not just cloud hyperscalers.
β‘Humanoid Convergence::With onboard NVIDIA compute now standard in new humanoid launches and unified gait policies emerging from research, the question is no longer whether humanoids can move β it’s whether their onboard intelligence can match their hardware.

Worth forwarding to a colleague? Pass it along.
Β© 2026 Physical AI, AD and Robotics Β· DriveTech AI. All rights reserved. Privacy Policy