The Race for a General Policy: Pi, RT-2, Groot and the Foundation Model Era in Robotics
Defining Robotics Foundation Models
Robotics foundation models represent a shift from task-specific control stacks to generalist vision-language-action architectures. Unlike traditional modular pipelines that separate perception, planning, and actuation, foundation models attempt to map raw sensor inputs directly to joint-level or wrist-level commands through large-scale training. The underlying premise is that scaling data, compute, and model capacity will yield policies capable of zero-shot or few-shot generalization across unstructured environments.
These models are typically trained on multimodal datasets comprising teleoperation trajectories, simulation rollouts, and real-world manipulation logs. The training objective remains fundamentally different from computer vision or natural language processing models. Robotics policies must satisfy strict temporal coherence, physical safety constraints, and deterministic actuation limits. A model that predicts plausible language or image tokens does not automatically translate to stable robot control. The gap between statistical correlation and physical execution remains the primary engineering hurdle.
The current landscape is dominated by three major initiatives: Physical Intelligence (Pi), Google DeepMind's RT-2 and Groot, and a broader ecosystem of open-weight and academic models. Each approaches the general policy problem with distinct data strategies, architectural choices, and deployment timelines. Evaluating them requires separating research milestones from commercial readiness.
The Current Landscape: Pi, RT-2, and Groot
Physical Intelligence (Pi)
Physical Intelligence, founded by former DeepMind and Google Brain researchers, has positioned itself at the intersection of foundation models and embodied AI. The company released Pi-0 and Pi-1, which are vision-language-action models trained on proprietary teleoperation datasets. Pi's approach emphasizes scalable data collection through human teleoperation and reinforcement learning from human feedback, aiming to produce policies that generalize across robot platforms without extensive per-task fine-tuning.
Grading Pi's claims requires strict adherence to deployment reality. As of the latest public disclosures, Pi operates primarily at the pilot and research deployment stage. The company has demonstrated generalist control on simulated and select real-world manipulator setups, but no mass-shipped hardware product ships with Pi models pre-installed. Independent verification of Pi's performance comes from technical blogs, open-weight model releases, and early pilot partnerships with academic and industrial labs. The company's roadmap targets broader hardware integration, but commercial shipping remains contingent on safety validation, latency optimization, and supply chain readiness for compliant actuators and compute modules.
Google DeepMind RT-2 and Groot
Google DeepMind's Robotics Transformer 2 (RT-2) was published as a peer-reviewed research system demonstrating that large vision-language models can be fine-tuned for robot control. RT-2 processes visual observations and language instructions to generate action tokens, bridging semantic understanding with physical execution. Subsequent work introduced Groot, a robot foundation model designed to scale data and model capacity for general-purpose manipulation. Groot focuses on improving cross-embodiment generalization and reducing the need for task-specific retraining.
RT-2 and Groot remain firmly in the research and pilot deployment tier. Google has published detailed methodology papers, open-sourced select model weights, and demonstrated capabilities on academic manipulators and simulated environments. The models have been integrated into research pilots using standard industrial arms and custom testbeds, but no commercial hardware ships with RT-2 or Groot as a default control stack. Independent reporting confirms that deployment is limited to controlled lab settings, academic collaborations, and early-stage industrial pilots. The grading hierarchy places these initiatives at pilot deployments second, with announcements last. Real-world reliability, fault tolerance, and integration costs remain the next validation steps.
Grading the Claims: Hardware, Pilots, and Announcements
Evaluating foundation model claims requires a strict hierarchy of evidence. Shipping hardware provides the most reliable signal because it forces engineers to confront real-world constraints: actuator latency, sensor drift, power management, and safety interlocks. Pilot deployments offer the second tier, demonstrating functional generalization in semi-controlled environments but still allowing for environmental tuning and human oversight. Announcements represent the lowest tier for grading, as they frequently precede years of integration work and do not guarantee operational readiness.
Applied to the current foundation model race:
- Shipping Hardware: No major foundation model currently ships pre-installed in commercial humanoid or manipulator platforms. Integration requires custom middleware, real-time inference optimization, and hardware certification.
- Pilot Deployments: Pi, RT-2, and Groot are actively used in research pilots and early industrial trials. These deployments validate generalization capabilities but require significant engineering overhead for deployment, monitoring, and fallback control.
- Announcements: Multiple vendors have announced foundation model roadmaps, open-weight releases, and partnerships. These announcements indicate strategic direction but do not substitute for deployment data or safety certification.
Independent verification remains essential. Peer-reviewed publications, open-weight model benchmarks, and pilot deployment reports provide the only reliable signals. Rendered concept videos and press releases should be treated as directional indicators, not operational proof.
India Availability and Pricing Realities
Foundation models themselves are software architectures, not standalone products. In India, availability depends on enterprise licensing, cloud inference access, or open-weight weight downloads combined with local compute infrastructure. The pricing structure typically follows a hybrid model: software licensing fees, compute costs for inference, and integration services for deployment.
For Indian enterprises and research institutions, landed costs break down as follows:
- Software Licensing: Enterprise foundation model licenses for robotics typically range from ₹15 lakh to ₹50 lakh annually per deployment, depending on usage tiers, support levels, and commercial vs. academic pricing.
- Compute Infrastructure: Real-time inference requires high-performance GPUs or specialized AI accelerators. Landed costs for compliant compute modules in India range from ₹8 lakh to ₹25 lakh per node, factoring in import duties, GST, and local distributor margins.
- Integration Services: Custom middleware development, safety validation, and hardware integration typically add ₹10 lakh to ₹30 lakh per pilot deployment. Indian system integrators charge premium rates for compliance certification and on-site commissioning.
- Hardware Components: Foundation models require compatible actuators, controllers, and sensors. Domestic assembly reduces landed costs, but high-precision harmonic drives and torque sensors remain largely imported. Complete pilot-grade manipulator platforms in India range from ₹25 lakh to ₹80 lakh, excluding integration and software licensing.
Availability is growing but remains concentrated in tier-1 tech hubs. Indian startups and research labs are piloting foundation models for pick-and-place, assembly, and quality inspection. However, full commercial rollout depends on reducing inference latency, improving safety certification pathways, and lowering integration costs. Until then, foundation models in India function as advanced research tools and early-stage pilots rather than production-ready solutions.
The Path to a General Policy
Achieving a true general policy requires solving three interconnected problems: data scaling, sim-to-real transfer, and safety validation. Data scaling depends on diverse, high-quality teleoperation and real-world logs. Sim-to-real transfer demands physics-accurate simulation and robust domain randomization. Safety validation requires deterministic fallback controls, formal verification, and extensive field testing.
The race to a general policy will not be won by announcements alone. It will be determined by pilot deployment metrics, independent benchmarking, and hardware integration readiness. Foundation models will gradually transition from research tools to commercial components as inference costs drop, safety frameworks mature, and integration standards stabilize. The next two years will likely show increased pilot deployments, clearer pricing models, and measurable ROI in structured environments. Generalization beyond controlled settings remains a longer-term objective requiring sustained engineering investment and rigorous validation.
References
- Google DeepMind. (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. Nature. https://www.nature.com/articles/s41586-023-06285-6
- Google DeepMind. (2024). Introducing Groot: A Robot Foundation Model for General-Purpose Manipulation. DeepMind Blog. https://deepmind.google/discover/blog/groot/
- Physical Intelligence. (2024). Pi-0 and Pi-1: Vision-Language-Action Models for Generalist Robot Control. Physical Intelligence Technical Reports. https://physicalintelligence.company/
- Google DeepMind. (2023). RT-2 Technical Blog and Model Weights. DeepMind Research. https://deepmind.google/discover/blog/rt-2/
- Independent Robotics Industry Reports. (2024). Foundation Model Deployment Benchmarks and Pilot Validation Studies. arXiv Preprints and Peer-Reviewed Publications. https://arxiv.org/search/?query=robotics+foundation+models+deployment
✓ Key takeaways
- •Hands-on view of The Race for a General Policy: Pi, RT-2, Groot and the Foundation Model Era in Robotics inside our Robotics Foundation Models library.
- •Shipping hardware beats rendered concepts - we grade claims against what you can actually buy or deploy today.
- •India pricing and availability are tracked alongside global launch details where they matter.
References
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Introducing Groot: A Robot Foundation Model for General-Purpose Manipulation
- Pi-0 and Pi-1: Vision-Language-Action Models for Generalist Robot Control
- RT-2 Technical Blog and Model Weights
- Foundation Model Deployment Benchmarks and Pilot Validation Studies
Related articles
More in Robotics Foundation Models →

