The Race to a General Policy: Pi, RT-2, and Groot in Context
The Push Toward a General Policy in Robotics
The robotics industry is undergoing a structural shift from bespoke, task-specific controllers to foundation models that aim to generalize across environments, tasks, and embodiments. The term "general policy" refers to a single neural network that can take raw sensory input, interpret natural language instructions, and output low-level joint commands without hard-coded rules for each new scenario. While the ambition is broad, the current reality is defined by constrained data pipelines, compute latency, and safety validation. Foundation models are not replacements for robot hardware; they are software layers that require real-world data, robust actuation stacks, and rigorous testing to function outside controlled environments.
This article evaluates the current state of robotics foundation models by grading claims against a strict hierarchy: shipping hardware first, pilot deployments second, and announcements last. We examine Pi (Physical Intelligence), RT-2 (Google DeepMind), and Groot (Figure AI), assess their technical claims, and map their availability and approximate pricing for the Indian market.
What Foundation Models Actually Do
Robotics foundation models, specifically Vision-Language-Action (VLA) architectures, unify perception, reasoning, and control into a single differentiable pipeline. Unlike traditional robotics stacks that separate computer vision, path planning, and inverse kinematics, VLA models are trained end-to-end on multimodal datasets. The training process typically involves:
- Real robot data: Teleoperation logs, joint trajectories, and tactile feedback from physical hardware.
- Web-scale vision-language data: Pretraining on image-text pairs to build world knowledge and language grounding.
- Simulation and domain randomization: Bridging the sim-to-real gap by exposing the model to varied lighting, textures, and object geometries.
The output is not a high-level plan but a sequence of low-level actions: joint angles, torque values, or gripper states. The model must operate in real time, typically under 100 milliseconds of inference latency, while maintaining stability and safety. Failure modes include hallucination of object states, overconfident policy outputs in unseen configurations, and compute bottlenecks when scaling to multi-sensor inputs.
Current Landscape: Pi, RT-2, and Groot
Three systems dominate current industry discussions around general policies: Pi, RT-2, and Groot. Each follows a VLA approach but differs in data strategy, deployment status, and commercialization path.
Physical Intelligence (Pi)
Physical Intelligence, spun out of Figure AI, developed the Pi-0 and Pi-0.5 models. Pi-0 was trained on a proprietary dataset of real-world manipulation episodes, emphasizing zero-shot task generalization. Pi-0.5 introduced diffusion-based action generation and improved temporal consistency. The model is designed to run on edge-compatible hardware, though it still requires substantial GPU resources for inference. Physical Intelligence has focused on manufacturing and logistics use cases, partnering with industrial automators for controlled deployments.
RT-2 (Robotics Transformer 2)
RT-2, published by Google DeepMind and UC Berkeley in Nature (2023), combined web-scale vision-language pretraining with robot interaction data. The model demonstrated improved compositional generalization and language grounding compared to its predecessor, RT-1. Google DeepMind open-sourced the weights, enabling academic and research replication. However, RT-2 remains primarily a research artifact. It has not been integrated into commercial hardware or scaled to production-grade inference pipelines. Its value lies in architectural insights rather than deployed capability.
Groot (Figure AI)
Figure AI introduced Groot as the foundation model powering the Figure 02 robot. Groot is trained on a continuous loop of real-world manipulation data, with a focus on scaling skills across different robot embodiments. The model is tightly coupled with Figure's hardware stack, including custom actuators, force-torque sensors, and safety controllers. Figure has moved Groot from announcement to pilot deployment, notably in BMW's manufacturing facilities, where it handles part picking, bin sorting, and quality inspection tasks.
Shipping Hardware vs. Pilots vs. Announcements
Evaluating foundation models requires separating software claims from hardware reality. The grading hierarchy clarifies where each system currently stands:
- Shipping hardware: Figure 02 (with Groot) has entered limited commercial production and pilot deployments at BMW. Physical Intelligence's Pi models run on third-party hardware in pilot settings but are not yet bundled as standalone shipping products. RT-2 has no associated commercial hardware.
- Pilot deployments: BMW manufacturing plants host active Figure 02 pilots. Academic labs and research consortia run RT-2 and Pi in controlled environments. These deployments validate generalization but do not represent scaled commercial availability.
- Announcements: Open-source weight releases, partnership MOUs, and conference demonstrations dominate the announcement tier. These signals indicate research direction and data strategy but do not confirm production readiness, safety certification, or ROI validation.
The race to a general policy is currently being won by data pipeline efficiency, not model architecture alone. Systems that can collect, clean, and label real-world robot data at scale will outperform those relying on simulation or synthetic datasets. Inference optimization, latency reduction, and safety validation remain the actual bottlenecks to commercialization.
India’s Position: Availability, Integration, and Pricing
India's robotics market is still dominated by industrial arms, collaborative robots, and AMRs. Foundation models are not commercially available as off-the-shelf products in India. Enterprises are currently relying on localized large language models, computer vision pipelines, and vendor-specific automation software. If foundation models enter the Indian market, adoption will follow three paths:
- Cloud API licensing: Remote inference via regional data centers. Estimated cost: ₹5–15 lakh per month for enterprise-tier access, depending on latency SLAs and concurrent robot fleets.
- On-prem private deployment: Edge GPUs or localized server racks for data sovereignty and low-latency control. Estimated cost: ₹20–50 lakh for initial hardware, plus ₹3–8 lakh per month for maintenance and model updates.
- Integration partnerships: Indian system integrators bundling foundation models with existing robot hardware. Pricing would be project-based, typically ₹10–30 lakh per deployment, excluding integration and validation.
Availability in India remains limited to R&D labs, technology parks, and early automation integrators. Regulatory considerations around data localization, AI safety guidelines, and manufacturing compliance will shape adoption timelines. Enterprises should treat foundation models as experimental layers until certified for industrial deployment.
The Gap Between Policy and Practice
General policies are not a single breakthrough but an incremental engineering challenge. Real-world constraints include:
- Compute latency: Vision-language-action models require substantial GPU resources. Edge deployment demands quantization, pruning, and hardware acceleration.
- Safety validation: Unseen scenarios trigger unpredictable outputs. Redundant safety controllers, force limiting, and human-in-the-loop oversight remain mandatory.
- Data curation: High-quality, labeled robot data is scarce. Synthetic data helps but does not replace real-world actuation feedback.
- ROI calculation: Foundation models increase flexibility but do not eliminate integration costs. Payback periods extend until deployment scales beyond pilot phases.
The industry is moving toward modular stacks: foundation models for high-level reasoning, deterministic controllers for safety-critical tasks, and simulation for continuous training. This hybrid approach balances generalization with reliability. Buyers should prioritize vendors with shipped hardware, verified pilot deployments, and transparent data pipelines over announcement-driven roadmaps.
References
- Physical Intelligence. "Pi-0 and Pi-0.5 Models." Accessed via official documentation and technical briefings. https://www.physicalintelligence.company/
- Google DeepMind & UC Berkeley. "RT-2: Vision-Language-Action Models Transfer Robotics Knowledge." Nature, 2023. https://www.nature.com/articles/s41586-023-06794-8
- Figure AI. "Groot Foundation Model and Figure 02 Robot." Official press materials and technical documentation. https://www.figure.ai/
- BMW Group. "Figure 02 Pilot Deployment at Manufacturing Facilities." Corporate press release. https://www.bmwgroup.com/en/news/general.html
- IEEE Spectrum. "The State of Robotics Foundation Models." Independent technical analysis. https://spectrum.ieee.org/robotics-foundation-models
- Physical Intelligence. "Technical Report: Pi-0.5." Research publication. https://www.physicalintelligence.company/research
✓ Key takeaways
- •Hands-on view of The Race to a General Policy: Pi, RT-2, and Groot in Context inside our Robotics Foundation Models library.
- •Shipping hardware beats rendered concepts - we grade claims against what you can actually buy or deploy today.
- •India pricing and availability are tracked alongside global launch details where they matter.
Related articles
More in Robotics Foundation Models →

