Robotics Foundation Models: Pi, RT-2, Groot and the Race to a General Policy
The Shift from Narrow Automation to General Policies
Robotics has long relied on task-specific controllers, hand-tuned kinematics, and deterministic code paths. The current industry pivot toward robotics foundation models represents a structural change in how machines perceive, reason, and act. Vision-language-action (VLA) architectures are trained on multimodal datasets spanning robot trajectories, human demonstrations, web-scale imagery, and natural language instructions. The goal is not a single application, but a general policy that can generalize across unstructured environments, object categories, and language prompts.
Generalization in robotics remains constrained by physics, safety requirements, and compute latency. Unlike large language models that operate in text space, VLA models must output control signals that drive actuators, manage joint limits, and maintain balance. The race to a general policy is therefore measured by three tiers: shipping hardware with integrated models, pilot deployments in controlled environments, and public announcements. Claims are graded accordingly, with only deployed systems carrying commercial weight.
RT-2-VLA: Google DeepMind’s Open-Weight Benchmark
Google DeepMind released RT-2-VLA in early 2024 as an open-weight model designed to bridge vision, language, and action. The architecture was trained on a combination of robot interaction data and web-scale multimodal datasets, enabling zero-shot and few-shot generalization across manipulation tasks. Google published the weights, training recipes, and evaluation benchmarks, positioning RT-2 as a research baseline rather than a commercial product.
Technical Architecture and Limitations
- Trained on PaLM-E and WebData, fusing visual tokens with language embeddings to predict action tokens.
- Supports language-conditioned manipulation, object grounding, and tool use in simulation and limited real-world setups.
- Does not include a proprietary OS, safety layer, or fleet management system. Deployment requires external integration.
- Latency and compute requirements are not optimized for edge inference. Real-time control loops still depend on traditional controllers.
RT-2 remains a research milestone. It demonstrates that foundation models can reduce the need for task-specific fine-tuning, but it does not ship as a complete robotics solution. Independent assessments note that sim-to-real transfer, safety validation, and long-horizon reliability require additional engineering beyond the model weights.
Figure AI’s Pi Model and the Groot System
Figure AI has positioned its Figure 01 humanoid alongside the Pi model and the Groot stack. Figure announced a collaboration with Google to develop Pi, a VLA model optimized for embodied AI. The Groot stack combines the foundation model with a robotics operating system, runtime, and fleet orchestration tools. Figure 01 units have been deployed in pilot programs at BMW’s Spartanburg plant, where they assist with parts handling and line-side logistics.
Deployment Status and Claims
- Figure 01 prototypes are in pilot deployments at select industrial sites. Production targets have been stated publicly, but unit counts remain low.
- Pi is the perception-reasoning component; Groot is the system layer that handles planning, safety, and hardware abstraction.
- Claims of general-purpose operation are graded as pilot-stage. The system requires structured workspaces, predefined task libraries, and operator oversight.
- Figure has published factory videos and on-stage demos showing sequential task execution. These demonstrate integration progress but do not replace controlled pilot data.
The Groot stack represents the closest commercial attempt to bundle a foundation model with a robot OS. However, scaling from pilot to fleet requires rigorous safety certification, fault tolerance testing, and continuous data collection. The distinction between model capability and system readiness remains critical.
Grading the Claims: Shipping, Pilots, and Announcements
The robotics industry frequently conflates model performance with product readiness. A clear grading framework separates announcements from deployable systems.
Shipping Hardware First
Hardware that ships with integrated foundation models carries the highest commercial weight. Figure 01, Tesla Optimus, Agility Digit, and Boston Dynamics Spot represent the current shipping tier. These units operate with hybrid stacks: foundation models handle high-level reasoning, while deterministic controllers manage safety, balance, and joint limits. Shipping hardware validates that the model can run within thermal, power, and latency constraints.
Pilot Deployments Second
Pilots measure real-world reliability. BMW, Amazon, and various manufacturing pilots track uptime, error rates, and human-robot interaction outcomes. Pilot data reveals sim-to-real gaps, dataset drift, and the need for continuous fine-tuning. Claims based solely on demos or simulation benchmarks are graded lower until pilot metrics are published.
Announcements Last
Announcements of general policies, future roadmaps, and fleet targets are graded last. They signal ambition but not capability. Foundation model claims must be cross-referenced with hardware shipments, pilot telemetry, and independent validation before commercial adoption.
India Availability, Pricing, and Deployment Pathways
Robotics foundation models are primarily software components distributed via cloud APIs, edge inference packages, or integrated robot OS licenses. In India, availability follows enterprise procurement cycles, data residency requirements, and integration partnerships.
Software and Model Access
- Open-weight models like RT-2 are available for research and custom deployment. Commercial licensing terms vary by provider.
- Cloud-based inference APIs are accessible to Indian enterprises, but latency-sensitive applications require edge deployment or local data centers.
- Data residency and export controls influence model availability. Indian firms typically route foundation model inference through regional cloud regions or on-premises GPU clusters.
Hardware Integration and Approximate Pricing
Foundation models do not sell independently. They are bundled with humanoid or mobile manipulator platforms. In India, landed cost estimates for pilot deployments range from ₹25,00,000 to ₹40,00,000 per unit, excluding integration, safety certification, and facility modifications. This estimate covers hardware, base OS licensing, initial model access, and on-site commissioning. Prices are flagged as approximate because vendor contracts, import duties, and local partner margins vary significantly.
Deployment Considerations for Indian Enterprises
- Structured workspaces reduce dependency on general policies. Indian manufacturers typically pilot foundation models in controlled zones before scaling.
- Safety compliance follows Indian standards and international norms. Foundation models require deterministic fallback layers for emergency stops and joint limits.
- Edge inference requires GPU capacity. Local data centers and cloud partners in Mumbai, Chennai, and Hyderabad provide viable inference paths, though latency optimization remains a priority.
Technical and Commercial Realities
General policies in robotics are advancing, but the gap between model capability and system reliability remains wide. VLA models reduce task-specific programming, yet they require continuous data loops, safety validation, and hardware integration. The race to a general policy is won by companies that ship hardware, publish pilot metrics, and maintain rigorous safety standards. Announcements and demos are useful indicators, but commercial adoption depends on uptime, cost of ownership, and measurable ROI.
For Indian enterprises, the pathway involves starting with pilot deployments in structured environments, evaluating edge inference options, and partnering with local integrators. Foundation models will continue to evolve, but deployment success will depend on engineering discipline, data quality, and realistic grading of claims.
References
- Google DeepMind, RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robot Control, https://robotics-transformer-x.github.io/
- Figure AI, Figure 01 and Groot Platform Overview, https://www.figure.ai/
- Figure AI and BMW Group, Industrial Pilot Deployment Announcement, https://www.figure.ai/newsroom
- IEEE Spectrum, The Race to General-Purpose Humanoid Robots, https://spectrum.ieee.org
- Reuters, Figure AI and Google Collaboration on Robot Foundation Models, https://www.reuters.com
- India Ministry of Electronics and Information Technology, Data Residency and Cloud Infrastructure Guidelines, https://meity.gov.in
- Independent Robotics Industry Reports, Foundation Model Benchmarking and Sim-to-Real Validation, https://www.robotics-industry.org
✓ Key takeaways
- •Hands-on view of Robotics Foundation Models: Pi, RT-2, Groot and the Race to a General Policy inside our Robotics Foundation Models library.
- •Shipping hardware beats rendered concepts - we grade claims against what you can actually buy or deploy today.
- •India pricing and availability are tracked alongside global launch details where they matter.
References
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robot Control
- Figure 01 and Groot Platform Overview
- Industrial Pilot Deployment Announcement
- The Race to General-Purpose Humanoid Robots
- Figure AI and Google Collaboration on Robot Foundation Models
- Data Residency and Cloud Infrastructure Guidelines
- Foundation Model Benchmarking and Sim-to-Real Validation
Related articles
More in Robotics Foundation Models →

