Light Origins Introduces Light-O1: A Breakthrough in Cross-Embodiment AI Training
The Launch of Light-O1 by Light Origins
Light Origins has recently introduced its pioneering product, the Light-O1, which stands as its first general-purpose embodied foundation model. This innovative release is part of the company's strategic mission to develop advanced foundation models tailored for Physical AI applications.
Human Action Learning from Videos
The Light-O1 model is designed to acquire reusable human action knowledge from structured human activities demonstrated in internet videos. The model then fine-tunes this knowledge for various robot embodiments and specific tasks. During extensive scaling experiments, it was consistently observed that increasing the pretraining budget significantly decreased prediction errors during subsequent adaptations.
The training process began with a foundational 4 billion-parameter model, from which several variations were developed, catered to six different pretraining budgets, ranging from 3.75 billion tokens to an impressive 120 billion tokens, the latter equating to approximately 100,000 hours of human action data. The resulting pretrained models were then fine-tuned for three different domains: public egocentric human data, data from the Unitree G1 robot, and data generated by Light Origins' in-house developed humanoid robot, LightBot. Notably, as the pretraining scale increased, both prediction loss and pose error rates for the models showed a significant decline, consistent with power-law behavior.
Real-World Applications of Light-O1
The Light-O1 showcase revealed remarkable capabilities of LightBot in executing multi-step household tasks. These include practical activities like retrieving slippers from a cabinet, sorting various types of waste even when the items were repositioned mid-task, and efficiently performing a towel handoff. Furthermore, Light-O1's application was demonstrated on the Unitree G1 robot, efficiently wiping surfaces and completing tasks in synchronized movements.
The Importance of Scalable Pretraining
The unique advantage of Light Origins lies in its methodology for gathering robot interaction data. Unlike traditional approaches, which require extensive hardware and manual data collection, Light Origins extracts structured 3D human movements from available internet videos, juxtaposing these actions with visual and textual representations and subsequently training an autoregressive model. This strategy aims to capture recurrent patterns of physical behavior, facilitating the adaptation of this knowledge for targeted applications.
Light-O1 effectively blends language-processing capabilities with visual and spatial comprehension, enabling robots to perform complex tasks with coordination. Beyond just practical implementations, Light Origins is also launching Light-O1-Preview—a model that interprets natural-language instructions, conceptualizes the required body movements, and outputs the necessary action sequences.
Scaling Towards Physical AGI
Looking ahead, Light Origins outlines its roadmap toward achieving Physical AGI through three main scaling paradigms: Scalable Pre-Training, Scalable Alignment, and Scalable Deployment. The introduction of Light-O1 underlines the company’s commitment to Scalable Pre-Training, utilizing large-scale human-action pretraining to establish a generalizable action prior that can be adapted across various robotic forms and tasks.
Earlier this month, alongside Light-O1, the introduction of LightNav-0 marked a significant development towards Scalable Alignment. This model harnesses a vast collection of real-world scenes sourced from the internet to create reusable simulated environments, yielding extensive navigation experience crucial for various robot types.
For their Scalable Deployment strategy, Light REACT evaluates recent physical interactions to deduce results influenced by external forces and environmental variables, adjusting the robots’ actions accordingly.
Building Infrastructure for the Future
Light Origins envisions a comprehensive, interlinked framework that integrates pretraining, alignment, and real-world application into a cohesive learning loop. To support this ambition, the company is expanding its data processing capabilities, now operating at a multi-thousand-GPU scale and achieving a weekly throughput of approximately 200,000 hours of video content.
According to Roger Jiang, the founder and CEO of Light Origins, the critical inquiry in the realm of Physical AI is whether there exists a pretraining signal that gains value with scale. The developments surrounding Light-O1 indicate that as the scale of human-action pretraining continues to grow, the prediction accuracies post-adaptation also improve across diverse robotic embodiments.
The significant funding acquired through a Pre-A round in August 2026 is strategically aimed at bolstering model training capabilities, enhancing multimodal data infrastructures, and fostering comprehensive software and hardware research and development.
About Light Origins
Founded in late 2024 by Roger Jiang, an accomplished figure previously associated with OpenAI and a contributor to ChatGPT, Light Origins is dedicated to developing foundation models for Physical AI. As the company moves toward the ambition of achieving Physical AGI, its models and deployment strategies remain intertwined within a continuous learning framework.