Seraphic and CrowdStrike: A New Chapter for Browser Security

Read more
close icon

September 1, 2026

From Pixels to Action: The Software Stack for Physical AI

Type

Deep Dives

Vikram Venkat

Vikram Venkat

“You see, but you do not observe,” Arthur Conan Doyle’s Sherlock Holmes admonishes Dr. Watson in A Scandal in Bohemia. In the physical AI world, cameras and sensors (as described in the earlier article) enable seeing; to move to observation and action, a unified software stack is essential.

At its core, physical AI is about closing the loop from signal to decision and action. To do so, a typical software stack involves multiple layers, each of which progressively evolves the system’s understanding of the world – reducing uncertainty while increasing the abstraction of the dynamic environment it operates in.

Layer cake: The physical AI software stack

Sensor fusion and data movement: The physical AI software stack sits atop the sensors mentioned earlier. The first layer of the software stack requires combining the data received from all these sensors. This requires timestamping inputs received from each sensor and synchronizing with other sensors’ inputs to ensure consistency across different data sources. This is a complex problem that requires very accurate time measurement and synchronization, as synchronization errors can propagate through downstream perception and decision-making. Typical systems are a hybrid of classical filtering methods and learned components. Additionally, this is a latency and scheduling problem, especially as the amount of data collected increases in bandwidth-heavy systems, and multiple different sensors need to be orchestrated in a time- and compute-efficient manner. 

While this is a relatively commoditized layer, efficient architectures that enable rapid and accurate multimodal sensing, paired with self-calibrating pipelines that integrate closely with the underlying hardware stack, can still deliver meaningful differentiation through improved performance and reliability.

Perception: This layer moves the system from signals toward the first semblance of meaning. The perception layer has two major components, geometric and semantic perception, which together help create a structured representation of the world across maps, surfaces, objects, and labels. This is a crucial bridge to move from vision to reasoning.

Geometric perception involves understanding the three-dimensional world the system operates in – identifying depth and coordinates while ensuring spatial consistency. Classical geometry pipelines required calibration and stable viewpoints; recent research has enabled a shift toward more foundational geometric models that enable dense 3D reconstructions with fewer calibration requirements. Semantic perception involves object detection, segmentation, recognition, tracking, and pose. State-of-the-art stacks include open-world, promptable perception with zero-shot generalization. Perception often includes mapping and creating a persistent memory of space that can be used downstream.

Many of the core techniques in this space are relatively mature, with current research increasingly focused on hybridizing classical and learned models. Defensibility at this layer can come from proprietary datasets – often a combination of real-world and simulated data that cover a wide range of edge cases and steady-state scenarios.

World models: This is where the stack first moves into reasoning, enabling prediction and anticipatory intelligence. World models build a predictive model of the environment, understanding how objects move, how the system’s actions affect the environment, and estimating future trajectories. This layer of the stack is still developing, with recent research focusing on latent world models integrated with reinforcement learning and on self-supervised multimodal prediction models. Ensuring reliability across long-horizon tasks is also a critical challenge.

This can serve as a central reasoning layer within the stack, maintaining representations of the environment that inform downstream planning and control steps. Consequently, its integration complexity and importance to system performance can also make this an especially promising area for differentiation and defensibility.

Planning, decision-making, and task execution: Having understood the world, the system can now take actions aligned with real-world objectives. These actions are often organized in a hierarchical pyramid. At the bottom are atomic actions, such as locomotion (e.g., driving or walking), manipulation (e.g., grasping or twisting), or material/energy transfer (e.g., applying heat or paint). These atomic actions can then be combined into skill primitives, such as pick-and-place in warehouses or tool use in industrial maintenance. Skills combine to execute task-level actions, such as assembling a component or cleaning a room. This layer closes the loop from sensing to action and is typically integrated with standard control systems.

Cutting-edge research has moved beyond basic algorithms to a broader, system-level approach – ensuring that actions can be completed within safety constraints and real-world limitations. Pre-trained Vision-Language-Action models, increasingly adapted through post-training and fine-tuning, with end-to-end policy control, are already beginning to be deployed in real-world systems. In addition, closed-loop data engines that enable relabeling, fine-tuning, and failure triage are emerging to drive autonomy. This requires significant contextualization across different use cases and workflows, as well as close hardware integration to ensure the efficacy of actions. Consequently, this layer of the stack can be highly defensible as well.

Alongside these, multiple different parts of the stack need to be developed to ensure real-world reliability. These include evaluation and observability, security systems that protect data in transit across the stack, and data pipelines that drive real-world data collection and closed-loop retraining. These can also be highly defensible elements of the stack and create the potential for differentiated, cross-vertical platforms. 

Building a winning physical AI platform

As evidenced throughout this series, multiple exciting themes are emerging across both the hardware and software layers of physical AI. Cutting-edge research is underway across the entire tech stack, as well as across verticals and use cases. Startups building in this space need to follow three key principles as they develop world-changing solutions.

Fusing the real and the simulated: Real-world data remains critical; however, it is also hard to collect and annotate. A major imperative for many builders in this space is to reach real-world deployment as safely and quickly as possible to collect actual data. These use cases should also be carefully selected to reflect the variety and complexity of environments and tasks. In addition, data pipelines need to be carefully built to integrate real and simulated data with correct weighting to drive accuracy and reduce costs. Companies that crack data collection, annotation, and integration will likely have a highly defensible position in the ecosystem.

Ensuring deep integrations: Companies that can build systems closely integrated across the entire stack will be well-positioned in the physical AI world. Accuracy, latency, cost, and other key parameters depend on how well the system’s components work together. Furthermore, given the preceding waves of automation, there are already several widely used systems and sensors; ensuring compatibility with these systems would simplify deployment and enable unified outcomes. Integration also drives stickiness, making solutions harder to displace. While vertical integration offers benefits, it does not necessarily mean that a startup needs to build all parts of the stack; simple plug-ins and efficient orchestration can achieve the same outcome.

Memory and context as differentiators: Similar to the generative AI world, memory and context are emerging as key moats. Startups that enable persistent memory of geometry, space, and actions, grounded in the context of the environment and use case, can enable more reliable and contextually appropriate actions. To do so, companies need to innovate in memory architecture and build ontologies and semantic maps tailored to the verticals they operate in.

The opportunity for physical AI

Physical AI has the potential to be truly disruptive and “Net New.” We believe value will accrue across horizontal layers – including sensing infrastructure, world models, and decision-making – as well as vertical-specific solutions. Some of the most advanced use cases today are in warehouse automation, logistics, and manufacturing; while healthcare, security, and many other service industries can benefit as collaboration, both human-machine and machine-machine, advances.

Physical AI offers a world of opportunity, and era-defining companies are yet to emerge. At Cota Capital, we are excited to continue exploring this evolving ecosystem. Our portfolio companies operate across different parts of the physical AI stack, including Rhombus, Atomathic AI, Activ Surgical, Quadric, FlowFuse, Ramen, and others. We look forward to partnering with additional founders building the next generation of physical AI. If you are building in this space, reach out to us.

times
#
# #