
大家都在押注更大的模型。但对于物理人工智能而言智能本身并非倍增器学习才是。彼得·路德维希2026年7月28日A billion machines will become autonomous or intelligent over the next ten years. Cars, trucks, tractors, mining haulers, defense systems, warehouse robots, humanoids—the physical economy will be rebuilt around software that perceives, decides, and acts.The prevailing assumption about how we get there goes something like this: models keep improving, world models mature, foundation models for robotics arrive, and autonomy falls out the other end. Intelligence is the whole game; scale the intelligence and the machines will follow.I’ve spent a decade helping Applied Intuition build the software infrastructure behind many of the world’s most ambitious physical AI programs, from software-defined vehicles and autonomous trucks to construction equipment, mining systems, defense platforms, and robotics. That vantage point has given me a front-row seat to where the industry is accelerating—and where it continues to slow itself down. I’m as bullish on the intelligence as anyone, but the prevailing assumption gets the math wrong. Deployed physical AI is a product of two variables: the capability of the models, and the capacity of the engineering system around them—i.e. how requirements become software, how software gets validated, and how validated systems get deployed, monitored, and improved. The industry has largely poured everything into the first variable while the second sits roughly where it was a decade ago, built for quarterly releases and hundred-person integration teams. Frontier intelligence running on a legacy engineering system doesn’t produce frontier outcomes, because the old engineering system is the limiting factor.The contrarian bet, then, isn’t against intelligence. It’s that the next order of magnitude in physical AI comes from making the engineering system as intelligent as the models it carries. The industry’s roadmap, however, rests on a handful of assumptions that once made sense but no longer match where physical AI is headed.Smarter Models Don’t Create Deployed MachinesThe gap between a capable model and a certified, operating machine is enormous, and model quality alone doesn’t close it. A model that’s 20% better in benchmark terms still has to be integrated with many other software components, tested across millions of scenario variations, traced against safety requirements, validated on hardware, rolled out to a fleet, and monitored in the field. In a typical program, that pipeline — not the model — sets the tempo. Teams take delivery of a meaningfully better model and then spend two quarters proving it’s safe to ship.This is why world models, as remarkable as they are, won’t get us to a fully autonomous future on their own. They are advancing faster than engineering organizations can keep up. Every leap in model capability relocates the bottleneck rather than eliminating it. The constraint moves downstream, from “can the machine perceive the world?” to “can we validate, integrate, and operate what the machine can now do?” A team whose validation cycle takes months is, in effect, throttling frontier AI down to the speed of its own process.Here’s the implication the industry hasn’t priced in: as models commoditize toward the frontier, two companies with access to the same intelligence will have wildly different outcomes. The difference will be determined by how fast their engineering systems can absorb what the models can do. Intelligence is becoming ubiquitous. The ability to operationalize it isn’t.Digital AI Doesn’t Transfer to Physical AIThe second assumption is subtler: that the agentic revolution happening in digital work will naturally extend to physical AI. Direct the coding agents and copilots at the autonomy stack, and the same productivity gains will follow.They won’t, because most digital AI stops at documents, conversations, and code. Physical AI work doesn’t live there. It lives in drive logs and sensor data, in simulation runs and hardware-in-the-loop test rigs, and in requirements databases and validation reports. It lives in fleet telemetry streaming from real vehicles on real roads and job sites. An agent that has never seen a disengagement, doesn’t know why a perception regression matters, and can’t trace a requirement to a test case, is not a productivity tool in this domain. It’s a liability with a friendly interface.Making agents genuinely capable in deploying physical AI itself is a frontier intelligence problem, one that is entirely different than training a bigger model. The agents need access to the actual data and tools of the trade—simulators, data pipelines, validation systems—through interfaces hardened enough to trust. They need embedded domain judgment and the accumulated knowledge of what “validated” actually means when the artifact ships into a multi-ton machine. And they need evaluation and governance built in, because in this domain a plausible answer and a correct answer can be separated by fatal consequences.When I say physical AI needs its own agentic platform, I’m talking about a platform that combines state-of-the-art models grounded in the data layer, tooling, and domain expertise of physical systems, with evaluation and governance native to the platform. It is not a chatbot bolted onto engineering tools, and it is not something you get by fine-tuning a general-purpose agent. It’s a different architecture, and it demands as much AI innovation as the models themselves.The future of physical AI wont be determined by the smartest models alone, but by the engineering systems that turn intelligence into autonomous machines at scale.Speed Doesn’t Compromise SafetyWhen discussing agentic capabilities in physical AI development, the reflexive objection is that agents have no place in safety-critical engineering. Automation means moving fast and perhaps getting a few things wrong. It’s a reasonable position when the thing in question weighs several tons.However, in safety-critical systems, the speed of your feedback loopisa safety mechanism. When validation takes weeks, teams test at irregular milestones. When it takes minutes, they test on every change. Problems surface earlier, when they’re cheap to fix. Coverage expands to orders of magnitude more scenarios. Requirements get implemented more accurately and verified more often. The slow, careful-looking process isn’t by its nature the safest one. It’s often the one where defects age quietly for months before anyone notices.What actually makes agents safe in this domain isn’t slowing them down. It’s drawing the line correctly. Automate development, validation, and operations workflows, so safety-critical issues get resolved faster, and refuse to automate certification, regulatory sign-off, and final engineering judgment. High-stakes agents propose; humans decide. Anything touching production systems runs behind approval gates. This isn’t a temporary concession while the models improve. It’s the correct permanent architecture for physical AI, in the same way that a well-designed autonomous vehicle has a defined operational domain rather than unlimited authority.When agents help build physical AI faster, that speed can make the end product safer.Models Don’t Compound. Systems Do.When you understand these assumptions are holding physical AI back, the logical next step is combining better intelligence with faster learning, creating an agentic flywheel where frontier models and frontier engineering systems feed each other.A continuously turning flywheel looks like this. A machine underperforms in the field. Agents mine the operational data to find out where and why, or synthetically generate the scenarios that expose the gap. Findings become requirements, and requirements become test cases. Test cases become validated software, which is deployed to the fleet. The fleet generates new data, which makes the models better, which makes the agents better, which speeds up the next turn of the loop. Every turn used to take months and a room full of specialists. Each stage that agents accelerate doesn’t just save time; it increases the number of turns, and each turn compounds both the speed and the intelligence. Smarter models turn the loop faster. A faster loop makes the models smarter. That’s the compounding effect the industry is leaving on the table when it treats intelligence as the whole game.We know this flywheel is real because we’ve been running it on ourselves. Applied Intuition has spent a decade at the frontier of physical intelligence — perception, simulation, validation, vehicle software across automotive, trucking, mining, agriculture, and defense. Over the past year we built an agentic platform grounded in that infrastructure. It’s called Dana, and our engineers have built more than a thousand internal apps and agents on it. With Dana, development cycles have become roughly 20x faster, with higher output quality. Deployments went from once every few weeks to multiple times a day. Applications that took months to build now take days or hours. And, tellingly, applications that would never have justified months of effort now get built regularly. When the cost of building drops by an order of magnitude, the set of things worth building expands by more than an order of magnitude.The intelligence made the engineering system possible; the engineering system made the intelligence matter. Neither alone gets you there.What This Means for the Next DecadeIf the industry keeps betting on intelligence alone, the next decade of physical AI looks like a slow one: dazzling demos, decade-long programs, and a widening gap between what machines can do in a lab and what’s actually operating in the world. The models will be extraordinary and the deployment curve will stay stubbornly flat, because every improvement will queue up behind engineering organizations that absorb change at last year’s speeds.Pair frontier intelligence with agentic engineering systems and the curve bends. Physical AI starts reaching its full potential. Farms that stabilize output through labor and climate shocks. Mines with continuous, safer extraction. Freight networks that self-route around disruption. Defense systems that hold under degraded conditions. Multi-hundred-billion-dollar markets converging on the same stack, with the learning loop at the center of all of them.Software ate the world by making it cheap to build applications for the digital economy. Physical AI will do the same for the physical one, but only if building intelligent machines becomes as fast and iterative as building software, without compromising the discipline safety-critical systems demand. That takes the best modelsanda reinvention of how we engineer, and the second half is the one almost nobody is building.Twenty years from now, we won’t remember which company had the best world model in 2027. We’ll remember which company figured out how to continuously turn intelligence into deployed systems. That’s the problem we’ve been working on.未来十年将有十亿台机器实现自主运行或具备智能。汽车、卡车、拖拉机、矿用运输车、国防系统、仓库机器人、人形机器人——实体经济将围绕能够感知、决策和行动的软件进行重建。目前普遍认为实现这一目标的过程大致如下模型不断改进世界模型日趋成熟机器人技术的基础模型最终形成自主性也随之而来。智能是关键所在提升智能规模机器自然会随之发展。过去十年我一直帮助Applied Intuition公司构建支撑众多全球最具雄心的物理人工智能项目的软件基础设施涵盖软件定义车辆、自动驾驶卡车、建筑设备、采矿系统、国防平台和机器人等领域。这段经历让我得以近距离观察行业的发展趋势以及它自身发展停滞不前的原因。我对人工智能的未来充满信心但目前普遍的假设存在逻辑错误。物理人工智能的部署取决于两个变量模型的能力以及围绕模型构建的工程系统的容量——即需求如何转化为软件、软件如何得到验证以及经过验证的系统如何部署、监控和改进。行业目前几乎将所有资源都投入到了第一个变量上而第二个变量却仍然停留在十年前的水平仍然沿用着季度发布和百人集成团队的模式。在老旧的工程系统上运行的前沿智能无法产生前沿成果因为老旧的工程系统本身就是制约因素。因此这种反主流观点并非反对智能本身而是认为物理人工智能的下一个飞跃阶段将来自于使工程系统本身达到与其承载的模型相同的智能水平。然而行业路线图却建立在一些曾经合理但如今已不再符合物理人工智能发展方向的假设之上。更智能的模型不会创建已部署的机器一个性能优异的模型与一台经过认证、可实际运行的机器之间存在着巨大的差距单靠模型质量的提升并不能弥合这一差距。即使一个模型在基准测试中提升了20%它仍然需要与其他众多软件组件集成在数百万种场景变化中进行测试对照安全要求进行验证在硬件上进行验证部署到整个机队并在现场进行监控。在一个典型的项目中决定项目进度的是整个流程而不是模型本身。团队拿到一个性能显著提升的模型后还需要花费两个季度的时间来证明其安全性才能最终交付使用。这就是为什么世界模型虽然卓越但仅靠它们本身无法带我们走向完全自主的未来。它们的进步速度远远超过了工程组织的响应速度。模型能力的每一次飞跃都只是将瓶颈转移到了其他地方而不是消除它。限制因素向下游转移从“机器能否感知世界”变成了“我们能否验证、集成并运行机器现在能够做到的事情”一个验证周期需要数月之久的团队实际上是在将前沿人工智能的速度限制在自身流程的速度之内。行业尚未充分考虑以下影响随着模型向前沿领域商品化两家拥有相同情报资源的公司最终会得出截然不同的结果。这种差异取决于它们的工程系统吸收模型功能的速度。情报正变得无处不在但将其转化为实际应用的能力却远未普及。数字人工智能无法转化为物理人工智能第二个假设更为微妙数字工作中正在发生的智能体革命自然会扩展到物理人工智能领域。引导编码智能体和副驾驶关注自主系统架构同样的生产力提升也将随之而来。他们不会因为大多数数字人工智能都止步于文档、对话和代码。而物理人工智能的工作并非如此。它存在于驾驶日志和传感器数据中存在于模拟运行和硬件在环测试平台中存在于需求数据库和验证报告中。它存在于来自真实道路和作业现场真实车辆的车队遥测数据流中。一个从未经历过用户脱离、不了解感知回归为何重要、也无法将需求追溯到测试用例的智能体在这个领域并非生产力工具。它只是一个界面友好的累赘。让智能体真正具备部署物理人工智能的能力是前沿智能领域的一大难题这与训练一个更大的模型截然不同。智能体需要通过足够可靠的接口访问实际数据和工具——模拟器、数据管道、验证系统等等。它们需要具备嵌入式领域判断能力以及在将产品交付给一台重达数吨的机器时“验证”的真正含义。此外它们还需要内置评估和治理机制因为在这个领域一个看似合理的答案和一个正确的答案之间可能存在致命的后果。我所说的物理人工智能需要其自身的代理平台指的是一个将基于物理系统数据层、工具和领域专业知识的尖端模型与平台原生评估和治理功能相结合的平台。它并非简单地将聊天机器人附加到工程工具上也不是通过微调通用代理就能实现的。它是一种不同的架构并且对人工智能创新提出了与模型本身同等的要求。物理人工智能的未来不仅仅取决于最智能的模型还取决于将智能大规模转化为自主机器的工程系统。速度并不影响安全性在讨论物理人工智能开发中的智能体能力时人们往往会反驳说智能体在安全攸关的工程领域没有立足之地。自动化意味着快速行动但也可能导致一些错误。当目标物体重达数吨时这种观点不无道理。然而在安全关键型系统中反馈循环的速度本身就是一种安全机制。当验证需要数周时间时团队只能在不规则的里程碑节点进行测试。而当验证只需几分钟时他们就能对每一次变更进行测试。问题会更早地被发现从而降低修复成本。测试覆盖范围也因此扩展到更多场景。需求能够得到更准确的实现并被更频繁地验证。缓慢而谨慎的流程本质上并非最安全的。在这种流程中缺陷往往会悄无声息地存在数月之久直到有人发现。在这个领域真正确保智能体安全的并非降低其运行速度而是正确划定界限。自动化开发、验证和运维工作流程以便更快地解决安全关键问题同时拒绝自动化认证、监管审批和最终工程判断。高风险智能体提出方案由人类做出决定。任何涉及生产系统的事项都必须经过审批流程。这并非模型改进期间的临时妥协而是物理人工智能的正确永久架构正如设计良好的自动驾驶汽车拥有明确的运行范围而非无限的权限一样。当智能体帮助更快地构建物理人工智能时这种速度可以使最终产品更安全。模型不会产生复合效应系统才会。当你理解这些假设阻碍了物理人工智能的发展时合乎逻辑的下一步就是将更智能的技能与更快的学习速度结合起来创建一个智能飞轮使前沿模型和前沿工程系统相互促进。一个持续运转的飞轮看起来是这样的一台机器在实际应用中性能不佳。智能体会挖掘运行数据找出问题所在及原因或者合成场景来暴露差距。发现的问题转化为需求需求转化为测试用例。测试用例转化为经过验证的软件并部署到整个系统中。系统生成新的数据从而改进模型改进智能体进而加快循环的下一轮。过去每一轮循环都需要数月时间并且需要一屋子的专家参与。智能体加速的每一个阶段不仅仅节省了时间它增加了循环次数而每一次循环都会同时提升速度和智能水平。更智能的模型能够更快地完成循环。更快的循环速度又使模型更加智能。这就是行业将智能视为全部时所忽略的复合效应。我们深知这种飞轮效应真实存在因为我们一直在自身实践中验证它。Applied Intuition 十年来一直致力于物理智能领域的前沿研究涵盖感知、仿真、验证以及汽车、卡车、采矿、农业和国防等行业的车辆软件。过去一年我们基于这一基础架构构建了一个智能体平台名为 Dana。我们的工程师已在其上开发了超过一千个内部应用程序和智能体。借助 Dana开发周期缩短了约 20 倍同时输出质量也显著提高。部署频率从几周一次提升到每天多次。过去需要数月才能构建的应用程序现在只需几天甚至几小时即可完成。更重要的是那些过去根本不值得花费数月时间开发的应用程序现在也开始定期构建。当构建成本降低一个数量级时值得构建的项目数量级也会相应增加。智慧使工程系统成为可能工程系统使智慧发挥作用。两者缺一不可。这对未来十年意味着什么如果业界继续只依赖智能那么未来十年物理人工智能的发展将会十分缓慢令人眼花缭乱的演示、长达十年的项目以及机器在实验室中的表现与实际应用之间日益扩大的差距。模型固然会非常出色但部署曲线却会始终保持平缓因为每一项改进都将排在那些以去年速度吸收变革的工程团队之后。将前沿智能与智能工程系统相结合曲线将发生转变。物理人工智能开始充分发挥其潜力。农场能够应对劳动力和气候冲击稳定产量。矿山能够持续、安全地开采。货运网络能够自动绕过中断。防御系统能够在恶劣条件下保持有效运行。数千亿美元的市场将汇聚到同一技术栈上而学习循环则是所有这些市场的核心。软件通过降低构建数字经济应用的成本彻底改变了世界。物理人工智能也将对物理世界产生同样的影响但这只有在构建智能机器的速度和迭代性能够与构建软件一样快并且不损害安全关键系统所要求的严谨性时才能实现。这需要最佳模型和对工程方式的彻底革新而后半部分几乎无人涉足。二十年后我们不会记得哪家公司在2027年拥有最佳的世界模型。我们会记住哪家公司找到了将情报持续转化为可部署系统的方法。这正是我们一直在努力解决的问题。