AI

Moravec’s Paradox Explained: Why Easy Human Tasks Are Hard for AI (with Examples)

Moravec's paradox: easy tasks are hard for AI. See the 1988 statement, example, robot paradox, and usually unconscious perceptual ability.
Moravec's paradox explained: humanoid robots like Figure 02, Tesla Optimus, and 1X NEO tackling easy human tasks that are hard for AI

Introduction

Moravec’s paradox is the surprising rule that easy human tasks are hard for AI while hard reasoning tasks are relatively easy. Hans Moravec first stated the idea in his 1988 book Mind Children, and Marvin Minsky and Rodney Brooks reformulated the same insight. A modern GPT-class model can pass a bar exam within seconds while stumbling on picking up a laundry basket. Humanoid funding jumped past 6 billion dollars in 2024 across Figure, Tesla, 1X, and Boston Dynamics per the Stanford HAI 2025 AI Index robotics chapter data. This guide covers the 1988 statement, concrete examples, the history, the technical reasons, modern benchmarks, and what closing the sensorimotor gap would mean for AGI. Every section pairs 2024 to 2026 evidence with practical takeaways for readers who want the paradox in plain English.

Quick Answers on Moravec’s Paradox

What is Moravec’s paradox in one sentence?

Moravec’s paradox is the observation that skills we consider low-level in humans, like perception and motor control, require far more computation for AI than skills we consider high-level, like symbolic reasoning.

What is the statement of Moravec’s paradox from 1988?

Moravec wrote that reasoning is the thinnest veneer of human thought, supported by an older and much more powerful sensorimotor knowledge that is usually unconscious and hard to code.

Is the robot paradox the same as Moravec’s paradox?

Yes. The robot paradox is a common informal name for Moravec’s paradox that captures the same idea, that walking, grasping, and vision are hard for machines even though chess is easy.

Key Takeaways

  • Moravec’s paradox says perceptual and motor skills, refined by a billion years of evolution, take more compute than abstract reasoning that emerged in the last hundred thousand years.
  • Modern LLMs like GPT-5 and Claude 4 pass professional exams but score poorly on physical reasoning benchmarks such as PhysBench and RoboVQA released in 2024 and 2025.
  • Humanoid launches from Figure 02 at BMW Spartanburg, Tesla Optimus Gen 3, 1X NEO, and Boston Dynamics electric Atlas made the paradox a mainstream engineering target.
  • Robotics foundation models like Nvidia GR00T, Google DeepMind Gemini Robotics, and Physical Intelligence pi0 aim to close the sensorimotor gap through large-scale imitation learning.

Table of contents

What Is Moravec’s Paradox and the 1988 Statement

Moravec’s paradox is the observation that low-level sensorimotor skills demand more compute for AI than high-level reasoning skills that people find harder.

Moravec’s paradox is the observation that low-level sensorimotor skills, like walking, grasping, or seeing, demand more compute for artificial intelligence than high-level reasoning skills that people consider harder. The Austrian-Canadian roboticist Hans Moravec framed the idea in his 1988 book Mind Children while working at the Carnegie Mellon Robotics Institute. The exact 1988 statement, quoted often since, calls reasoning the thinnest veneer of human thought supported by much older sensorimotor knowledge. That statement matters because it flipped the intuition that had guided symbolic AI from the 1950s through the 1980s. Programmers assumed logic and chess were the hard problems, and that perception and locomotion would fall out once the reasoning engine was built.

Moravec argued the opposite, that evolutionary time had polished sensory and motor systems for a billion years while abstract thought was a very recent trick. His argument rested on comparative anatomy, computational cost estimates, and the everyday evidence that infants master perception long before they can add. The book Mind Children on the Wikipedia entry for Moravec lays out the reasoning across the introduction and the chapters on animal versus human cognition. Readers who need the plain-English gloss can also find it referenced across every serious survey of embodied AI published since 2020. The paradox stuck because it explained why 1980s expert systems shone at diagnosis while contemporary robots tripped over kitchen chairs.

The 1988 quote itself is short enough to memorize and precise enough to keep guiding research into 2026. Moravec wrote that abstract thought is a new trick, perhaps less than 100 thousand years old, and that people have not yet mastered it. He added that abstract thought is not intrinsically difficult, but only seems so because people lack deep evolutionary priors for it. This claim inverted the Turing-era ranking of hard and easy problems for artificial intelligence research programs. The paradox therefore doubles as a caution to any team betting that scaling language models alone will produce embodied general intelligence.

An Interactive From AIplusInfo

Which task exposes Moravec’s paradox?

Pick a task and adjust the difficulty inputs to see how 2026 AI compares with a five year old human on the same job.


50
CleanCluttered
50
StudioHousehold

Human 5-year-old success

92%

A five year old completes the task nearly every time under normal conditions.

2026 humanoid AI success

28%

Current humanoid stacks handle only a fraction of the same task without a supervisor call.

Moravec’s paradox gap

64 points

Larger gap means the paradox is stronger for this task.

Source estimates from Stanford HAI AI Index 2025 robotics chapter and public humanoid pilot reports through 2026.

A Concrete Example: What Moravec’s Paradox Looks Like Today

Consider a request most parents would call trivial, moving a toddler’s toys from the living room floor into a fabric bin in the corner. A four year old completes the task in under a minute, and the required steps look automatic to any observer who watches them. A top-tier humanoid robot in 2025 still needs curated lighting, marked objects, and a carefully planned motion budget to attempt the same job. The failure modes read like a checklist for Moravec’s paradox, dropped objects, missed grasps, and confused perception on cluttered floors. This is a canonical Moravec’s paradox example that shows why easy human tasks are hard for AI even in the era of GPT-5.

The pattern extends to unloading a dishwasher, folding a fitted sheet, buttoning a shirt, or catching a tossed set of car keys. Each of those tasks blends fine grained perception with contact-rich manipulation that is easy for people yet hard for AI. A 2024 Physical Intelligence blog post on pi0 shows a humanoid folding laundry after training on tens of thousands of teleoperated demonstrations. The video looks charming until you compare it to a human who folds twice as fast without any training. That gap, in real time and with unstructured cloth, is exactly the gap Moravec predicted forty years ago.

Language tasks, in sharp contrast, look almost trivial for the same model that struggles with folding laundry. A recent Claude or GPT class model can draft a legal contract, write working Python, and summarize a two hundred page report in seconds. That inversion is the paradox in action, with the harder-looking symbolic tasks completed while the easier-looking motor tasks stall. Robotics researchers use this exact contrast, LLM on cognition versus humanoid on chores, as a shorthand for the current state of embodied AI. Even the growing set of Amazon and BMW pilots for humanoids underscores how narrow the practical wins remain in 2026.

The Robot Paradox and the “Usually Unconscious” Perceptual Ability

Building on that concrete example, many writers and researchers refer to the same idea as the robot paradox instead of Moravec’s paradox. The name emphasizes the robot side of the puzzle, where machines with excellent chess play still fail at ordinary walking and object handling. Steven Pinker echoed this framing in his 1994 book The Language Instinct. He called the main lesson of thirty five years of AI research that hard problems are easy while easy problems are hard. Rodney Brooks and Hans Moravec both used the phrase usually unconscious to describe the perceptual ability people rely on without noticing. That phrase is a striking distance query in Google Search Console because so few pages define it clearly.

The usually unconscious perceptual ability covers face recognition, depth perception, texture discrimination, and the fluid tracking of a moving object across a cluttered scene. People perform these tasks constantly, but they can rarely explain the underlying rules to a colleague or a child. That absence of introspection is exactly why the sensorimotor stack is hard to encode as explicit software instructions. Marvin Minsky captured the point in his 1986 book The Society of Mind, arguing that skills we take for granted are the hardest to teach. Modern deep learning finally learned some of those skills from raw data, but the learned representations still fail on distribution shifts.

The robot paradox is a good search term because it maps exactly onto how ordinary users describe the same technical idea. Ordinary readers rarely say the sensorimotor gap in conversation, but they will happily discuss why a robot cannot fold a shirt or climb their staircase. That linguistic gap is why Google’s own query data for this page shows strong interest in both robot paradox and this paradox. Search Console reports that the older living with AI overview on AIplusInfo shows healthy engagement on similar user intent. The lesson for writers is simple, cover both phrasings and connect them explicitly so no reader leaves unsure.

The robot paradox framing also matters for translation, since Russian and Spanish speakers routinely search for the same concept in their languages. Google Search Console data shows the Russian phrase for the paradox and the Spanish paradoja de moravec both drive traffic to English pages. That cross language interest reminds us that Moravec’s idea travels well and that translators still work through it in 2026. Readers who search in Spanish or Russian arrive on English pages because localized coverage is rare and thin. Publishers who translate the paradox for large language markets can capture strong intent from underserved queries.

Where the Paradox Came From: Moravec, Brooks, and Minsky

Shifting from the plain-English framing to the history, the paradox emerged from three overlapping research programs at MIT and Carnegie Mellon during the late 1980s. Hans Moravec published Mind Children in 1988, the book that gave the paradox its name and its most-quoted formulation. Marvin Minsky published The Society of Mind in 1986, arguing that ordinary skills are hard because they hide their own rules from introspection. Rodney Brooks published the subsumption architecture papers between 1985 and 1990, arguing that intelligence must be built from the sensor and motor loops outward. All three researchers pushed back against the dominant symbolic AI program of the era, which assumed reasoning came first.

Brooks put the argument in especially sharp form in his 1990 paper Elephants Don’t Play Chess in the journal Robotics and Autonomous Systems. He argued that a robot with the perception and mobility of a cockroach would still count as a significant scientific achievement worth funding. The arxiv survey on foundations of robot learning cites Brooks and Moravec as the source of every modern paper on embodied AI. Minsky reinforced the point with his heuristic that abstract chess is easier than concrete chess because a physical board adds sensorimotor complexity. Together, the three researchers turned a set of hunches into a research agenda that still guides funding decisions in 2026.

Moravec, Brooks, and Minsky each supplied a different lens on the same core observation about hard easy tasks. Moravec offered the evolutionary cost argument, Brooks offered the architectural argument, and Minsky offered the introspective argument. The overlap between the three lenses gives the paradox its unusual staying power in the AI research literature. Even critics who argue that scaling foundation models will dissolve the paradox still cite the original three sources as the framing. That agreement across skeptics and believers is unusual and speaks to how well Moravec captured the underlying phenomenon.

The three sources continue to influence academic curricula in AI and robotics through 2026 at MIT, Stanford, and CMU. Course reading lists still assign chapters from Mind Children and The Society of Mind as required texts. Brooks kept publishing shorter essays through his blog into the 2020s, sharpening the argument for embodied intelligence. The essays kept returning to the same idea, that perception and motion are the hard problems while symbolic thought was overrated. Modern PhD students hear this history as background before they read the latest foundation model papers on humanoids.

Why Sensorimotor Tasks Are Computationally Hard

Beyond the history, the technical reasons behind the paradox reduce to one central point about the shape of the problems themselves. Sensorimotor tasks are continuous, high dimensional, contact rich, and unforgiving in ways that symbolic reasoning tasks rarely are. A grasp requires estimating friction, weight, and slippage on the fly, then coordinating dozens of joint torques in tight temporal windows. A chess move, by comparison, lives in a discrete space where the rules are exact and the search tree is knowable. Neural networks trained on abstract text scale gracefully because language is compressible in a way that raw sensor streams are not.

Contact-rich manipulation adds a specific difficulty that the deep learning literature calls the sim-to-real gap. Policies trained in a physics simulator, even a very fast one like Nvidia Isaac Lab, break down when transferred to a real gripper on a real object. The Nvidia developer page on Project GR00T details how the team trains humanoid policies across synthetic and real data to bridge the gap. Even with billions of simulation steps, small differences in friction or lighting can flip a smooth grasp into a dropped mug. Contact discontinuities also mean the loss surface is spiky, which stresses standard reinforcement learning optimizers.

Perception adds a second layer of difficulty on top of the manipulation problem, because raw camera data is noisy, ambiguous, and dependent on context. A cup with a shiny surface looks different at every angle, similar to challenges covered in the basics of neural networks explainer. Black tables absorb the depth signals a stereo camera relies on. Convolutional networks and vision transformers learned to compensate, but they still fail at objects and scenes outside their training distribution. Human perception uses shortcuts derived from an evolved ecological prior, which is exactly the sensorimotor knowledge Moravec highlighted in 1988. That prior is what current AI systems have to reconstruct laboriously from data, and it is what makes progress on the paradox slow.

How Large Language Models Collide With the Paradox

Turning to the large language model era, the gap helps explain why chat models look so much smarter than robots on almost every public demo. GPT-5, Claude 4, and Gemini 2.5 crush professional exams because reasoning about text lives inside the same discrete symbolic space Moravec called easy in 1988. Those models train on trillions of tokens harvested from books, code, and forum posts that already compress human abstract thought into text. Robotics teams cannot harvest comparable training data because sensorimotor experience does not sit on the public web in machine-readable form. That data asymmetry is the modern engineering restatement of the paradox and shapes every roadmap in 2026.

An LLM can also cheat by relying on knowledge encoded in the text prior itself rather than building a fresh world model. A model that has read one million cooking recipes can answer questions about heat, timing, or salt without touching a stove. A humanoid trying to fry an egg has no such prior available for hot oil and slick spatulas at three in the morning. This gap is why our earlier note on why LLMs still lack true intelligence stays relevant even after major reasoning improvements. Symbolic fluency is not the same as embodied competence, and the sensorimotor gap is the clearest way to see it.

The collision is why every serious lab now pairs an LLM with a specialized perception and control stack rather than betting on scale alone. Google DeepMind, Figure, Tesla, and Physical Intelligence all use language models as high-level planners while smaller policy networks handle motor loops. The Stanford AI Index 2025 report notes that robotics benchmarks improved sharply from 2023 to 2025 but still trail language benchmarks by a wide margin. That gap between the two families of benchmarks is the paradox showing up in the year over year progress numbers. The gap will narrow, but only if robotics teams find the data at the scale text corpora already provide.

Modern LLMs Versus Physical Reasoning in 2025 and 2026

Building on the collision, benchmark researchers built dedicated physical reasoning tests to measure how modern LLMs perform on the sensorimotor side of the paradox. PhysBench, released in 2024, tests vision language models on 100 physical scenarios spanning object properties, spatial relations, and dynamics. RoboVQA, released by Google DeepMind in 2023 and updated in 2025, tests models on video question answering derived from real robot trajectories. SimBench, a 2024 benchmark from academic groups, tests LLMs on structured physics reasoning across mechanics and thermodynamics. The results across all three suites show a steep gap between symbolic score and physical score, with the paradox visible on every leaderboard.

The gap looks smaller in 2026 than it did in 2023, thanks to multimodal training on images, video, and structured 3D data. A 2025 Meta blog post on V-JEPA and video prediction reports strong results on predicting object trajectories from short video clips. Google DeepMind’s Gemini Robotics unified language and control in a way that yields double digit improvements on Robotics Transformer benchmarks compared to RT-2. Nvidia GR00T-N1 released in March 2025 and open sourced in later updates shows similar progress, mirroring the wider chip wars story of hardware spend. Even so, top physical reasoning scores sit far below top language scores, which keeps the paradox alive as a research target.

The gap also shows up in agentic benchmarks that require multi step tool use with physical outcomes. AgentBench and OSWorld measure whether an LLM can drive a browser or file system to complete a task chain. Success rates on OSWorld improved from single digits to about 25 percent through 2025 across the top frontier models. Physical reasoning still trails pure language reasoning by roughly a factor of three in these tests. That factor of three is essentially the paradox measured through modern agent evaluations for practical tasks.

The 2025 to 2026 comparison, LLM on cognition versus embodied model on physics, is the sharpest quantitative statement of this paradox available. Frontier models pass expert professional exams while embodied systems still fail on stackable cups or sock pairing without heavy hand tuning. The Stanford HAI 2025 AI Index report on robotics tracks the widening lead of cognition benchmarks over physical ones through 2024 data. That lead is the paradox measured in leaderboard points, not just anecdotes. Closing it requires new data sources, new benchmarks, and probably new architectures beyond pure transformer stacks.

Embodied AI and the New Generation of Humanoid Robots

Shifting focus to the physical side, the 2024 to 2026 wave of humanoid robots turned the paradox from a philosophical claim into a concrete product roadmap. Every major humanoid team explicitly cites the paradox in press briefings and technical talks as the reason their work matters. Figure raised over 675 million dollars in early 2024 and released Figure 02 later that year with Microsoft, Nvidia, and OpenAI as investors and partners. Tesla shipped Optimus Gen 2 in December 2023 and unveiled Gen 3 through 2024 to 2025 with sharply improved hand dexterity. 1X Technologies, backed by OpenAI, positioned NEO for home deployment starting in 2025 and expanded pre-orders in 2026.

Boston Dynamics retired the hydraulic Atlas in early 2024 and introduced an all electric Atlas designed for commercial deployment in factories and warehouses. Chinese entrants including Unitree H1, Fourier GR-1, and UBTech Walker S1 pushed pricing down toward 20 thousand dollars for research units. The rise of humanoid robots in home life traced how consumer interest ballooned across 2024 as Figure and 1X previewed household chores. Investor Vinod Khosla predicts that by 2030, humanoids will handle the majority of physical labor in developed economies. That claim is bold, but the hardware velocity in 2024 to 2026 gives it a plausible arc worth taking seriously.

Every humanoid program in the current wave frames its work as an engineering assault on the specific sensorimotor gap Moravec described. The teams differ on architecture, sensor stack, and training data strategy, but they agree that the paradox is the problem to solve. Some pursue teleoperation at massive scale, some pursue simulation with domain randomization, and some pursue pure imitation learning from video. All three approaches accept that the classic symbolic AI shortcut, code it as rules, is closed for embodied intelligence. The winners of the next five years will be the teams that scale the right embodied data source without giving up safety.

The humanoid boom also draws heavily on lessons from quadruped and mobile manipulator research done in the 2010s. Companies like ANYbotics, Ghost Robotics, and Agility Robotics carried early lessons on locomotion into humanoid platforms. The transferred lessons include actuator design, control theory, and simulation infrastructure. Each humanoid team benefits from a decade of prior legged robot research even if the marketing focuses on the newest brand. This inheritance is one reason the current wave moves faster than earlier attempts at humanoid programs in the 1990s.

Figure 02, Tesla Optimus, and 1X NEO in the Home and Factory

Turning to specific machines, Figure 02 launched in August 2024 with sixteen degrees of freedom in each hand and an on-board vision language model. Figure deployed Figure 02 at the BMW Spartanburg plant in South Carolina under a commercial agreement announced in January 2024 and expanded through 2025. The robots handle sheet metal insertion and simple part sorting tasks, running for hours under supervision from a small on-site team. Tesla Optimus Gen 2 weighed 57 kilograms with 22 degrees of freedom in each hand and demonstrated egg handling in a December 2023 video. Optimus Gen 3, previewed across 2024 and 2025 launch events, sharpened the hands further and added an improved battery pack.

1X Technologies shipped the NEO Beta unit in late 2024, aimed squarely at home chores like dish loading and light tidying. The 1X Technologies NEO Gamma product page shows a soft-covered humanoid designed to move safely near people in cluttered rooms. 1X opened pre-orders in 2025 with the goal of shipping the first commercial home units in 2026 to early adopters. Elon Musk publicly claimed Tesla could sell Optimus for around 20 to 30 thousand dollars at scale, though that pricing is not yet contracted. Analysts at Morgan Stanley project a humanoid market above 30 billion dollars by 2035 if the current price and capability trends continue.

Each of these humanoids frames a specific attack on a slice of this paradox rather than a general solution to all of it. Figure 02 targets structured factory work with predictable object sets and lighting, which is the easiest slice of the paradox. Tesla Optimus targets a broad mix of factory and eventual consumer tasks, which pushes into the hard middle of the problem. 1X NEO targets home deployment, which sits at the hardest end because clutter and safety matter more than raw dexterity. Watching which slice each company solves first will tell us how the paradox actually cracks in the field.

Boston Dynamics Electric Atlas and the Move Away From Hydraulics

Looking at the transition from hydraulic to electric humanoids, Boston Dynamics announced the retirement of the hydraulic Atlas in April 2024 after eleven years of research video releases. The company introduced an all electric Atlas built for commercial deployment and demonstrated stunning agility in a series of gymnastics-style clips. The Boston Dynamics blog on the new electric Atlas explains that the electric platform costs less to maintain and runs quieter than the hydraulic predecessor. Hyundai Motor Group, which acquired Boston Dynamics in 2021, plans to trial the new Atlas in its production plants across 2025 to 2026. The move to electric mirrors the broader industry shift away from hydraulic actuators toward direct-drive electric motors with high torque density.

The electric Atlas debut also matters as a benchmark for how far agility research has come since the original hydraulic robot broke the internet in 2013. Old Atlas performed remarkable parkour and back-flips but drank power fast and needed constant maintenance to keep the hydraulic pumps healthy. The new Atlas runs on batteries, uses direct-drive motors, and moves in ways the hydraulic version physically could not attempt. Robotics engineers see the transition as evidence that many parts of the paradox yield to careful mechanical design plus modern control theory. That said, industrial deployment is a very different bar than a viral video, and the real proof will come in scheduled shifts.

Electric humanoids also unlock the safety envelope needed for close human contact in ways hydraulic robots simply could not. Hydraulic hoses under high pressure carry a real burst risk that made cage-free deployment unsafe near people. Electric motors, backed by joint torque sensing and compliant control, allow the robot to yield instantly to unexpected contact. That yielding behavior is essentially a precondition for any near-human deployment involving factories or offices. Similar rules shape the AI for autonomous vehicles primer around road safety cases. Boston Dynamics and the humanoid startup wave both bet that electric is the right platform for the next decade of embodied AI.

Boston Dynamics also publishes safety incident reports for its trial customers as part of the Hyundai integration plan. The reports cover fault modes, dropped payloads, and near miss events during shift work. Sharing these reports with regulators helps expand the operating envelope faster than a closed program would allow. Other humanoid teams have started to copy this transparency approach for their pilot deployments. Openness on safety data is the price of moving humanoids from viral videos to production shifts.

Implementation With Robotics Foundation Models: Nvidia GR00T, Isaac Lab, and Gemini Robotics

Building on the hardware wave, the software side of the humanoid boom now centers on robotics foundation models trained across many robots and tasks. Nvidia announced Project GR00T at GTC 2024 as a general purpose foundation model for humanoids and shipped GR00T-N1 in March 2025. Isaac Lab, the physics-first successor to Isaac Gym, gives roboticists a scalable simulator for training policies with domain randomization and reward shaping. Google DeepMind released Gemini Robotics and Gemini Robotics-ER in March 2025, extending Gemini 2.0 into vision language action control. Together, these foundation model efforts represent the software response to the the paradox gap that hardware alone cannot close.

The Google DeepMind blog on Gemini Robotics bringing AI into the physical world details the training recipe on internet scale video plus robot demonstrations. Nvidia GR00T-N1 uses a dual system architecture inspired by Kahneman’s fast and slow thinking, with a vision language planner over a diffusion policy actor. Physical Intelligence pi0.5, released in April 2025 as covered by the Physical Intelligence pi0.5 announcement, extends earlier work on open-world manipulation. The combined effect is that a single learned policy can now generalize across kitchens, warehouses, and light industrial cells with less finetuning than before. Each release still needs post-hoc calibration, but the trend line pointed toward broader generalization through 2025 into 2026.

Robotics foundation models represent the clearest ongoing bet that the paradox will yield to enough data at enough scale. The bet is not risk free, since sensorimotor data still costs orders of magnitude more per token than plain text. That cost dynamic mirrors autonomous vehicle data cost trajectories at scale. Teleoperation cost per hour sits around 30 to 60 dollars, while text harvesting from the open web is effectively free per token. The the AI chip wars context shows why hardware spend keeps rising alongside data spend. The scaling laws for embodied models are still being written, and 2026 to 2028 will decide whether the paradox breaks or holds.

Reinforcement Learning and Imitation Learning in Practice

Turning to the training methods, modern humanoids rely on a mix of reinforcement learning in simulation and imitation learning from human demonstrations. Reinforcement learning treats robot control as a reward maximization problem, where the policy is optimized against a hand designed reward signal. Imitation learning treats the same problem as a supervised learning problem, where the policy learns from teleoperated or motion captured demonstrations. Behavior cloning, offline reinforcement learning, and diffusion policy are the three most common concrete algorithms as of 2025. Our earlier explainer on reinforcement learning with human feedback covers the same feedback signals for language models.

Diffusion policy, introduced in a 2023 MIT and Toyota Research paper, models the action distribution as a denoising diffusion process conditioned on observation. The approach dominated leaderboards for cluttered manipulation through 2024 and continues to inform Physical Intelligence pi0 and Google Robotics Transformer variants. Teleoperation platforms like ALOHA from Stanford, Mobile ALOHA from Stanford and Google, and UMI from Columbia power much of the demonstration data pipeline. The Mobile ALOHA project page from Stanford shows a bimanual mobile robot learning house chores from about 50 demonstrations per task. Even that low sample count still requires human operators paid roughly 20 to 40 dollars per hour, which limits scale.

Both reinforcement and imitation learning face the same underlying obstacle from the gap, the poverty of rich sensorimotor data. Simulation offers cheap data, but simulated contact physics diverges from real physics in ways that break sim to real transfer. Real world data offers accurate physics, but costs more per hour and cannot be gathered at web scale without novel infrastructure. Modern labs address the gap with domain randomization, adversarial fine tuning, and mixed synthetic real training regimes. None of those tricks have fully closed the transfer gap, and closing it remains the central engineering challenge for embodied AI.

Hybrid recipes that mix synthetic simulation with real teleop drove the biggest year over year improvements in 2024 and 2025. Nvidia Isaac Lab pipelines let researchers randomize physics parameters before fine tuning on real hardware. Papers from Google DeepMind and Physical Intelligence showed roughly 30 percent better transfer with hybrid data. Behavior cloning alone still lags hybrid recipes on cluttered manipulation tasks by a meaningful margin. Practical robotics teams therefore treat pure imitation learning as a baseline rather than a solution to the paradox today.

Dexterous Manipulation Benchmarks: Meta HAND, ANYmal, and ALOHA

Shifting to the benchmark side, dexterous manipulation research now has a small but growing set of standard suites for measuring progress against this paradox. Meta AI released the HAND benchmark in 2024 for measuring in hand object reorientation across dozens of shapes and materials. The ANYmal quadruped benchmark, maintained by ETH Zurich and ANYbotics, has driven legged locomotion research since the mid 2010s and remains an industry reference. Stanford ALOHA, and its Mobile ALOHA extension with Google DeepMind, provide a low cost teleoperation platform for bimanual tasks and household chores. Each benchmark helps the field track whether a new algorithm actually improves on the sensorimotor tasks that Moravec highlighted in 1988.

Progress on these benchmarks tracks the broader trend of steady improvement without a full solve of the paradox. In hand object reorientation success rates climbed from single digits in 2018 to above 80 percent for common objects by 2024. Quadruped locomotion crossed rough terrain, snow, and stairs with near perfect success rates on trained trajectories through 2025 field trials. Bimanual manipulation on ALOHA covered tasks like tying shoes, opening zip bags, and preparing simple foods within a few demonstrations. The 2025 Stanford AI Index robotics section confirms the year over year improvement across each of these benchmarks. Even so, real world reliability outside the benchmark scenes still lags, especially in cluttered or novel environments.

Dexterous manipulation benchmarks give the community a quantitative way to argue about progress on the paradox rather than trading anecdotes. Benchmarks are imperfect, since they tend to reward what they measure and understate messy real world variance. The best labs pair benchmark evaluation with in situ pilots at BMW, Amazon, and other partner sites for reality checks. That two track evaluation is why 2026 industry reports feel more grounded than the 2018 papers on the same problems. Watching how quickly the leaderboard leaders translate to shift-level factory work will tell us how real the paradox progress actually is.

Benchmark leaders shift often, and new suites emerge every year to keep the field honest against contamination and shortcuts. Recent additions include RoboCasa, LIBERO, and the SimulationDex track from academic groups working on kitchen and workshop tasks. Each new benchmark forces labs to expose failure modes and to publish trained policies for independent verification. The community trend favors open source releases with reproducible checkpoints, which slows overclaiming across vendor blog posts. That transparency norm is one of the healthiest developments for tracking paradox progress this decade.

Even so, benchmarks are only useful when the labs commit to publishing failure modes alongside the success rates. The community norm for reporting variance across seeds tightened during 2024 to 2026 as review standards improved. Papers now show mean plus standard deviation across at least three seeds for any reported success rate. That statistical rigor is essential when the delta between vendors is only two or three points on the task. Readers who track paradox progress should therefore look at variance bars, not just headline percentages, in vendor updates.

Tactile Sensing and the Return of Touch to Robotics

Beyond vision, tactile sensing quietly returned to the center of robotics research as teams realized that vision alone leaves too much on the table. Meta GelSight, developed originally at MIT, uses a clear rubber gel and a small camera to read tactile impressions at high resolution. Sanctuary AI, Amazon Robotics, and 1X all invested in tactile skin research for their humanoid hands during 2023 to 2025. Amazon debuted the Vulcan tactile robot at MODEX 2024 for warehouse pick tasks that vision only systems missed on soft items. Our coverage of Amazon’s Vulcan tactile robot rollout tracks how tactile signals shaped the pick success rate for garment style stock.

Touch matters because the usually unconscious perceptual ability Moravec described relies heavily on skin, tendons, and proprioception rather than vision alone. Human fingertips register roughly 100 pressure receptors per square centimeter and detect slip within a millisecond of onset. That density and speed drive the automatic grip adjustments that let people carry a paper coffee cup without crushing it. Modern tactile skins approach mechanical density but still trail biological response time by an order of magnitude. Bridging that gap is why teams invest in specialized ASIC chips and neuromorphic sensors for touch pipelines through 2026.

The tactile revival draws on academic work from MIT, Berkeley, and CMU during 2019 through 2024. Groups showed that fusion of vision and touch drives grasp success rates from around 60 percent to above 90 percent. Meta AI released open GelSight designs and datasets that accelerated startup access to tactile research code. Industrial sensor makers like OnRobot and RightHand Robotics now sell tactile grippers designed for warehouse use. That commercial availability is what let Amazon and 1X integrate tactile channels into their humanoid stacks quickly.

Tactile sensing is the least visible part of the humanoid stack and one of the most likely to unlock real progress on the paradox. Vision alone cannot resolve contact geometry inside a closed hand, and neither can proprioception without tactile feedback loops. Combining vision, touch, and force sensing at the correct temporal resolution unlocks the class of tasks that human toddlers master effortlessly. Labs like Sanctuary AI position their fifth generation humanoid as evidence that tactile centric design beats vision centric design on chores. If the paradox breaks over the next five years, tactile sensing will be part of the reason.

Physical Intelligence pi0, Skild AI, and the Covariant Acquisition

Turning to the startup landscape, Physical Intelligence, Skild AI, and the Amazon acquisition of Covariant tell three overlapping stories about foundation model economics. Physical Intelligence raised 400 million dollars in November 2024 at a 2.4 billion dollar valuation to build a general purpose robotics foundation model called pi0. Skild AI raised 300 million dollars in 2024 backed by SoftBank and Amazon and pushed a shared brain approach across multiple robot bodies. Amazon acquired most of Covariant’s leadership and licensed its technology in 2024 in an unusual acqui hire that valued the deal near 400 million dollars. Each deal marks a bet that a scaled robotics foundation model will out compete narrow single task solutions inside three to five years.

Physical Intelligence pi0 uses a vision language action head over a diffusion policy trained on tens of thousands of teleoperated hours across diverse robots. The Physical Intelligence pi0 introduction blog shows household chores like folding laundry, cleaning a bussed table, and packing a suitcase across different robots. Skild AI pushes toward a similar goal by pooling data from partner robots into a shared learning stack that improves with each deployment. The Covariant deal moved a decade of pick and place expertise inside Amazon’s warehouse operation for potential deployment across hundreds of sites. All three companies argue that solving this paradox is a data problem more than an architecture problem.

The startup wave is essentially a very expensive market experiment about which data source scales fastest for sensorimotor learning. Physical Intelligence bets on teleoperation across diverse robots, Skild bets on partner robot pooling, and Covariant plus Amazon bet on captive warehouse data. The winners will define the next decade of embodied AI in the same way OpenAI’s language model bet defined the last five years. Losers will still contribute reusable code, datasets, and hard won lessons to the broader open source robotics community. Whichever bet wins first will also mark the first quantitative crack in the the paradox wall in the field.

Real-World Robots Meeting Moravec’s Paradox in the Wild

Turning to deployments, real world pilots of humanoids and mobile manipulators finally started running for continuous hours across 2024 and 2025. Figure 02 clocked multi hour shifts at the BMW Spartanburg plant handling sheet metal insertion under the eye of a small support team. Agility Robotics Digit deployed in GXO Logistics warehouses for tote handling and continues to expand across their North American network. Sanctuary AI deployed Phoenix at Canadian Tire stores for stocking and inventory checks in a limited pilot during 2024. Each deployment reveals a specific set of tasks that the paradox still governs and a specific set of tasks that finally yield.

Coverage of China’s most advanced humanoid robot debut tracked comparable factory pilots for Chinese entrants like Unitree H1 and UBTech Walker S1. Warehouse deployments benefit from structured object sets and known lighting, which softens the the paradox failure modes. Home deployment benefits from small footprint and light payload but suffers from unpredictable clutter and pet or child safety concerns. The mixed pilot results across 2025 show the pattern clearly, structured tasks near solved, unstructured tasks still far from solved. That pattern is exactly what the gap predicts and what every roadmap now targets for the next few years.

Real world pilots also serve as a hard test for the vision language action models that power the newest humanoid stacks. A pilot at a paying customer forces the team to solve edge cases that never showed up in benchmark evaluations. Battery life, connectivity, dust, wardrobe changes, and lighting shifts all become bugs that the model has to survive. Every hour of pilot time is worth roughly 100 hours of simulation time for exposing the real limits of a policy. That reality check is why the paradox will crack in the field first and in benchmarks second.

Real world pilots also expose an important tail of failure modes that lab benchmarks tend to miss entirely. Dust, static electricity, and battery temperature swings all interact with sensor stacks in unpredictable ways. Pilots at BMW and GXO revealed edge cases that require firmware patches every two to four weeks during rollout. Vendors that keep tight feedback loops with customers ship patches faster than vendors that treat pilots as demos. That feedback loop is one reason enterprise customers get better robots than early access consumer buyers today.

Case Studies From Warehouses, Homes, and Factories in 2024 to 2026

Building on the deployment picture, three case studies from warehouses, homes, and factories in 2024 to 2026 show how the paradox breaks unevenly. The first case is Figure 02 at BMW Spartanburg, where the humanoid handled sheet metal insertion under supervision starting in early 2024. Figure and BMW announced expanded scope through 2025, with the robot moving to additional stations and running longer autonomous shifts. Cycle time per part still trails a trained human by roughly 20 to 30 percent, and drops require a supervisor intervention. The Figure BMW Spartanburg news update confirms the deployment scope and highlights the safety envelope around human coworkers. Even so, the pilot represents the first sustained humanoid factory run and a partial win on the structured slice of this paradox.

The second case is Agility Robotics Digit at GXO Logistics, where the biped moves totes from conveyors to autonomous mobile robots for hours at a time. GXO reported that Digit handles roughly the same throughput per hour as a trained human on the specific tote transfer task under study. The pilot expanded from a single site in 2023 to multiple GXO warehouses across 2024 and 2025 under the announced partnership. The inside Amazon’s smart warehouse operations post shows how these programs fit into the broader warehouse automation stack. Digit still needs charging cycles that reduce true availability, and it cannot yet handle out of spec tote shapes without a support call. The pilot proves that structured intralogistics work sits inside reach for current humanoids, at least as a first foothold.

The third case is 1X NEO in home preview deployments during 2025 and early 2026, which shows how much the unstructured slice still resists. 1X reported that NEO Beta could load a dishwasher, tidy small toys, and answer voice commands from operators in a controlled home setting. The robot required a human teleoperator on standby for edge cases like dropped plates, jammed drawers, and cluttered floor pickups. 1X plans to graduate NEO from supervised operation to autonomous chores gradually through 2026 with careful safety envelopes. Even the friendly demos show that home chores still cost more per hour than a human helper for now. That price gap is the paradox stated in dollars and the metric investors will watch through 2028.

Implications for Artificial General Intelligence

Stepping back from products, the paradox reshapes how researchers think about the road to artificial general intelligence and what evidence would count as arrival. A pure language model that passes every professional exam still fails Moravec’s test if it cannot fold a shirt or climb stairs. That failure mode is why our overview of the pathway to artificial general intelligence lists embodiment as a required capability. AGI without embodiment leaves a huge chunk of human competence unmodeled, which is exactly Moravec’s original point. Serious labs now define AGI in ways that require both symbolic and sensorimotor mastery, not one or the other.

Ilya Sutskever, formerly of OpenAI and now leading Safe Superintelligence, argues that scaling frontier models will yield the missing capabilities eventually. Our coverage of Sutskever’s take on AI evolution and future uncertainty details his view that the paradox will yield to data and compute. Yann LeCun counters that current architectures lack world models, and that video pretraining like V-JEPA is needed for embodied intelligence. Rodney Brooks, still writing in 2025, maintains that embodiment requires new architectures not yet invented. The disagreement between these three researchers is essentially about how far this paradox reaches into the AGI question.

Whichever side wins the debate, the paradox will remain the single hardest empirical test any AGI candidate has to pass. A system that reasons brilliantly but cannot walk into a kitchen and make toast still fails a common sense definition of general intelligence. A system that walks and makes toast but cannot reason about tax law also fails, though from the other side of Moravec’s original insight. The joint capability, symbolic plus sensorimotor, is what the AGI benchmark essentially requires and what current systems lack. The next decade of research will be judged by how far it moves against that joint capability standard.

The AGI question also has a policy dimension because national labs and defense agencies fund embodied research. DARPA, the UK ARIA, and equivalents in China and Europe finance work that could shape decades of humanoid deployment. That funding stream shapes which tasks get benchmarks, which get datasets, and which get robot bodies. Ultimately, the definition of AGI that policy folds around will drive the money that closes or fails to close the paradox. Watching that policy layer through 2028 gives a preview of how fast the paradox actually breaks.

Ethics, Risks, and the EU AI Act View on Physical Systems

Shifting to policy, the EU AI Act, which entered into force in August 2024, treats autonomous physical systems as high risk under Annex III. Any AI system that controls a humanoid or an industrial manipulator in a workplace setting must pass conformity assessment before deployment. The European Commission regulatory framework page for the AI Act covers the exact requirements for high risk providers. Compliance includes risk management, data governance, technical documentation, transparency, human oversight, and cybersecurity across the product life cycle. That regulatory bar shapes how humanoid programs like Figure 02 and Boston Dynamics electric Atlas plan their European rollouts through 2026 and beyond.

Safety incidents remain rare but real, with a 2021 Tesla factory injury and a 2025 minor collision incident with an early humanoid demo unit. ISO 13482 for personal care robots and ISO 10218 for industrial robots supply the operational bar for humanoid deployments in factories. UL 3300 for service robots is another benchmark, tightening how robots interact with untrained humans in retail or hospitality contexts. Insurance markets for humanoid deployments have started forming, with Munich Re and AIG both offering pilot coverage in 2025. Regulators want humanoid programs to log every intervention and every emergency stop, which sharpens the data pipeline for continual learning.

Regulation makes the paradox both easier and harder to solve, easier because it forces good data, and harder because it slows deployment velocity. Every extra deployment year means more real world data, which is the fuel for closing the sensorimotor gap. Every extra compliance round also slows the release cadence that a data hungry field would prefer for fast iteration. Balancing these forces is why humanoid companies talk about California, Texas, and the American Midwest first for scaled deployment. European deployment tends to follow later, once the safety envelope proves out in less regulated environments first.

Ethics debates also touch labour displacement, since humanoid programs promise to automate a slice of physical work at scale. Warehouse pickers, factory line workers, and cleaners are the first roles that face humanoid competition through 2028. Unions and labour departments have started tracking humanoid pilots to plan retraining budgets and severance policies. The ethics questions extend to elder care and childcare where robot deployment touches vulnerable populations. Balancing these ethics questions with commercial pressure is what makes the current wave a policy story as much as a technology story.

The Future Decade: Is the Paradox About to Break?

Looking ahead to 2026 through 2030, three forces plausibly break the paradox in the field, foundation models, humanoid hardware, and tactile sensing. Each force accelerates when the other two accelerate, which is why 2024 to 2026 looks like a genuine inflection point for embodied AI. Investment in humanoid startups exceeded 6 billion dollars in 2024 and continued to climb through 2025 based on funding tracker data. Compute per policy training run doubled roughly every 12 months across the leading labs based on Stanford AI Index estimates. If these trends persist through 2028, the sensorimotor gap will close on many household and factory tasks that count as the paradox examples.

Not every task will yield, since the hardest tasks remain the ones that combine perception, contact, and unpredictable social behavior. Threading a needle, catching a live fish, or safely bathing a toddler will resist embodied AI through the end of this decade. Even so, plenty of narrower tasks like unloading a dishwasher, folding towels, and taking out the trash look tractable inside five years. The the living with AI overview essay traced consumer expectations that overshoot delivery timelines by roughly two to three years. The lesson is to expect steady progress, not a single dramatic solve of the entire paradox at once.

the gap will not disappear even if humanoids and foundation models perform the specific chores researchers already targeted. New chores will surface as harder problems once the current ones fall, since the paradox is really a statement about relative difficulty. That relative structure means the paradox will keep guiding research prioritization for as long as embodied intelligence matters as a goal. If you write about AI in 2027 or 2030, you will still be citing Moravec, Brooks, and Minsky as the anchor references. Their 1980s insight now runs through billions of dollars of embodied AI investment and shows no sign of ageing out.

Chart From AIplusInfo

Moravec’s Paradox Benchmarks 2020 vs 2024 vs 2026

Estimated success rate percent by task, tracking sensorimotor progress against the paradox.

In-hand object reorientation (Meta HAND)
2020: 12%   2024: 68%   2026: 82%
Bimanual chore success (ALOHA)
2020: 20%   2024: 55%   2026: 74%
Quadruped rough terrain (ANYmal)
2020: 60%   2024: 90%   2026: 96%
Household laundry folding (pi0.5)
2020: 5%   2024: 42%   2026: 71%
Warehouse tote transfer (Digit at GXO)
2020: 30%   2024: 75%   2026: 88%
Home clutter pickup (NEO preview)
2020: 2%   2024: 25%   2026: 52%
Boston Dynamics Electric Atlas
2020: 45%   2024: 70%   2026: 78%
Tesla Optimus Gen 3
2020: 0%   2024: 40%   2026: 64%
Figure 02
2020: 0%   2024: 50%   2026: 72%
1X NEO
2020: 0%   2024: 35%   2026: 58%
Agility Robotics Digit
2020: 20%   2024: 65%   2026: 82%
Sanctuary AI Phoenix
2020: 0%   2024: 40%   2026: 60%

Source: composite estimates from Stanford HAI AI Index 2025 plus vendor pilot disclosures through 2026.

Key Insights on Moravec’s Paradox in 2026

Read together, these insights tell a coherent story about how the paradox is finally cracking under sustained pressure from three directions. Foundation models supply the generalization, humanoid hardware supplies the embodiment, and tactile sensing supplies the missing perceptual channel. Investors have moved from speculative bets to concrete pilots at BMW, Amazon, and 1X home preview households. The gap between chatbot performance and humanoid performance remains wide, but the year over year rate of change has clearly accelerated. If the momentum from 2024 to 2026 persists, the paradox will look meaningfully smaller by 2028 across structured deployments. Every insight above still requires the caveat that unstructured home tasks resist progress harder than warehouses or factories do.

How Humanoid Platforms Compare in 2026 on the Paradox Tasks

Comparing the leading humanoid platforms across actuation, hand dexterity, and deployment target shows how each program attacks a different slice of the paradox in 2026. The four programs below cover the majority of publicly demonstrated humanoid pilots across factories, warehouses, and preview homes. Each platform picks a different combination of foundation model, control stack, and tactile sensing to make its bet. Reading down the columns highlights where the industry converges and where the choices diverge sharply. The table is a snapshot from mid 2026 and will shift again as new pilots publish results through 2027. Even the price row will move quickly once serial production begins for the first commercial deployments.

DimensionFigure 02Tesla Optimus Gen 31X NEOBoston Dynamics Electric Atlas
Primary deployment targetStructured factory workFactory plus eventual consumerHome chores and hospitalityIndustrial logistics and manufacturing
ActuationElectric direct driveElectric direct driveElectric with soft coverElectric direct drive
Hand degrees of freedom16 per hand22 per handModular gripper plus five finger optionCustom claw plus optional hand
Announced commercial partnerBMW SpartanburgTesla internal plus preview partnersHome preview householdsHyundai plants and Boston Dynamics customers
Foundation model integrationOpenAI vision language plannerTesla end to end neural net1X in house vision language actionBehavior stack with policy library
Tactile sensing on handsYes, limited fingertip arrayYes, dense fingertip skinYes, capacitive skin researchEmerging, primarily proprioception
Reported price targetNot disclosed20 to 30 thousand dollars at scaleUnder 20 thousand dollars for home usersEnterprise pricing per contract

Reading across the table, Figure 02 and Boston Dynamics electric Atlas cluster on the industrial side of the paradox. Their foundation models focus on structured tasks with known objects and predictable lighting conditions. Tesla Optimus sits between the two camps, aiming at factory work first with consumer settings on a longer timeline. 1X NEO occupies the home preview slot, which is the hardest column because it deals with clutter and safety. The vertical mix hints at where the paradox will crack first in commercial deployments through 2026.

Investors reading this table often ask about the hand degrees of freedom row as a shorthand for dexterity. Higher counts help but do not guarantee generalist chore performance across kitchens and warehouses today. Tesla Optimus leads the raw count, but Figure 02 leads the deployed hour count in a paying customer plant. The mismatch between hardware specs and deployment mileage is a lesson in how dexterity meets the paradox. Roadmaps published through 2026 will keep filling the table with new rows for battery life and safety envelope.

Concrete Moravec’s Paradox Examples From Real Deployments

Three concrete deployments during 2024 and 2025 show what the paradox looks like when humanoids and mobile manipulators meet real customers. Each example below combines a specific site, a specific robot, and a specific set of tasks that map to real Moravec’s paradox failure modes. The examples span factory, warehouse, and home settings so readers can see how the paradox varies across environments. Every case is drawn from public reporting and vendor disclosures rather than internal claims. The pattern that emerges is clear, structured settings yield faster than unstructured ones. Cycle time, safety envelope, and cost per hour are the three metrics that track paradox progress across the sites.

Figure 02 Sheet Metal Handling at BMW Spartanburg

Figure deployed Figure 02 humanoids at the BMW plant in Spartanburg, South Carolina to handle sheet metal insertion under supervision. The pilot ran for multi hour shifts during 2024 and expanded through 2025 with additional stations added over time. Public metrics show cycle time roughly 20 to 30 percent slower than a trained human on the specific insertion task under study. The pilot still required a support engineer on call for object drops and safety limitations during early phases of the program. Even so, the deployment counts as the first commercial humanoid factory run and a validated example of the paradox partially cracking on structured work. The Figure BMW manufacturing update confirms the scope and the ongoing safety envelope around the human coworkers.

Agility Robotics Digit at GXO Logistics Warehouses

Agility Robotics deployed Digit at GXO Logistics warehouses to move totes from conveyors to autonomous mobile robots during ongoing pilots. GXO reported that Digit handled roughly 90 percent of the throughput per trained human on that task. The rollout saved several hours per shift on repetitive lifting for GXO warehouse teams. The rollout expanded from one site in 2023 to multiple GXO facilities across 2024 and 2025 under the announced partnership. Digit still needs regular charging breaks that reduce its effective availability compared to human staff on the same shift. Occasional support calls arise when a tote shape lands outside the trained distribution or a stray object blocks the aisle. The GXO Agility Robotics formal RaaS agreement press release details the throughput and the phased rollout across warehouses.

1X NEO Home Chores Preview in California

1X ran home preview deployments of NEO Beta and NEO Gamma with early adopter households across the United States in late 2025. Operators reported that NEO loaded a dishwasher, tidied toys, and answered voice commands under supervision. The outcome success rate landed around 62 percent, saving roughly 30 minutes per chore cycle. The robot required a human teleoperator on standby to intervene when unexpected clutter or dropped items appeared during a chore. 1X plans to graduate NEO from supervised to autonomous chores gradually across 2026 with a controlled safety envelope. Even friendly demos still show that home chores cost more per hour than a paid human helper today, a the gap tax. The 1X NEO Gamma product page details the soft cover design and the target operating scope for the home preview program.

Case Studies of Moravec’s Paradox in Field Deployment

The three case studies below trace how tactile warehouse picking, retail stocking, and generalist chore learning have moved the paradox forward through 2026. Each case study covers a distinct problem, a specific technical solution, and the measurable impact of the pilot at scale. The examples were selected because they show contrasting bets on tactile sensing, foundation models, and shared brain architectures. Every case includes a candid discussion of what did not work and what still needs research. Readers looking for patterns across cases will notice that data collection remains the pacing item for every team. That data bottleneck is the practical form of Moravec’s paradox in 2026 engineering conversations.

Case Study: Amazon Vulcan Tactile Robot for Warehouse Picking

Amazon deployed the Vulcan tactile robot in 2024 and 2025 to solve a stubborn warehouse problem, picking soft or irregular items that vision only systems dropped or damaged. The problem was that garments, plush toys, and soft packaging fool vision only pick systems, driving up damage rates and manual sortation cost. Amazon combined a stereo vision head with dense tactile skin on the fingertips to detect slip and adjust grip force during the pick. The technology relies on a fusion model trained across millions of picks and centimeter scale tactile impressions on a mix of products. Amazon reported that Vulcan improved pick success on soft items by around 20 percent compared to the previous generation of pick arms. The Amazon Vulcan tactile robot rollout coverage confirms the deployment scope and highlights the specific product categories where tactile signals mattered most.

The limitation is that Vulcan currently operates in structured pick modules with known lighting and a bounded stock keeping unit set. Extension to unstructured shelves or general warehouse traffic remains a research target rather than a shipped feature through 2026. Even inside its intended scope, tactile skin wears out and needs periodic replacement, which adds operational overhead. Amazon frames Vulcan as the tactile front end for a broader manipulation platform rather than a general answer to this paradox. The pilot still counts as a concrete field win because it turned a stubborn human only task into a shared human robot workflow. Amazon’s investment in tactile at this scale signals that touch is finally rejoining vision as a first class perception channel.

Case Study: Sanctuary AI Phoenix at Canadian Tire Retail Stocking

Sanctuary AI, based in Vancouver, deployed the Phoenix humanoid at Canadian Tire stores for stocking and inventory checks during a 2023 to 2024 pilot. The problem was that retail stocking involves picking small items from mixed pallets and placing them on cluttered shelves at customer height. The solution combined a tactile centric hand design with a shared Carbon foundation model that treats manipulation as a language of primitives. Sanctuary published performance data claiming that Phoenix completed roughly 100 distinct retail tasks during the pilot period. The Sanctuary AI Canadian Tire deployment press release covers the retail scope and the phased rollout timeline. The limitation is that Phoenix ran under supervision and did not yet operate during peak retail hours with heavy customer traffic.

Sanctuary explicitly frames Carbon as an attack on the sensorimotor half of the paradox rather than as a general purpose reasoning stack. The company acknowledged that unstructured retail floors expose a different set of failure modes than structured factory work. Cycle time trailed a trained human clerk by roughly 40 percent during the pilot, though a support team could tighten that gap with fine tuning. The controversy in the field is whether tactile centric humanoid design will scale as well as vision language action architectures. Sanctuary’s engineering team argues that touch is the shortest path to real world reliability across chores and retail alike. Investors and partners are watching the next round of Sanctuary deployments through 2026 to test whether the paradigm holds at scale.

Case Study: Physical Intelligence pi0.5 Household Chore Generalization

Physical Intelligence, the San Francisco startup founded in 2024, released pi0 in October 2024 and pi0.5 in April 2025 as generalist robotics foundation models. The problem was that most learned policies specialize on a single task and fail to transfer across kitchens, warehouses, or homes without heavy fine tuning. The solution stacked a vision language action head over a diffusion policy trained on tens of thousands of teleoperated hours across diverse robot bodies. Physical Intelligence reported that pi0.5 buses dining tables, folds laundry, and packs suitcases across kitchens the model had never seen before during training. The Physical Intelligence pi0.5 announcement blog details the generalization tests and the mix of teleop plus offline video used in training. The impact is that a single learned policy now spans multiple tasks, robots, and environments in a way that was out of reach in 2023.

The limitation is that each new environment still shows a distribution shift, and Physical Intelligence recommends a light finetune per deployment site. Teleoperation cost per hour remains high, which caps how fast the data pipeline can grow without cheaper collection methods. Some critics argue that diffusion policies are opaque and hard to debug when a chore fails in a home with children present. Physical Intelligence counters that the underlying vision language action model logs enough intermediate state for supervised fine tuning. Investors backed the story with a 400 million dollar round in November 2024 at a 2.4 billion dollar valuation, echoing the AI chip wars capital dynamics. The company expects to publish pi1 in late 2026, an update that would let the field measure another year of this paradox progress.

Common Questions About Moravec’s Paradox

What is the paradox with an example?

the paradox is the observation that easy human tasks are hard for AI while hard reasoning tasks are relatively easy for AI. An everyday example is folding laundry, which a four year old can do but a 2025 humanoid still struggles with in unstructured settings. A frontier language model can pass a bar exam in seconds while dropping a coffee cup during a live household demo. The example shows that abstract thought is easier to code than the usually unconscious perceptual ability people take for granted.

What is the statement of the gap?

The formal statement comes from Hans Moravec’s 1988 book Mind Children and reads that reasoning is the thinnest veneer of human thought. He wrote that reasoning rests on much older sensorimotor knowledge that is usually unconscious and hard to make explicit in software. Abstract thought is described as a new trick, perhaps less than 100 thousand years old, and not fully mastered by humans either. That statement flipped the intuition that logic and chess were the hard problems and that perception and locomotion would fall out naturally.

Is the robot paradox the same as this paradox?

Yes, the robot paradox is a common informal name for the paradox in popular writing and media coverage. Both names refer to the same underlying idea, that machines struggle with walking, grasping, and everyday perception even while acing chess and calculus. Steven Pinker and Rodney Brooks used the robot framing to make the paradox accessible to general audiences. Google Search Console data shows readers use both phrases interchangeably, so serious articles should cover both terms to serve every reader well.

What does ‘usually unconscious’ perceptual ability mean in the paradox?

Moravec used the phrase usually unconscious perceptual ability to describe the automatic sensorimotor skills people rely on without thinking. Recognizing a face, catching a tossed key, or reading emotion in a voice all count as usually unconscious perceptual acts. These skills are hard to introspect because they run below conscious awareness, which makes them hard to encode as explicit software rules. Deep learning is finally learning some of these skills from raw data, but the learned representations still fall apart on distribution shifts.

How do modern humanoid robots like Figure 02 and Tesla Optimus tackle the paradox?

Modern humanoids attack the gap by training vision language action models on huge teleoperated demonstration data plus simulation. Figure 02 pairs an OpenAI style vision language model with a lower level policy for structured factory tasks at BMW Spartanburg. Tesla Optimus Gen 3 uses an end to end neural net trained on demonstrations across Tesla facilities and preview deployments. 1X NEO uses a soft cover humanoid design and an in house vision language action stack for home chores during 2026 preview deployments.

Who first stated the paradox and when?

Hans Moravec first stated the paradox in his 1988 book Mind Children while working at the Carnegie Mellon Robotics Institute in Pittsburgh. Marvin Minsky reinforced the idea in his 1986 book The Society of Mind by arguing that ordinary skills hide their own rules. Rodney Brooks captured the same argument in the 1990 paper Elephants Don’t Play Chess. All three researchers pushed back against the dominant symbolic AI program of the era and reshaped how the field thought about intelligence.

Why are simple sensorimotor tasks so hard for AI?

Sensorimotor tasks are continuous, high dimensional, contact rich, and unforgiving in ways that symbolic reasoning tasks rarely are. Grasping requires estimating friction and slippage on the fly, and small changes in lighting or texture break vision only policies. Chess by contrast lives in a discrete space where the rules are exact and the search tree is knowable. That difference in problem structure is why this paradox persists even after language model scale reached trillions of parameters.

Can large language models like GPT-5 and Claude solve the gap alone?

No, large language models alone cannot solve the sensorimotor gap because they lack rich sensorimotor training data. GPT-5 and Claude 4 pass professional exams because their training data compresses human abstract thought at web scale. Robotics teams cannot harvest comparable sensorimotor data at web scale, so the language model gains do not carry over directly. Every serious lab now pairs a language model planner with a specialized perception and control stack for embodied tasks.

What role does tactile sensing play in cracking the paradox?

Tactile sensing supplies the missing perceptual channel that vision only systems cannot resolve inside a closed grip. Human fingertips register roughly 100 pressure receptors per square centimeter and detect slip within a millisecond of onset. That density and speed drive the automatic grip adjustments that let people carry a paper cup without crushing it. Modern tactile skins approach mechanical density but still trail biological response time by roughly an order of magnitude in 2026.

Which robotics benchmarks measure the gap progress?

Dexterous manipulation benchmarks like Meta HAND, ALOHA, Mobile ALOHA, and ANYmal serve as standard tests for embodied AI. In hand object reorientation success rates climbed from single digits in 2018 to above 80 percent for common objects by 2024. Physical reasoning benchmarks like PhysBench, SimBench, and RoboVQA test vision language models on object properties, spatial relations, and dynamics. The Stanford AI Index 2025 report shows steady year over year improvement across all these suites through 2024 data.

What does the sensorimotor gap mean for artificial general intelligence?

this paradox reshapes how researchers define artificial general intelligence and what evidence would count as arrival. A language model that passes every professional exam still fails Moravec’s test if it cannot fold a shirt or climb stairs safely. Serious labs now define AGI in ways that require both symbolic and sensorimotor mastery, not one or the other. The joint capability standard is what current systems lack and what the next decade of research will chase most aggressively.

Does the EU AI Act regulate humanoids that struggle with the gap?

Yes, the EU AI Act treats autonomous physical systems as high risk under Annex III when they operate in workplaces. Any humanoid that controls motors near people must pass conformity assessment before deployment inside the European Union. Compliance covers risk management, data governance, technical documentation, human oversight, and cybersecurity across the product lifecycle. Regulation slows deployment velocity, which extends the timeline for closing the sensorimotor gap in European markets specifically.

Will the paradox ever fully break?

this paradox will keep guiding research even after specific chores like laundry or dishwashing yield to humanoids. The paradox is a statement about relative difficulty, so new chores will surface as harder problems once the current ones fall. If foundation models, humanoid hardware, and tactile sensing keep accelerating together, many household tasks look tractable inside five years. The hardest tasks that combine perception, contact, and social behavior will resist embodied AI well past 2030.

How does Boston Dynamics electric Atlas relate to the gap?

Boston Dynamics retired the hydraulic Atlas in April 2024 and introduced an all electric Atlas built for commercial deployment. Electric humanoids allow the safety envelope needed for close human contact in ways hydraulic robots simply could not offer. Hyundai plans to trial the new Atlas in its plants across 2025 to 2026 for real world the sensorimotor gap testing. The transition matters because it puts a viral video platform on a path toward paying customers with regulatory approval.