From Robot Demo to Real Deployment: An Industrial Robotics Evidence Checklist
A robot can look flawless for ninety seconds and still be nowhere near ready for a factory or warehouse. A polished video answers a narrow question: can the machine complete this behavior under the conditions shown? A production deployment must answer a much less forgiving set of questions. Can it keep pace for a full shift, recover when the work is imperfect, coexist with people, connect to the surrounding operation, and leave behind records that help a team understand what happened?
This distinction matters more as mobile manipulators and humanoid robots move into industrial trials. The shape of the machine is not the deciding factor. A fixed arm, a collaborative arm, a mobile case handler, and a biped all face the same operational test: the robot has to become a dependable part of a larger process. Hardware is only one layer. The deployment also includes tooling, work presentation, safety functions, software, network behavior, fleet supervision, maintenance, training, and ownership inside the customer organization.
The practical way to judge progress is not to ask whether a demonstration was impressive. Ask what evidence exists. The checklist below uses only official company reports and major official documentation. Vendor figures are identified as vendor-reported results, not treated as independent audits. That attribution is important because a credible evaluation separates an observable operational fact from a supplier’s interpretation of it.
1. Start with a production-shaped task
A demo often begins with a capability, then searches for an attractive scene. A deployment begins in the opposite direction. It identifies a bounded piece of work with a clear input, output, takt requirement, exception path, and human owner. “Handle boxes” is not a production specification. “Unload eligible cases from these trailer types onto this conveyor during these shifts” is much closer.
Look for a written task envelope. It should describe object dimensions and weights, presentation variability, lighting, floor conditions, reach limits, required placement tolerance, upstream arrival pattern, and downstream capacity. It should also say what is excluded. An honest exclusion is useful evidence. It shows that the team knows where autonomous operation ends and where another process must take over.
Figure’s official account of its Figure 02 work at BMW Group Plant Spartanburg offers a concrete example of production-shaped measurement. Figure says the use case loaded three sheet-metal parts onto a welding fixture. The company identified three critical KPIs: total cycle time, correct loading of all three parts, and human interventions. It reported an 84-second total cycle requirement, a 37-second load phase, a target above 99 percent successful placement per shift, and a goal of zero interventions per shift. Those details are more informative than a statement that the robot can perform pick and place. They connect the motion to the line’s acceptance conditions.
Before approving a trial, require the same clarity. What event starts a cycle? What counts as complete? When does the system time out? What conditions cause a safe stop? Who deals with a rejected object? If these questions have no precise answers, the project is still exploring a behavior rather than operating a process.

2. Demand shift-level evidence, not a best run
Peak speed is seductive because it fits neatly in a caption. Operations teams need distributions. They need completed cycles per hour by shift, intervention frequency, downtime by cause, restart time, and the share of incoming work that falls inside the robot’s envelope. A single successful cycle tells you nothing about drift, thermal limits, depleted batteries, dirty sensors, changing illumination, congested aisles, or awkward objects that arrive late in the shift.
The strongest public deployment reports expose duration as well as output. Figure says its BMW deployment ran ten-hour shifts from Monday through Friday and accumulated more than 1,250 runtime hours. It also reports more than 90,000 loaded parts and a contribution to production of more than 30,000 X3 vehicles. These are Figure’s own figures, but their structure is useful: runtime, schedule, task output, and connection to finished production appear together.
Agility Robotics uses a different cumulative unit. In November 2025, the company reported that Digit had moved more than 100,000 totes at GXO’s Flowery Branch facility. Agility describes the work as part of an existing operational workflow, including interaction with proprietary facility systems, rather than an isolated behavior. The milestone does not reveal every reliability statistic a buyer would need, but it is stronger evidence than a staged tote transfer because it represents repeated work in a live facility.
Boston Dynamics and DHL provide another useful pattern. In a May 2025 joint announcement published by Boston Dynamics, the companies said deployed Stretch robots had reached case unloading rates of up to 700 cases per hour. The same announcement states that commercial introduction at DHL Supply Chain began in North America in 2023 and later expanded to the United Kingdom and continental Europe. “Up to” is not an average and should never be rewritten as one. Still, a named task, customer, deployment history, and measured rate create a trail that evaluators can investigate.
For an internal gate review, ask for at least median and lower-percentile throughput, not only the fastest hour. Break downtime into robot faults, blocked upstream flow, full downstream equipment, planned breaks, network loss, and operator-caused pauses. Otherwise the team may optimize the robot while the cell remains unproductive, or blame the robot for time when no eligible work was available.
3. Count interventions and test recovery
Autonomy is not a yes-or-no label. A robot that finishes 99 cycles alone and needs a specialist for the next one has a different operating model from a robot that needs a nearby associate every ten cycles but can be reset in seconds. Both the frequency and the cost of intervention matter.
Define intervention levels before collecting results. A simple operator reset might be Level 1. Clearing an object or adjusting work presentation might be Level 2. Remote engineering support could be Level 3. A hardware replacement or safety investigation could be Level 4. Record who performed the action, how long normal flow was interrupted, and whether work in process had to be discarded or rerouted.
Recovery tests should be deliberate. Present a missing part, a shifted container, an unreachable object, a blocked path, a communications interruption, and an emergency stop followed by the approved restart procedure. Confirm that the machine reaches a safe state, communicates the reason in language an operator can act on, and resumes without corrupting task state. A deployment that only works while nothing goes wrong is a long demonstration.
Figure’s BMW account is revealing because it discusses a failure-prone subsystem rather than presenting an entirely frictionless story. Figure identifies the forearm as Figure 02’s top hardware failure point at the site and says lessons from that subsystem informed a redesign of Figure 03 wrist electronics. The company says it removed a distribution board and dynamic cabling so each wrist motor controller communicates directly with the main computer. Whether or not another project uses similar hardware, the general evidence pattern is valuable: field failure, identified cause area, design response, and a path to verifying the revised system.
4. Treat safety as an application property
A safe robot component does not automatically produce a safe application. The gripper, payload, fixtures, conveyors, traffic pattern, maintenance access, and foreseeable human behavior all affect risk. This is why a production review should reject vague statements such as “the robot is collaborative” as a complete safety case.
The official ISO page for ISO 10218-2:2025 describes requirements for industrial robot applications and robot cells across integration, commissioning, operation, maintenance, decommissioning, and disposal. It also emphasizes the integration of the robot with end effectors and other system components. The practical lesson is straightforward: safety work follows the complete application lifecycle, not just the robot purchase.
Require a documented risk assessment for the actual site and use case, with identified hazards, protective measures, validation records, residual risks, and operating procedures. Include normal production, teaching, clearing jams, cleaning, tool changes, battery or energy isolation, maintenance, and decommissioning. Verify emergency stop behavior and restart conditions at the integrated cell level. Record changes, because a new payload, faster motion, altered route, or revised gripper can invalidate an earlier assumption.
Safety evidence also includes usability. Operators need unambiguous status indications and a known escalation path. Maintenance staff need controlled access and isolation procedures. Supervisors need to know when the robot is unavailable and how work will continue. If safe recovery routinely requires a robotics engineer, scale will be constrained even when the underlying protective functions work correctly.
5. Verify integration with the whole operation
Real work crosses boundaries. A warehouse robot may receive jobs from a warehouse management or execution layer, negotiate access to a conveyor, coordinate with autonomous mobile robots, and report completion. A factory robot may need part identity, fixture state, line permission, quality results, and traceability records. Manual job entry during a demo can hide most of this burden.
Ask for an interface inventory with owners, schemas, timing assumptions, authentication, retry behavior, and behavior during partial failure. Test duplicate messages, stale jobs, unavailable downstream equipment, clock errors, and delayed acknowledgments. Make idempotency explicit so a retried command does not create a duplicate move or duplicate production record.
Middleware configuration can affect whether the robot merely works on a lab network or behaves predictably on a site network. The official ROS 2 Jazzy documentation explains that Quality of Service policies cover history, queue depth, reliability, durability, deadline, lifespan, and liveliness. It also documents events for missed deadlines, lost liveliness, and incompatible QoS. If ROS 2 is part of a system, those policies should be selected and tested per data stream rather than accepted blindly. A high-rate sensor feed may tolerate lost samples, while a task-state transition may require a different delivery policy.
Integration should also be observable. Correlate a business job ID with robot actions, cell events, errors, and final disposition. Use synchronized timestamps. Keep enough diagnostic context to reproduce failures without recording more personal or operational data than necessary. This resembles the discipline required when teams build AI agents that actually work in production: tool access is only the beginning, while state, permissions, recovery, and evaluation determine whether the system can be trusted.

6. Prove operability, maintenance, and fleet control
The buyer is not acquiring a video. The buyer is accepting a long-lived asset with software versions, replacement parts, calibration needs, batteries, wear items, security credentials, and support obligations. A serious deployment plan therefore names the people who can start, stop, inspect, reset, and maintain the system on every operating shift.
Ask what the operator sees when performance degrades. Is the alert actionable? Can first-line staff distinguish blocked work from a robot fault? What is the expected time to acknowledge, diagnose, and restore service? Which repairs happen on site, which require a field technician, and which require depot return? Are spare modules stocked near the facility? Is there a tested rollback if a software update reduces performance?
Fleet operation adds another layer. Teams need controlled rollout groups, configuration tracking, health monitoring, job allocation, audit logs, and a way to remove a unit from service without confusing the work scheduler. Agility says its Arc platform supports facility mapping, workflow definition, operational management, and troubleshooting for Digit fleets. That is an official product claim, not proof that every deployment has the same configuration. It does, however, show the categories of capability a buyer should demand and validate.
Cybersecurity belongs in the same review. Official ROS 2 security documentation says ROS 2 can use underlying DDS security capabilities for authentication, encryption, and access-control policies, with configuration files for graph participants. Enabling a feature is not the same as completing a threat model. Inventory identities and certificates, restrict privileges, define credential rotation and revocation, protect update paths, log administrative actions, and decide how the cell behaves if a cloud or remote-support connection is lost.
7. Build an acceptance gate that can say no
A pilot needs exit criteria before installation. Without them, every result can be reframed as learning and the trial can continue indefinitely. Set thresholds for safe operation, throughput, quality, intervention rate, availability, recovery time, eligible-work coverage, and support response. Specify the measurement window and excluded downtime in advance.
The gate should compare the automated process with the real alternative, including the surrounding labor and equipment. Count loading, exception handling, inspection, supervision, maintenance, floor space, integration effort, spares, and expected upgrades. Do not hide these items behind a headline cycle rate. Equally, do not ignore improvements that are central to the use case, such as reducing work in difficult trailer conditions. Boston Dynamics and DHL explicitly connect Stretch deployment with reducing physically demanding unloading in hot or cold trailers. That outcome should be measured with an agreed operational indicator rather than left as a slogan.
Use staged gates. First, validate the task envelope off line. Next, run on site without controlling production. Then operate a limited shift with an approved fallback. Extend to representative shifts and product mixes. Finally, transfer routine ownership to operations and maintenance. A technical team standing beside the machine can be appropriate during commissioning, but it should not be mistaken for the steady-state staffing model.
The decision can be “scale,” “revise and repeat,” or “stop.” A stop is not necessarily a failure. It may show that work presentation needs redesign, the current robot is mismatched to the task, or integration cost exceeds the value. Honest gates protect both the plant and the robotics supplier from an endless showcase.
8. Read public deployment claims with precision
Official examples are useful, but they are still selected by the organizations publishing them. Read the verbs carefully. “Tested,” “piloted,” “deployed,” “commercially introduced,” and “planned” describe different levels of commitment. A memorandum of understanding is not the same thing as completed installations. Boston Dynamics’ 2025 announcement says the DHL agreement paves the way for more than 1,000 additional units. The completed evidence in that release is the prior deployment history and reported unloading performance, while the additional units are forward-looking.
Likewise, cumulative item counts need context. Ask how many robots contributed, over what period, across which shifts, with what intervention rate, and against which eligible work. A large total proves repetition but does not by itself establish availability or economics. A long runtime number is stronger when paired with fault categories and output quality. A customer logo is not an acceptance test.
This careful reading is also useful when assessing software that acts in digital environments. Our review of AI browser agents similarly distinguishes an advertised capability from the controls and fit needed for actual work. In physical automation, the stakes include moving machinery, production interruption, and maintenance access, so the evidence bar must be at least as explicit.
A compact deployment evidence pack
Before calling a robotics project a production deployment, request one evidence pack that an operations leader can review without reconstructing the story from slide decks:
- Use-case specification: task boundary, eligible work, exclusions, takt, quality requirement, and fallback flow.
- Safety file: application risk assessment, protective measures, validation, procedures, training, residual risks, and change history.
- Performance report: shift-level throughput distribution, quality, availability, interventions, recovery time, and downtime causes.
- Integration map: upstream and downstream systems, interface ownership, data contracts, timing, retries, and degraded modes.
- Operations plan: roles by shift, alerts, escalation, remote support, maintenance intervals, spares, and service targets.
- Software and security record: versions, configuration, identities, permissions, update process, rollback, logs, and incident response.
- Acceptance decision: thresholds, measurement window, actual results, open risks, accountable signatories, and the next gate.
A demo proves possibility. A deployment proves repeatability inside someone else’s constraints. The most persuasive robot is therefore not the one with the most cinematic motion. It is the one whose team can show where it works, where it does not, how often people step in, what happens after a fault, how safety was validated, and who owns Monday morning.
Frequently asked questions
What is the clearest sign that a robot has moved beyond a demo?
The clearest sign is sustained operation against predefined production acceptance criteria, with recorded throughput, quality, interventions, downtime, and recovery. A named customer or a large item count helps, but neither replaces shift-level evidence and an approved operating process.
How long should an industrial robot pilot run?
There is no universal duration. It should run long enough to cover representative shifts, work mixes, operators, environmental variation, maintenance events, and credible faults. The correct endpoint is evidence coverage, not an arbitrary number of calendar days.
Does a collaborative or humanoid form make the application safe?
No. Safety depends on the integrated application, including tooling, payloads, motion, surrounding equipment, access, and foreseeable human interaction. ISO 10218-2:2025 addresses industrial robot applications and cells across integration and their lifecycle, which is why a site-specific risk assessment and validation are essential.
Which metric matters most when evaluating production readiness?
No single metric is sufficient. Throughput without quality is misleading, and availability without eligible-work coverage can hide a narrow task envelope. A useful minimum set combines output quality, cycle-time distribution, availability, intervention frequency, recovery time, and safety performance.
Sources
- Figure AI, “F.02 Contributed to the Production of 30,000 Cars at BMW”. Official company deployment report.
- Agility Robotics, “Digit Moves Over 100,000 Totes in Commercial Deployment”. Official company deployment report.
- Agility Robotics, “Safeguarding America’s Humanoid Future”. Official company material describing deployments and Arc fleet functions.
- Boston Dynamics, “DHL Group Signs MOU with Boston Dynamics for Additional 1,000-Robot Deployment”. Official joint deployment announcement.
- International Organization for Standardization, ISO 10218-2:2025. Official standard overview for industrial robot applications and cells.
- ROS 2 Jazzy documentation, “Quality of Service settings”. Official Open Robotics documentation.
- ROS 2 Humble documentation, “ROS 2 Security”. Official Open Robotics documentation.
