Deployment & safety · Pilot toolkit
Robot pilot scorecard: decide whether to scale, change or stop
A practical scorecard for measuring task fit, interventions, safety observations, staff experience, support demand and cost inputs without inventing ROI.
Reviewed · Editorial review
A robot pilot should end with a decision, not a highlight reel. The useful options are to scale the tested workflow, change the system and test again, or stop because the fit is poor. That decision needs a baseline, representative conditions and a record of the human work around the robot.
This scorecard can be used after the robot deployment checklist. It is a measurement template, not proof that any model will deliver a particular result. Set the measures before the pilot and keep safety or acceptance gates separate from a blended score.
Write the decision before the measures
Complete this sentence:
At the end of the pilot, we will decide whether to ___ for the task ___ at the site ___ under the conditions ___.
Name the decision owner and date. State which changes can be approved inside the pilot and which require a new scope. If the real decision is only whether to run a larger trial, say so; do not present it as approval for full deployment.
Record the baseline
Describe how the task works now, using the same boundaries planned for the pilot. Record volume, operating window, people involved, exceptions and current quality or service measure. Note seasonal or unusual conditions.
The baseline is not automatically a case for automation. It is a reference for understanding change. Do not use a single unusually bad shift, an estimate nobody owns or a process different from the pilot route.
Where cost is relevant, list inputs rather than announcing ROI: staff time by role, equipment, consumables, support, space, downtime and management time. Decide who will review the financial method. Keep assumptions visible.
Define non-negotiable gates
Some results should not be averaged away by a strong score elsewhere. Define pass or stop conditions for:
- approved operating boundary and controls;
- required task or cleaning acceptance;
- data and privacy controls;
- emergency stop and recovery;
- unacceptable damage, near miss or repeated unsafe state;
- critical integration or security failure;
- lack of competent supervision or support.
If a gate fails, pause and follow the agreed review process. Do not add more points for novelty or speed to compensate.
The Health and Safety Executive’s risk-assessment guidance provides the general framework for identifying hazards and controls. A pilot scorecard does not replace competent risk assessment.
Use six scorecard sections
Score each section only against pre-agreed criteria. A simple red/amber/green rating is often more honest than a precise percentage unsupported by the data.
1. Task completion
Record attempted and accepted tasks. Define the denominator before testing. An intervention may still lead to completion, so show autonomous or normal completion separately from assisted completion.
Useful fields:
- attempts;
- accepted completions;
- assisted completions;
- failures by defined reason;
- skipped or excluded tasks;
- rework required;
- cycle or service time distribution where meaningful.
Avoid selecting only successful demonstrations. Record the whole agreed window.
2. Interventions and exceptions
Log every human action beyond normal loading, unloading or supervision. Give it a category: route blockage, perception, integration, payload, battery, user request, software, hardware or unknown.
Record duration and role, but do not treat every minute as a cash saving or cost without a reviewed method. More important is whether interventions are predictable, documented and manageable by the intended team.
3. Safety and operating control
Record control checks, stops, near misses, route deviations, unauthorised approaches and conditions outside the approved envelope. Separate an effective protective stop from a fault: stopping may show a control worked, but frequent stops can still make the workflow impractical.
Note who reviewed each event and whether the pilot resumed, changed or stopped. Do not downgrade the significance of an event to protect the business case.
4. Reliability and support demand
Track operating time, planned service, charging, faults, recovery and supplier support. State how uptime is calculated. Excluding every pause or setup period can make the figure meaningless.
Capture:
- scheduled operating window;
- time ready for the approved task;
- planned and unplanned downtime;
- restart or recovery count;
- support contacts and response time;
- configuration changes;
- unresolved defects.
Use the exact model, software and site configuration in the report so the result is not misapplied to another setup.
5. People and workflow fit
Ask the people who operate, supervise, work beside and receive the service. Use a short consistent set of questions and allow negative feedback.
Examples:
- Was the robot’s state understandable?
- Could you get human help when needed?
- Which new work did the deployment create?
- Which step became easier or harder?
- Did the route or interface exclude anyone?
- Would the proposed operating procedure work on a normal shift?
Do not turn a few comments into a quantified satisfaction claim. Report the number and roles of respondents and preserve disagreement.
6. Commercial and operational inputs
List the complete resources required to reproduce the pilot result:
- robot hire, service or purchase route;
- operator and supervisor time;
- mapping, programming and integration;
- tooling, dock and infrastructure;
- delivery, setup and removal;
- licences, connectivity and cloud services;
- consumables, maintenance and spares;
- training and support;
- internal project and change-management time.
If the pilot used free engineering time, unusual supplier attendance or a cleared route, show it. Those inputs may not exist at scale. The rent, buy or RaaS guide can help compare commercial responsibility after task fit is established.
Set red, amber and green criteria
For each section, define the thresholds before data collection.
- Green: meets the agreed condition under representative operation with manageable support.
- Amber: potentially workable, but a named change and retest are required.
- Red: fails a gate, misses the required task or creates an unacceptable dependency.
Avoid averaging colours into a single “82% successful” claim. A pilot with green novelty and red task acceptance is red for the defined decision.
Make the final decision traceable
Use one of four outcomes:
- Scale the tested workflow. State the same conditions, configuration and controls that must be preserved.
- Change and retest. Name the change, owner, evidence needed and next decision date.
- Use a different robot or commercial route. Explain which requirement caused the mismatch.
- Stop. Record why and what was learned so the same unsuitable idea is not revived without new evidence.
Attach the baseline, raw logs, configuration, participant roles, exceptions, open defects and scorecard definitions. Separate supplier claims from pilot observations. If a result will be used publicly, obtain permission and review the method and wording.
Copyable pilot scorecard
- Decision and owner:
- Task, site and operating window:
- Robot and exact configuration:
- Baseline and source:
- Non-negotiable gates:
- Task completion — R/A/G and evidence:
- Interventions — R/A/G and evidence:
- Safety/control — R/A/G and evidence:
- Reliability/support — R/A/G and evidence:
- People/workflow — R/A/G and evidence:
- Commercial inputs — R/A/G and evidence:
- Open defects and assumptions:
- Decision: scale / change / alternative / stop:
- Approver and date:
Use this record to structure a trial with an authorised supplier or integrator. Share the task, site and decision you need to make if you want help locating relevant Robotysys research; no pilot or supplier offer is implied.
Next step
Ready to choose a robot?
Compare suitable models or send us the job. We will help you narrow it down.