AI-generated editorial illustration by China Made & Tech. It depicts no real robot, supplier, factory, worksite, safety mark, test, certificate, country, or performance result.
By China Made & Tech Team. Independent English field guide to China’s niche hardware brands, hidden champions, founders, factory towns, and supplier clusters.
China’s humanoid-robot standards story can sound simpler than it is. A ministry announces a standards-system guide. A benchmark method becomes effective. A city calls for real-world training sites. A university publishes an acceptance notice for a proof-of-concept centre. A vendor then says it is “standards-ready,” “deployment-ready,” or “ready for procurement.” The documents are real. The conclusion may still be unsupported.
The buyer’s problem is not that there are too few records. It is that the records answer different questions. A consultation notice tells you that an industry guide is still being discussed. A benchmark method tells you how a system can be assessed in a specified testing frame. A site-validation programme tells you that a user unit or third party should define conditions and issue a report for a real scenario. A public procurement notice tells you that a named transaction moved through named stages. None is a universal certificate that a particular humanoid robot will complete your task, at your site, with your people, under your commercial terms.
That is especially important for China-linked procurement. The word standard can mean a published technical document, a draft policy architecture, a product requirement, a test method, a buyer specification, a training-site condition, an acceptance checklist, or an internal operating procedure. The word deployment can mean a demonstration, a trial, an installed machine, a validated use case, a paid service, or a fleet operating repeatedly. The word acceptance can mean a contractual milestone without telling an outside reader what task was measured, how it was measured, or what happened after handover.
The useful question is therefore not, “Is this Chinese humanoid robot compliant with the new standards?” It is: which document is current, which exact system and task did it cover, who set the pass conditions, what result was recorded, and what does the purchase contract actually accept?
On 25 August 2026, China’s Ministry of Industry and Information Technology (MIIT) published the National Humanoid Robot Industry Standard System Construction Guide (2026 edition) as a draft for public comment, with the consultation open through 23 September. That is an important policy signal. It is not, by itself, a binding technical standard, a market-access approval, a supplier certificate, or a buyer’s acceptance result. MIIT’s consultation notice is explicit about the procedural status.
This article is a desk-research procurement guide, not legal advice, a safety assessment, a product test, a supplier audit, or an endorsement of any robot maker. It does not decide whether a named robot is safe, capable, economical, standards-compliant, or suitable for a particular worksite. It gives a buyer and a China-side supplier a better evidence conversation: five separate files for policy status, benchmark method, site conditions, scenario validation, and transaction acceptance.
Is China’s August 2026 humanoid-robot standards guide already a binding standard?
No. MIIT published a draft guide for public consultation, not a notice that a final binding standard had taken effect. The comment period ran from 25 August to 23 September 2026. A buyer should record the document version and status rather than treat the headline as a product certificate. The official notice is the appropriate source for that distinction.
Is there a current embodied-intelligence benchmark method in China?
Yes, but its scope matters. The national standards-information platform lists YD/T 6770-2026, Artificial Intelligence — Key foundational technology — Embodied intelligence benchmark testing methods, as effective from 1 June 2026. Its stated scope covers benchmark frameworks, metrics and methods in simulated and real environments, including environment setup, task-library construction, testing and metric calculation for one system’s perception-decision-execution chain. That describes a benchmark boundary, not a universal workcell acceptance or safety finding. Read the standard-register entry.
What does the 2026 real-world-training action ask for?
MIIT and the State-owned Assets Supervision and Administration Commission say that a user unit, or an entrusted third party, should set scenario-specific validation procedures and achievement conditions; assess real task success, efficiency improvement, safety and reliability, and economic feasibility; and issue an application-verification report. Those are programme instructions for a named scenario. They do not publish a universally achieved result for a product category or supplier. The action notice is the source.
A standards headline is a document event, not a purchase decision
The August consultation matters because it shows that China’s policy institutions are trying to organise humanoid-robot standardisation systematically. For manufacturers, component makers, systems integrators, research groups, users, insurers, and procurement teams, that direction matters. A buyer who ignores it may miss a changing vocabulary for data, testing, safety, identity, lifecycle management, or operating practices.
But direction is not the same thing as present proof. A draft guide has a different legal and operational position from an in-force standard. An in-force standard has a different position from a test report. A test report has a different position from a contract acceptance. A contract acceptance has a different position from sustained work under ordinary operating variation. Collapsing those layers is how a reasonable procurement conversation becomes an untestable sales promise.
Start with the most mundane discipline: save the document as a record, not a slogan. The file should show its exact title, issuer, publication date, status, consultation window, link, and the reason it matters for the contemplated use case. The MIIT notice says the 2026 guide is a draft and publicly seeks comments. It does not say that every product described as humanoid, embodied, intelligent, or domestically made has passed a new national test. It does not list approved models. It does not attach a universal acceptance threshold for a warehouse, factory, restaurant, hospital, laboratory, or emergency-response site.
That is not a weakness in the notice. It is the normal boundary of a consultation. A standards-system guide is a way to coordinate work across many possible topics. It can tell a market what kinds of documents, tests, terminology, and institutional structures are being considered. It cannot substitute for the document that specifies the exact task a buyer is purchasing.
The procurement error usually begins with a sentence that sounds harmless: “The robot meets China’s new standards.” Ask the seller to turn that sentence into five answers.
- Which exact standard, draft, guide, plan, or technical specification is meant?
- What is its current status—draft, published, effective, planned, local, industry, national, internal, or contractual?
- Which clause, capability, component, data field, test, or task does the seller say is relevant?
- Which exact product configuration, software version, peripheral set, and worksite does the claim cover?
- What evidence supports that statement, and what statement does the evidence not support?
If the answer remains a slide-deck phrase, the buyer has a policy signal rather than a procurement record. The appropriate next action is not to decide that the vendor is unsuitable. It is to narrow the claim. A vendor may be able to provide a meaningful benchmark report, integration document, safety design, task video, environment requirement, test protocol, or service plan. Those materials can be valuable. They should simply be filed under their own headings.
The same discipline protects China-side sellers. A supplier that says too much may accidentally take responsibility for the final integrator’s worksite, the buyer’s process design, third-party equipment, human operating procedures, or a later contract remedy. A supplier that says too little can leave a buyer unable to assess a serious proposal. The five-file approach gives both sides a way to be specific without pretending that a standards consultation eliminates ordinary systems engineering and commercial allocation.
A draft standards guide is not a pass certificate
The first file is the policy-status file. Its job is modest but essential: tell the reader what public document exists now and what procedural status it has.
In this case, MIIT’s 25 August 2026 notice announces the public consultation for the National Humanoid Robot Industry Standard System Construction Guide (2026 edition). The consultation period is dated. That date matters because readers may encounter a news article, a social-media post, a trade-show presentation, or a vendor deck after the consultation window and assume that the underlying guide was final on the day they saw it. The safe record is the notice itself, not the headline around it.
The difference between a policy architecture and a product pass certificate is more than legal vocabulary. A standards-system document can organise future work across foundational topics, systems, components, applications, safety, data, and governance. A purchase acceptance, by contrast, normally needs a concrete object: an identified configuration performing identified work at an identified location against identified conditions. One document can help shape the other. Neither replaces the other.
Think of a rail network plan. It can show a future hierarchy of corridors, but it is not a train timetable, a station-access map, or a passenger’s trip record. In the same way, a humanoid-robot standards-system guide can indicate the direction of institutional work but does not tell a buyer whether an individual robot can pick one package, conduct one inspection, carry one tray, operate one machine, or recover from one edge case. The relevant standard may not even be final or directly applicable to the proposed work.
For a buyer, the policy-status file should therefore include a short “claim translation” table. Do not use it as a legal conclusion; use it as a request list.
| Seller phrase | Evidence question | Record that can answer it |
|---|---|---|
| “Aligned with the new China standards” | Which final or effective document, and which relevant topic? | Policy-status file and technical cross-reference |
| “Prepared for future regulation” | What internal work has actually been completed? | Versioned engineering or quality record |
| “Standards-ready platform” | Ready for which test, task, site, or contract requirement? | Benchmark or site-validation file |
| “Approved for deployment” | Approved by whom, for what named scenario, and under what contract? | Application-verification and acceptance files |
The buyer should also distinguish three different clocks.
The first is the policy clock: a consultation can open, close, change, and lead to later formal work. The August guide belongs here. The second is the technical clock: standards and methods can be published, become effective, be revised, be supplemented, or be superseded. The benchmark method belongs here. The third is the transaction clock: a tender, award, contract, installation, test, acceptance, warranty, and service response occur on dates set by a specific deal. No policy clock can complete a transaction clock automatically.
This matters during fast-moving market cycles. When a new robot category attracts attention, suppliers and buyers are under pressure to compress all three clocks. A buyer wants an answer before a budget window closes. A supplier wants an order before a competitor gets a pilot. An integrator wants to promise a task before all interfaces are frozen. A policy headline can feel like permission to skip the slow work. It is not.
The more useful response is to place the consultation notice in the policy-status file and create a review trigger: revisit the file if the guide’s status, text, related formal standards, or relevant local requirements change. That is a much more defensible use of current public information than treating a draft as an all-purpose badge.
A benchmark method is not a workcell acceptance
The second file is the benchmark-method file. It asks a narrower question: what test method is relevant to the system claim, and what does that method actually measure?
The National Public Service Platform for Standards Information lists YD/T 6770-2026 as an industry standard published on 11 March 2026 and effective on 1 June 2026. Its stated scope is useful because it is concrete: it describes benchmark testing in simulated and real environments, including the test environment, task library, test process, and metric calculation for the full perception-decision-execution chain of a single embodied-intelligence system. The platform entry supports those boundaries.
That is meaningful progress for a buyer. It says that “embodied intelligence” need not remain a vague marketing category. There is a named method-oriented record that discusses how a system can be framed for assessment. It points the conversation toward test environments, tasks, procedures, and metrics rather than a choreography video or a generic intelligence claim.
It does not settle everything a worksite needs. A benchmark can be well-designed and still be incomplete for a particular purchase. Consider several questions a real buyer may have:
- Is the test environment physically comparable to the buyer’s work area?
- Is the object mix, geometry, speed, lighting, floor condition, temperature, dust, noise, connectivity, and human traffic comparable?
- Is the task one continuous sequence or a collection of selected subtasks?
- What event counts as success, partial success, abort, human intervention, safety stop, recovery, or failure?
- Was the robot in the same hardware configuration, payload range, battery state, software release, tool setup, and network mode as the quoted system?
- How many runs were conducted, over what period, and how were exceptions logged?
- Which performance measure matters to the buyer: accuracy, cycle time, throughput, intervention rate, availability, recovery time, traceability, ergonomics, energy use, or something else?
The existence of a benchmark method does not answer those questions because they belong to the test plan and the worksite. That is not an indictment of the method. It is a reminder that a method should be read with its scope visible.
The phrase “real environment” needs similar care. A record can describe a real environment without being your environment. A test in a real warehouse, workshop, clinic, retail area, or laboratory may still differ materially from the target operation. It may have another layout, another object population, different work rules, a different handoff to humans, a different shift pattern, a different acceptable intervention level, or a different consequence of a missed step. A real-world test is stronger than a pure render for some questions; it is not portable proof for every question.
This is why the benchmark-method file should contain more than a standard number. It should contain a method-to-claim map. The map connects each commercial statement to the test record needed to support it.
| Commercial statement | Method question | Evidence needed before relying on it |
|---|---|---|
| “The robot can perform task X” | What is task X, exactly? | Task definition, start/end state, permitted intervention, configuration |
| “The system is reliable” | Reliable under which period and exception rules? | Run log, failure taxonomy, recovery treatment, observation window |
| “The system is safe” | Safe under which hazard analysis and operating conditions? | Site-specific safety assessment and operating procedure |
| “The project saves labour” | Compared with which baseline and implementation cost? | Buyer-owned economic model and measured site baseline |
| “The platform passed a benchmark” | Which benchmark and which version? | Benchmark report, method reference, data scope, limitations |
The standard number itself illustrates the discipline. YD/T 6770-2026 is a stable identifier for a method record. It is not a robot serial number, a product passport, a procurement lot number, a site identifier, a maintenance agreement, or a customer reference. Those may need to be linked in a particular project, but they should never be silently substituted for one another.
The real test belongs to a named site
The third file is the site-condition file, and the fourth is the scenario-validation report. They are closely related but not interchangeable. The site-condition file describes the environment into which a robot is being introduced. The report describes what happened when a named system was assessed against named conditions.
MIIT and SASAC’s 2026 real-world-training action makes this distinction unusually clear. The action is directed at regular deployment in real production and living environments. It calls for real scenario units in industrial, service, and special domains. It says that a user unit, or an entrusted third party, should develop application-validation procedures and achievement conditions based on the characteristics of the scenario. It then names the kinds of criteria to assess: real task success, efficiency improvement, safety and reliability, and economic feasibility, followed by an application-verification report. The official notice also describes a 2026 policy goal of more than 100 high-value scenarios and ten-thousand-unit-scale deployment capacity. That is a programme target, not a current scorecard for individual vendors.
The action’s useful procurement lesson is not that every buyer should copy a government programme. It is that the user site must enter the evidence chain. The user unit knows the actual task objective, the process constraints, the exception paths, the people who will interact with the system, and the cost of failure. A robot manufacturer may know the platform. An integrator may know the interfaces. A component supplier may know a sensor or actuator. None can create a credible worksite acceptance condition without the user’s operational facts.
Take an apparently simple material-handling task. “Move boxes from point A to point B” is not a procurement specification. A useful site-condition file must answer questions such as:
- What exactly is inside the box, and how much can the weight, surface, shape, or centre of gravity vary?
- Where does the box start, and what counts as a valid final placement?
- Is the route fixed, partially structured, shared with people, shared with mobile equipment, or subject to frequent obstructions?
- What happens if the barcode cannot be read, an item is damaged, a path is blocked, a human enters the area, or the robot loses a network service?
- Which other machines, conveyors, doors, safety devices, inventory systems, or human confirmations are part of the workcell?
- Which changes to the environment are permissible, who pays for them, and who owns any new data or process logic?
- What is the fallback process when the robot stops, cannot complete a task, needs charging, needs remote support, or needs a hardware replacement?
Without those answers, a buyer has a robot category and a desired use case, not a validation plan. A system can look capable in a demonstration because the start state, objects, environment, and human assistance are controlled. That is not deceptive by itself. It becomes deceptive only when those controlled conditions are hidden while the buyer is asked to infer an open-ended operational result.
Beijing’s June 2026 call for real-world-training applications provides a useful public example of what a site screen can look like. It says proposed sites should have clear demand, clear operating conditions, a high degree of standardisation, and economic feasibility. It adds that sites should be adaptable for training, testing, and verification. The Beijing notice does not certify any robot. It does show that the condition of the operating site is an explicit selection issue—not background scenery.
For a buyer, “high degree of standardisation” should prompt discussion, not false comfort. It can mean that a work process has enough repeatable structure to define a task, collect a baseline, set a test, and interpret an exception. It does not mean the buyer’s location is risk-free or that every task should be automated. A site may be standardised in one zone and chaotic in another. A task may be repeatable in daylight but differ across shifts. A successful pilot may depend on a supervisor who is not present during normal operations. The site-condition file needs to capture those boundaries.
The vendor also needs protection from an under-specified site. If a purchaser says “the robot must work in our warehouse,” that can hide unbounded expectations: unknown floor transitions, unknown aisle traffic, unknown network coverage, unknown object variation, changing SKU policies, changing staffing, or a dependency on third-party equipment. A supplier should not answer that with a universal capability promise. It should reply with a proposed site survey, interface list, task definition, exception taxonomy, and staged test protocol.
The scenario-validation report should then show what was actually assessed. It need not be a public document. In many commercial projects, it contains confidential process details. But the buyer and supplier should agree on its existence, owner, version, scope, and decision effect. It should state, at a minimum:
- the exact system configuration and software version;
- the named worksite, workcell, and task definition;
- the preconditions and allowable environment changes;
- the test window, run rules, and exception treatment;
- the pass, conditional-pass, remediation, and fail conditions;
- the data owner and the treatment of logs, video, operational data, and confidential process information;
- the observed result, limitations, open issues, and signatories; and
- the relationship between the report and any next-stage purchase, expansion, payment, warranty, or exit decision.
This list is an editorial framework, not a statutory form. Its strength is that it makes the handoffs visible. The buyer can see whether the claimed task is the tested task. The supplier can see whether a pass condition was agreed before testing. The integrator can see which interfaces are still open. The legal or procurement team can see what result, if any, triggers an obligation.
Programme targets do not become vendor results
The national real-world action includes large, attention-grabbing numbers: by the end of 2026, it aims to consolidate more than 100 high-value application scenarios and support ten-thousand-unit-scale deployment capacity. These figures are useful as a policy ambition. They suggest that officials want work to move beyond isolated demonstrations toward repeatable scenarios and the conditions for larger deployment.
They are not a shipment total, an installed-base figure, a productivity result, a safety result, or a procurement recommendation. The grammar matters. An action plan can set a target; a buyer should not rewrite that target as “China has already deployed ten thousand humanoid robots successfully,” or “a vendor participating in this ecosystem therefore has proven our task.” Neither conclusion follows from the cited notice.
This distinction is valuable for planning. A buyer can use a target to decide that a category deserves monitoring, a small discovery budget, or a structured request for information. A seller can use it to understand which kinds of evidence may become more important: real scenario data, task definition, user-unit collaboration, validation procedures, lifecycle responsibility, safety mechanisms, and operating experience. But a target should never replace an acceptance condition.
The good procurement question is: What needs to be true before our own expansion decision? It may be a threshold for task completion, intervention frequency, system recovery, traceability, human acceptance, safety operating procedure, data governance, integration stability, or unit economics. The target should be selected by the buyer’s actual process and written into the contract or validation plan. It should not be imported from a policy headline merely because it sounds authoritative.
This is also why comparisons across projects are dangerous. One pilot may have a narrow, highly structured task and a protected space. Another may involve a variable task, shared human traffic, or more costly consequences of interruption. A statement such as “robot A was deployed in a factory” tells us little until it names the task, site, system configuration, observation period, intervention rules, and business effect. A national policy target tells us even less about any one deployment, although it may tell us something important about policy direction.
Build the acceptance file before you accept the claim
The fifth file is the transaction-acceptance file. Its purpose is to connect all other documents to a particular purchase, lease, service agreement, pilot, or expansion. It must be separate because a technically interesting benchmark or a promising site report does not decide who supplies what, when it is delivered, who installs it, what payment is due, who owns data, what counts as acceptance, how a defect is handled, or how the project exits.
Henan’s public-procurement platform offers a simple illustration of why the stages should remain distinct. A notice for Zhengzhou University’s embodied-intelligence and humanoid-robot proof-of-concept centre labels an acceptance-report notice and separately lists procurement, result, contract, and acceptance stages for the named project. The public record is not a transferable performance report. It does not tell an outside buyer the task success rate, uptime, safety outcome, economic result, or suitability of any different worksite. It does show that a purchase process contains more than the moment a system is announced or awarded.
That separation is not bureaucracy for its own sake. It creates a place for commercial truth. A procurement file should identify the object of purchase: hardware, software, tooling, integration, training, service, spare parts, support, data services, remote operations, maintenance, site changes, and any third-party dependencies. It should identify the parties: manufacturer, distributor, integrator, user unit, third-party tester, insurer, maintenance provider, and relevant interface owners. It should identify the dates: delivery, installation, test, acceptance, expansion, warranty, review, and exit.
In a humanoid-robot project, the hardest issues often live between those headings. Who is responsible if a model update changes task behaviour? Who approves a change to the workcell? Who keeps and accesses operation logs? Who handles a change in the buyer’s SKU or process? Who supplies a part if a subsystem fails? Who has authority to stop operation? What happens when a pilot passes technical validation but misses the buyer’s economic case? What happens if the robot is technically available but the integration partner cannot support the site?
No general standard number answers every one of those questions. A supplier can legitimately reference relevant standards or benchmarks, but the contract still needs a specific allocation. A buyer can legitimately require evidence, but it should specify whether it needs a pre-delivery benchmark, an on-site acceptance test, a phased service-level report, a warranty response plan, or a later expansion gate. A third-party assessor can contribute, but its mandate must say what it did and did not assess.
The acceptance file should contain a claim ledger. This is not a compliance spreadsheet with a green tick beside every broad promise. It is a simple list that keeps commercial statements tied to a source and a decision.
| Claim under discussion | Supporting file | What it can establish | What it cannot establish by itself |
|---|---|---|---|
| “The policy direction is becoming more structured” | Policy-status file | A dated consultation or effective public record exists | A product pass, safety result, or buyer approval |
| “The platform can be benchmarked” | Benchmark-method file | A method boundary and relevant test design | A buyer’s exact workcell result |
| “The scenario has been validated” | Site and validation files | A named system was assessed under named conditions | A different site, task, configuration, or period |
| “The purchase has been accepted” | Transaction-acceptance file | A named contract stage occurred | A general claim about market-wide capability |
| “The system should now scale” | Expansion decision record | A buyer’s next decision and its conditions | A permanent result under all future changes |
File 1: policy status
This file starts with the August MIIT consultation notice, but it should not end there. It should record the relevant public document’s status and a review trigger. If a supplier cites a guide, standard, local policy, industry method, or internal norm, add the exact source and its status. If the seller cannot provide the source, record the claim as unverified rather than arguing about its marketing language.
The buyer’s decision is simple: does the document change what we need to request, monitor, or plan for? It may. Does it decide that we should accept a robot? It does not.
File 2: benchmark method
This file should identify the method, version, system configuration, environment boundary, task library, test process, metric calculation, and report owner. YD/T 6770-2026 gives a useful public frame because it explicitly connects benchmark testing with environments, tasks, procedures, and metrics. The buyer should ask which parts of that frame are relevant to the proposed system and which parts remain to be specified for the site.
Do not let a method title substitute for a report. “Tested under YD/T 6770-2026” is a starting question, not a conclusion. Ask for the report, its scope, system identity, configuration, data window, metrics, exclusions, deviations, and sign-off. If the supplier cannot share all raw evidence because of confidentiality, it may be able to supply a bounded summary or arrange a controlled review. The commercial solution can vary; the scope question should not disappear.
File 3: site conditions
This file belongs partly to the user unit. It should document the task, physical layout, interface inventory, objects, environment variation, human interaction, safety operating procedure, network and power assumptions, required site changes, and fallback process. Beijing’s public call is a helpful reminder that a training and verification site needs clear demand, clear operating conditions, standardisation, and economic feasibility. Those conditions are not a universal product certificate. They are an instruction to make the environment legible.
The buyer should own or co-own this document because it describes the buyer’s operation. A vendor may contribute a requirements template, but it cannot invent the actual process constraints. If the site is not clear enough to specify, it may not be clear enough to accept a robot for the intended task.
File 4: scenario-validation report
This file records the test that matters to the proposed deployment. The MIIT–SASAC action says user units or entrusted third parties should set the procedure and achievement conditions, assess task success, efficiency improvement, safety and reliability, and economic feasibility, then issue a report. A private buyer does not need to reproduce the programme word for word. The logic is still strong: define the conditions before asking for an answer.
The report should make failure legible. If a run required human intervention, did the test treat that as success, recovery, or failure? If a robot completed the task slowly, was the result commercially useful? If a safety stop occurred, what triggered it and how was the event handled? If the system worked after an environment modification, was that modification within the agreed project scope? The report should not conceal these questions behind a single pass/fail slide.
File 5: transaction acceptance
This file ties the evidence to money, responsibility, and continuation. It states the deliverable, acceptance conditions, test procedure, data and confidentiality treatment, payment milestones, warranty or service commitments, change-control process, and exit path. It should say whether a validation report is a precondition for payment, a decision input, a limited pilot result, or a prerequisite for a larger roll-out.
This is where broad “standards-ready” language becomes a real commercial statement or is discarded. If a supplier’s claim matters enough to influence purchase, it matters enough to have an agreed contractual scope. If it cannot be defined, it should not trigger a final acceptance decision.
Turn “deployment ready” into five procurement questions
The five-file framework can feel abstract until it is used in a live conversation. Here is a practical question sequence for a buyer evaluating a China-linked humanoid-robot proposal.
Question one: what are we buying?
Name the hardware, software, tools, integration work, service coverage, replacement parts, remote support, training, and site adaptation. A robotic body is often only part of the deliverable. If the seller’s proposal depends on a specific end effector, vision stack, cloud function, map, safety layer, data pipeline, or third-party machine interface, put it inside the stated system boundary.
Question two: what document does the seller mean by “standard”?
Request the exact document and its status. If the answer points to MIIT’s August draft guide, record it as a consultation-stage policy reference. If it points to YD/T 6770-2026, record it as a benchmark method and request the relevant test materials. If it points to an internal specification, label it internal. The category tells the buyer what follow-up is needed.
Question three: what exact scenario is being proposed?
Ask for the task definition, start and end states, objects, environment, people, interfaces, expected exception types, and permitted operational changes. This is the site-condition file. It should be specific enough that a third party could understand what was supposed to happen without relying on the seller’s memory.
Question four: who defined success and who reports it?
The official real-world action puts scenario procedures and achievement conditions with the user unit or entrusted third party. Apply the same principle. Agree on the test protocol before the pilot. Agree on evidence handling before the test. Agree on who signs the report before the result appears. A report created after a commercial dispute begins is much weaker than a protocol agreed before the first run.
Question five: what happens if the claim is only partly true?
This is the contractual question. A system may show promise without meeting a full expansion condition. It may pass a narrow task but not the intended range. It may work technically but require site changes that alter the economic case. The purchase agreement should allow conditional results, remediation, retest, scope reduction, staged payment, or exit. A project that has only “success” and “failure” can encourage both sides to hide useful information.
These questions are a purchasing discipline, not a demand that every small pilot use the paperwork of a national programme. The file can be lightweight for a discovery project and more formal for a production-critical operation. The principle remains: the strength of a claim should match the strength and scope of the record behind it.
Common shortcuts that create false confidence
“The guide exists, so the robot is compliant”
The MIIT notice establishes that a draft guide entered public consultation. It does not establish compliance for an individual robot. Treating it as a certificate is a category error.
“The benchmark is real, so the task is proven”
A current benchmark method is meaningful, especially when it describes task libraries, test procedures, and metric calculation. But its scope is not automatically the buyer’s workcell. Ask what was tested, where, how, with which configuration, and what the report excludes.
“The robot was demonstrated in a real setting, so it is ready for our setting”
Real environments vary. A meaningful demonstration can show potential and can justify a more structured pilot. It cannot eliminate the need for a site-condition file or an agreed validation procedure.
“The policy goal is large, so the market result is already established”
The action’s more-than-100-scenario and ten-thousand-unit-scale figures are end-2026 targets. They are not a reported installed base, supplier score, or proof of the buyer’s return.
“A project had acceptance, so any similar purchase will work”
The Zhengzhou University notice shows process stages for one named public project. The stage record itself does not transfer task performance, safety, availability, integration quality, or economics to another customer. Compare only like with like—and even then, ask for the underlying scope.
“A buyer checklist can replace engineering”
It cannot. The point of a file is not to turn a complex system into a green spreadsheet. It is to make unknowns visible early enough for the buyer, supplier, integrator, and site owner to decide what to test, what to change, and what to place in the contract.
Method and limitations
This article is desk research based on current public records from MIIT, the national standards-information platform, Beijing’s municipal economic and information authority, and Henan’s public procurement platform. It is not legal advice, a safety assessment, a standards certification, a product test, a supplier audit, a factory visit, a customer interview, or a finding about a named robot’s real-world performance.
The August 2026 guide is a public-consultation draft and can change. YD/T 6770-2026 defines a stated benchmark-method scope, not a site-specific acceptance result. The national real-world action outlines a verification process and policy targets, not a universal threshold or completed vendor scorecard. The public procurement example records stages, not transferable task outcomes. Before making a commitment, readers should consult the current original documents and obtain appropriate technical, safety, legal, procurement, and contractual advice for the exact transaction.
Frequently asked questions
Does an effective benchmark method make a robot “standards ready”?
No. An effective benchmark method can provide a defined way to frame testing. It does not, by itself, identify a product configuration, publish a worksite result, certify safety, or settle contractual acceptance. Ask for the exact method, report, configuration, task, environment, and limitation statement.
Is a public draft guide useless for procurement?
No. It is useful as a policy-status signal and a review trigger. It can inform the buyer’s request for information and help explain why records need to be versioned. It becomes misleading only when someone treats consultation status as a product approval or a final technical requirement.
What should a buyer ask for before a pilot?
Ask for a system boundary, task definition, site-condition file, proposed benchmark evidence, agreed validation procedure, achievement conditions, data and confidentiality plan, service and change-control plan, and a written statement of what the pilot can and cannot decide. Scale the detail to the project’s risk, but do not omit the task and site boundary.
Who should define the pass condition?
The user site must be involved because it owns the operational objective and the consequences of failure. The supplier and integrator should contribute technical feasibility and test design. An independent third party can help where appropriate. The important point is to agree on the rule before measuring the result, and to attach it to the named task and system configuration.
Does acceptance mean the project has proven long-term economics?
Not necessarily. Acceptance is a transaction-stage term whose meaning depends on the contract. It may confirm delivery, installation, completion of a test, or another agreed milestone. Long-term economic and operational performance may need a separate observation period and an expansion decision.