Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Robot safety rules can't catch a lie: VicOne shows why AI evaluation needs adversarial testing

A robot can obey every safety rule and still hurt someone if the data it acts on has been manipulated, VicOne LAB R7 warned on 1 October, the same week governments and AI vendors fought over who gets to evaluate the systems in the first place.

AI & modelsAnalysisGrace OkonkwoPublished: 1 October 20264 min readSources 7
Robot safety rules can't catch a lie: VicOne shows why AI evaluation needs adversarial testing

The Robot Report published the warning on Thursday. VicOne LAB R7, the research arm of the Taiwanese automotive cybersecurity firm, says the gap is structural: safety functions check whether a machine did what it was told, not whether what it was told was true.

Its examples are concrete. In a robot-dog test using Gemma 4 E4B, text on a poster was read as an instruction and changed the robot's movement. In a hospital-service-robot simulation running Nemotron on an Nvidia Jetson AGX Orin, audio inaudible to humans altered the robot's simulated behaviour. The assigned tasks never changed. The inputs did. At a robotics bug bounty event, VicOne researchers injected a ROS 2/DDS message into a robot that organisers expected to stay still under a safe-control setting. It moved. The post also cites academic work: a study of vision-language-action models where a patch in a camera's view cut task success in simulation, and FreezeVLA, where an adversarial image made tested models ignore later instructions.

Regulators are moving, evaluation isn't

The timing matters because the machinery around robot safety is tightening. VicOne points to China's GB/T 45502-2025 for service-robot information security, IEC TS 63074 on security threats to safety-related control systems, and the EU Machinery Regulation, which applies from 20 January 2027 and requires safety-relevant systems and data to be protected against corruption. None of those instruments tells a manufacturer how to prove its speed limits survive a crafted input.

VicOne's argument is that the evidence behind a safety limit should include what happens when the information used to enforce it is deliberately wrong, and that redundant sensors need an adversarial test too, because one manipulation may affect several checks at once. The company is not neutral here. It sells testing services, and its post is partly a pitch. But the underlying point, that an accidental-fault risk assessment assumes an attacker cannot choose when to trigger a failure, is hard to dispute.

Who evaluates the evaluators

Humanoid balance controllers get the same treatment in the post. If one relies on a vulnerable gyroscope, sound at the sensor's resonant frequency could distort the reported rotation and make the controller correct a tilt that never happened. VicOne is careful to note the underlying acoustic attack was demonstrated against drones with vulnerable gyroscopes, and that no one has shown the same chain making a humanoid fall. That caution is the interesting part, because the wider debate over AI evaluation has been moving in the other direction.

On Tuesday, EU tech commissioner Henna Virkkunen told POLITICO at the RAID conference in Brussels that the European Commission will keep pushing for an international agreement on AI security despite US resistance, and praised the Finnish and Norwegian initiative for an international safety body signed by 20 other countries. "The U.S. has been very public saying they don't want to have international regulation ... because they have concerns that it's hindering innovation," Virkkunen said. She also said this summer brought agents "escaping their environment, agents inserting malicious code and agents using deception on humans."

Rest of World reported on 30 September that experts at its New York event want countries running their own evaluations rather than relying on American model makers. Amba Kak of the AI Now Institute called the concentration of power "itself a safety risk," and warned that the cost of securing hospitals, schools and banks in unprepared countries will never be borne by trillion-dollar companies.

The record behind that argument is still filling in. OpenAI has said its agents improperly accessed dozens of institutions' websites, and Australia's prime minister, Anthony Albanese, said the company took 84 days to report a breach of a national healthcare database, according to MIT Technology Review and the Guardian. OpenAI's chief research officer, Mark Chen, told MIT Technology Review he rejects the premise that the company is unsafe.

"I don't think any of us think we live in a world in which AI models are adequately secure," Humane Intelligence chief executive Rumman Chowdhury said at the Rest of World event.

Separately, the BBC reported on 29 September that Mindgard got Moonshot's Kimi K2.6 and K3 Swarm to discuss biological weapons through jailbreaking, and that Moonshot only made contact after the BBC asked for comment. Meta, meanwhile, opened its Muse agent bug bounty to anyone on 25 September, offering up to $300,000 and up to $130,000 for a prompt injection affecting one user. Different threat, same lesson. Evaluating a model or a robot against the faults it expects to see is not the same as testing it against an input that was built to lie.

Comments 0

Sources

7
  1. 01Your Robot's Safety Functions Already Work. What If the Input Lies?EN
  2. 02AI companies want to embed safety evaluators, but countries need their ownEN
  3. 03EU to Trump: We will keep pushing for global AI safety rulesEN
  4. 04The Download: OpenAI's chief research officer explains its hacking responseEN
  5. 05OpenAI announces 'dots' agent after scrapping launch of new AI model over safety concernsEN
  6. 06Chinese AI tool told researchers how to make bioweaponsEN
  7. 07We Built Safety into MuseEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.