A second scanner enters the ClawHub pipeline

OpenClaw has added Tencent’s AI-Infra-Guard, or AIG, to ClawScan, the open-source command-line system used to review skills and plugins uploaded to ClawHub. The 2 October announcement means every submission now passes through both AIG and NVIDIA’s SkillSpector before it receives a final security assessment. The scanners run independently rather than blending their outputs at the start.

That design matters because agent skills can combine natural-language instructions, scripts, dependencies and network behaviour. A scanner that focuses mainly on prompt content may miss a risky package or data flow, while a code-oriented tool may overlook instruction hijacking or memory poisoning. OpenClaw says an AI judge examines both sets of findings alongside the submitted files, preserving disagreement as evidence rather than forcing the tools into one view.

What Tencent AIG contributes

Tencent describes AIG as an LLM-driven, multi-stage audit and vulnerability-review pipeline. OpenClaw says it can trace relationships between a skill’s instructions and the scripts, dependencies and data flows behind them. Its review spans nine risk categories, including instruction hijacking, memory poisoning, remote payload execution, unauthorised access, persistence and insecure dependencies.

The integration is relevant beyond one marketplace. Agent packages can ask a model to execute commands, read files or connect to external services, which gives a seemingly small bundle a broad operational reach. A review process therefore needs to examine what code can do as well as what documentation claims it will do. Multiple scanners can add coverage, although they can also generate conflicting findings that still require careful adjudication.

Benchmark results need context

OpenClaw and Tencent evaluated the combined pipeline on a fixed 556-case subset of SkillTrustBench, a public benchmark developed by Tencent and the Chinese University of Hong Kong, Shenzhen. The cases cover benign, suspicious and malicious skills across nine categories. OpenClaw reports that ClawScan matched 86.9 per cent of benchmark labels and correctly classified 98.6 per cent of malicious cases.

The malicious-case result is encouraging, but it does not mean ClawHub submissions are guaranteed safe. A benchmark is a sample, and attackers can create behaviours outside its categories or conceal risk in dependencies and remote content. The overall agreement figure also shows that classification remains imperfect. Buyers should read the security signal as one control in a layered process, not as a replacement for permissions, isolation, provenance checks and runtime monitoring.

Disagreement becomes a feedback mechanism

Tencent found that AIG and SkillSpector surfaced different risks even when both used the same AI model. OpenClaw says it will share anonymised disagreement cases and confirmed false positives with Tencent. Those examples can become regression tests and inform changes to AIG’s detection rules, while ClawScan provides a place to evaluate whether the revisions actually improve results.

This feedback loop is one of the more important elements of the announcement. Security scanners can degrade if they are tuned only to maximise a headline score, especially when benign but unusual skills resemble malicious ones. Tracking false positives alongside missed attacks helps maintain usability. Publishing the benchmark and keeping ClawScan open also gives contributors a route to compare alternative scanners, models and review prompts against the same cases.

What operators should still verify

ClawHub users should continue to inspect a skill’s requested access, scripts and dependencies before deployment. A strong review result cannot predict every runtime input or later upstream change. Organisations can reduce exposure by pinning versions, limiting filesystem and network permissions, testing in an isolated environment and monitoring the actions an agent takes after installation.

For OpenClaw, the addition of AIG makes the marketplace review more evidence-rich and diversifies the security tooling behind it. The practical test will be how the pipeline handles new attack patterns and whether its published signals help users make better decisions without overwhelming them with warnings. The announced regression process gives the project a credible method for improving, but ongoing transparency about misses and false alarms will matter as much as the launch benchmark.

Marketplace maintainers can make the signal more useful by explaining which findings are deterministic, which depend on a model judgement and which remain unresolved. A submission that passes today may still become risky if a remote dependency changes, so periodic re-scanning and alerts for material differences should accompany upload-time review. Organisations operating private registries could also run the open ClawScan tool in their own release process, preserving results with the exact skill version. That creates a verifiable trail between what was assessed and what was installed. The Tencent integration strengthens the front door; version control, sandboxing and runtime observation still protect everything that happens afterwards.

OpenClaw should publish future benchmark runs with unchanged test splits and documented scanner versions. That will make improvements comparable and help users distinguish genuine detection gains from changes in the evaluation setup. Independent reproduction would make the evidence stronger still.