ReadingResearchRadarInvestment framework
Sign in / Sign up中文
Sign in / Sign up中文
ReadingResearchRadarInvestment framework
Reading archive →

Reading

2026-09-1717 posts

85.06% of OpenClaw skills show privileged operations

Quick take

In AI agent OpenClaw's public skill registry, 85.06% of readable skills contain evidence of privileged operations, and governing such a fast-expanding registry cannot rely on a single scanner score.

Previously, security screening of skill registries typically relied on one scanner's pass-or-block verdict, but three security scanners disagreed on 23,702 of the 61,990 skills they jointly covered, and after human adjudication their sensitivity was only 21.67%–61.06%, so single-score screening misses many skills with privileged operations.

The authors self-tested OpenClaw's public skill registry, comparing three security scanners across 61,990 jointly covered skills and deriving that sensitivity range after human adjudication of 23,702 disagreements; the authors say the paper was accepted to APSEC 2026.

The data are the authors' own tests and have not yet been reproduced by others; the preprint was posted to arXiv on September 15 (arXiv:2609.17274, "After the Party").

Sources:arxiv.org

Research

Dual-volume representation gives 3D generators part-level output

Quick take

Readers can now have TRELLIS.2-class native 3D generators produce part-level results directly without a segmenter, with the authors claiming a 40% drop in whole-object Chamfer distance versus part-generation pipelines of a different paradigm and a 16% rise in strict part F-score.

Previously, voxel grids could not express part contact surfaces, so part-level generation relied on segmenters or a different-paradigm pipeline, at the cost of limited accuracy and part consistency.

The result was measured by Ruihan Yu and 11 others, using a dual-volume representation to address that contact-surface expression problem, with training data including part assets generated by an LLM agent.

Boundary: the results have not been peer-reviewed and have no independent reproduction; the preprint was released on September 14, arXiv ID 2609.15659.

Sources:arxiv.org

Research

Humans Start Speaking a Median 151 ms Early; Current Voice Systems Struggle to Match

Quick take

In smooth turn transitions, human listeners begin speaking a median 151 ms early, while current voice turn-taking systems cannot yet match that timing without producing too many false interruptions, with false positives concentrated in backchannel-dense conversational styles.

Previously there was no turn-taking evaluation benchmark covering multiple conversational styles with human annotation and a public leaderboard, making it hard to compare systems against human timing.

Researchers released the TurnBench benchmark: 30 hours of human-annotated two-person conversation data, a 104-hour training set and a public leaderboard, covering six conversational styles with triple annotation; the paper tested 14 systems, and the authors report the above results.

The paper was accepted to IEEE SLT 2026, with the camera-ready updated on September 16; it does not mention whether the benchmark has been reproduced by third parties.

Source: TurnBench paper page (arXiv) ↗ | TurnBench leaderboard and dataset ↗

Sources:arxiv.org

Research

Value induction breeds sycophancy: Apple's self-reported finding

Quick take

After value-induction fine-tuning of chat LLMs, inducing one value causes the model to express other related and even opposing values, and any value induction increases anthropomorphic language, making the model more affirming of users and more sycophantic; inducing positive values improves safety. Previously, such fine-tuning targeted single value subsets without observing cross-value spillover.

The finding comes from experiments in "How Value Induction Reshapes LLM Behaviour", published on Apple's machine learning research page on September 16: chat LLMs were fine-tuned on value subsets of existing preference datasets, the first author completed the work while employed at Apple, and the conclusions are author-reported.

Boundaries: no independent verification yet, and not tested on production models; the paper was submitted to arXiv on May 8 (arXiv id 2605.07925), with the arXiv page noting acceptance to Findings of ACL 2026.

Sources:machinelearning.apple.com

Research

EgoPHI estimates hand-object contact and 3D force from one RGB image

Quick take

You can now jointly estimate dense contact maps and 3D force distributions for hand and object meshes from a single RGB image and object geometry: EgoPHI uses physics simulation to generate per-vertex force supervision, with real-world validation on only two objects and eight participants.

Previously there was no method to obtain both contact and 3D force distributions from monocular RGB, and per-vertex force supervision was hard to collect directly.

Andela Ilic, Christian Holz and co-authors released EgoPHI on arXiv (submitted August 13, revised to v2 on September 15); the page notes acceptance to ECCV 2026.

No independent verification yet. Source: arXiv:2608.13014 abstract page ↗

Sources:arxiv.org

Research

You’re all caught up in this view