LesenForschungRadarAnlageframework
Anmelden / Registrieren
Anmelden / Registrieren
LesenForschungRadarAnlageframework
Lesearchiv →

Lesen

2026-09-1717 Beiträge

Cross-domain association boosts human creative originality, but seven LLMs overall do not benefit

Kurzfassung
Verifiziert 2026-09-17 22:43 GMT+8

Random cross-domain sources reliably boost the creative originality of products designed by humans, but seven large language models overall did not benefit from the same cross-domain association intervention; only when the source and target were semantically distant enough did the highest-rated LLM benefit.

It was previously unclear whether both humans and machines could gain originality from cross-domain association. A five-person team including Liu, Dubova and Griffiths had human subjects and seven LLMs design products using random cross-domain sources, finding that the LLMs' ideas were overall more original than humans' yet did not benefit from the intervention. Humans tend to transfer surface features, while LLMs tend to transfer structural and functional properties. These are the authors' self-reported experimental findings.

The results are self-reported by the authors, appear in arXiv preprint 2603.19087 (first submitted March 19, revised to v3 on September 15), and have not yet been independently reproduced.

Source: arXiv:2603.19087 abstract page ↗

Quellen:arxiv.org

Forschung

Apple's Glyph lifts column-description retrieval NDCG@10 to 0.92

Kurzfassung
Verifiziert 2026-09-17 22:19 GMT+8

Column description generation and sensitivity tagging for enterprise data catalogs can now be done automatically by collaborative LLM agents: Apple self-reports that after fine-tuning a 6-layer MiniLM, same-tag retrieval NDCG@10 on an internal held-out set rose from 0.55 to 0.92.

Previously this work relied on manual effort, costly and hard to scale; Glyph's description agent retrieves the pipeline source code that generates the column on demand, and the tagging agent runs three strategies in parallel over a 275-leaf classification ontology, fused with RRF.

The measurement comes from an Apple Machine Learning Research blog post published September 16, with the paper submitted to arXiv on September 9.

What it does not cover and who has not yet reproduced it are not stated in the original.

Quellen:machinelearning.apple.com

Forschung

Visual cues boost video-planned robot navigation, CueNav reports 70% narrow-passage success

Kurzfassung
Verifiziert 2026-09-17 21:36 GMT+8

Giving a video model visual cues — a bird's-eye view and partial body view — can substantially raise robot navigation success: the authors self-report that maze navigation success nearly doubles versus a cue-free planner, narrow-passage success reaches 70%, and they demonstrate zero-shot semantic-conditioned navigation and deployment of the same planner across robot platforms. Previously, video models planned navigation without cues before an inverse dynamics model converted predicted video into actions, with notably lower success.

The method, called CueNav, comes from a TUM/MIRMI and MIT CSAIL team; all figures above are author self-reported, the abstract does not state whether tests were simulation or real hardware, results are independently unverified, and the project page shows code not yet released; the preprint was submitted to arXiv on September 15 and revised September 16 (arXiv:2609.16737).

Quellen:arxiv.org

Forschung

DM³-Nav self-reports: decentralized multi-robot semantic navigation matches or beats centralized baselines

Kurzfassung
Verifiziert 2026-09-17 21:25 GMT+8

Multiple robots can perform semantic navigation without a central coordinator or shared global state, relying only on pairwise ad hoc communication to exchange local maps and navigation intents, and the team deployed it on two robots with onboard perception in an office — a result self-reported by a Northeastern University team in DM³-Nav.

Multi-robot semantic navigation previously typically relied on a central coordinator and shared global state; the costs of centralized approaches in communication and scalability are what this work tries to avoid.

The authors self-report that on the HM3Dv0.2 and GOAT-Bench benchmarks, the method matches or exceeds centralized and shared-map baselines; the measurement was done first-party by the Northeastern University team.

The results have no independent verification; the preprint was first submitted on April 23, revised to v3 on September 16, and arXiv lists it as accepted to IROS 2026 (arXiv:2604.22014).

Quellen:arxiv.org

Forschung

Causal Mamba hits 3.32 PESQ in real time

Kurzfassung
Verifiziert 2026-09-17 20:46 GMT+8

Real-time speech enhancement can reach 3.32 PESQ under a 25 ms algorithmic latency constraint — that is the self-reported result of RT-SEMamba by Rong Chao et al. on Voicebank-DEMAND, built on causal time-frequency Mamba blocks.

Previously real-time denoising forced a trade-off between latency and quality, and an 8-layer teacher model was good but compute-heavy. The authors use progressive knowledge distillation to compress the 8-layer teacher into a 1-layer student, raising PESQ from a naive 1-layer baseline of 3.06 to 3.18, with steady-state RTF unchanged and a 2.64x speedup over the teacher. The authors say the fixed-size recurrent state makes long-duration inference cheaper in memory and bandwidth.

All of the above are the authors' self-reported benchmark results, not yet reproduced by a third party. The preprint was first submitted on August 12, revised to v2 on September 16, and the arXiv page notes acceptance at INTERSPEECH 2026 (arXiv:2608.12099).

Quellen:arxiv.org

Forschung

ACT encoder ablation's 35%→2% drop fails to reproduce

Kurzfassung
Verifiziert 2026-09-17 19:26 GMT+8

A re-run of the CVAE encoder ablation for the robot imitation learning method ACT did not reproduce the reported 35%→2% success-rate drop: as relayed by Bo Kang, the original paper claimed that removing the encoder dropped average success rate on two simulation tasks from 35% to 2%, but re-running on the original code did not reproduce the drop.

Previously readers had only the original paper's ablation conclusion, with no way to know that training duration and checkpoint selection alone can flip which policy wins, for unknown reasons.

The re-run was performed by Bo Kang on the original code (arXiv preprint, submitted September 15, revised September 16, id 2609.16745); the author also reports that the latent variable is zeroed at inference, that skipping the encoder improves training throughput, and has released code and evaluation tools.

The result is not peer-reviewed and has not been reproduced by others; it covers two simulation tasks only, not real-robot experiments.

arXiv abstract page (primary source) ↗、HTML full text v2 (primary source) ↗

Quellen:arxiv.org

Forschung

ADORE unifies global and local explanations with first/second derivatives; authors claim it beats LIME and SHAP

Kurzfassung
Verifiziert 2026-09-17 19:00 GMT+8

Readers can now obtain global feature importance and local sample contributions within a single framework: ADORE uses first- and second-order derivatives to characterize nonlinear feature interactions, combined with randomized SVD and dynamic sparsity detection, covering tabular, text, and image data, with a Python package open-sourced.

Previously, getting both global and local explanations required separate tools such as LIME and SHAP; the authors self-report that ADORE outperforms both in interaction modeling and computational efficiency, though this comparison is the authors' own experimental conclusion.

The measurement was made by the authors themselves (first-party self-report): the benchmarks for their interaction-modeling and computational-efficiency comparison were LIME and SHAP, with ADORE reported as superior.

Boundary: the comparison has not been reproduced by third parties, and coverage is limited to the tabular, text, and image data the authors tested; Lemen Chao and two co-authors submitted it to arXiv on September 15 (v2 revised September 16, id 2609.17171).

Quellen:arxiv.org

Forschung

Lightweight road segmentation hits 97.23% MaxF, 68.73 FPS on Jetson

Kurzfassung
Verifiziert 2026-09-17 18:46 GMT+8

A lightweight vision-LiDAR fusion network, LiteViLNet, self-reports 97.23±0.15% MaxF on the KITTI Road benchmark with only 14.04M parameters in the full model, and runs at 68.73 FPS under TensorRT FP16 inference on a Jetson Orin NX.

Previous lightweight road segmentation approaches often traded off between accuracy and embedded real-time performance, struggling to achieve both.

The network consists of MobileNetV3 plus a 0.12M-parameter geometric encoder; the authors self-report 97.23±0.15% MaxF on the KITTI Road benchmark, 22.18 FPS for the model alone under PyTorch FP16 inference on a Jetson Orin NX, and separately measured 68.73 FPS under TensorRT FP16. These are the authors' self-reported public benchmark results, with no independent verification yet.

The preprint LiteViLNet by Daojie Peng et al., v1 submitted on May 20, v3 revised on September 16, arXiv:2605.21007.

Quellen:arxiv.org

Forschung

EdiTikZ self-reported: 9B beats GPT-5.6-Sol in human eval

Kurzfassung
Verifiziert 2026-09-17 18:09 GMT+8

EdiTikZ, 4B/9B models built on Qwen3.5, can edit scientific TikZ figures; the authors self-report that in a human evaluation with 9 raters and 4,320 ratings, the 9B model scored above GPT-5.6-Sol and on par with Gemini-3.1-Pro.

No dedicated model for scientific TikZ figure editing existed before; this work mined 391,000 pairs of TikZ revisions from arXiv, GitHub and TeX SE and inferred 781,000 editing instructions for training.

The measurement is a self-reported human evaluation: 9 raters, 4,320 ratings in total, with the 9B model above GPT-5.6-Sol and comparable to Gemini-3.1-Pro; the results are not peer-reviewed.

Boundary: covers only scientific TikZ figure editing, with no third-party replication; the preprint was first posted September 1 and updated September 14 (arXiv:2609.01409), models and data are released, and the full code is said to be coming soon.

Quellen:arxiv.org

Forschung

85.06% of OpenClaw skills show privileged operations

Kurzfassung
Verifiziert 2026-09-17 16:33 GMT+8

In AI agent OpenClaw's public skill registry, 85.06% of readable skills contain evidence of privileged operations, and governing such a fast-expanding registry cannot rely on a single scanner score.

Previously, security screening of skill registries typically relied on one scanner's pass-or-block verdict, but three security scanners disagreed on 23,702 of the 61,990 skills they jointly covered, and after human adjudication their sensitivity was only 21.67%–61.06%, so single-score screening misses many skills with privileged operations.

The authors self-tested OpenClaw's public skill registry, comparing three security scanners across 61,990 jointly covered skills and deriving that sensitivity range after human adjudication of 23,702 disagreements; the authors say the paper was accepted to APSEC 2026.

The data are the authors' own tests and have not yet been reproduced by others; the preprint was posted to arXiv on September 15 (arXiv:2609.17274, "After the Party").

Quellen:arxiv.org

Forschung

Dual-volume representation gives 3D generators part-level output

Kurzfassung
Verifiziert 2026-09-17 15:51 GMT+8

Readers can now have TRELLIS.2-class native 3D generators produce part-level results directly without a segmenter, with the authors claiming a 40% drop in whole-object Chamfer distance versus part-generation pipelines of a different paradigm and a 16% rise in strict part F-score.

Previously, voxel grids could not express part contact surfaces, so part-level generation relied on segmenters or a different-paradigm pipeline, at the cost of limited accuracy and part consistency.

The result was measured by Ruihan Yu and 11 others, using a dual-volume representation to address that contact-surface expression problem, with training data including part assets generated by an LLM agent.

Boundary: the results have not been peer-reviewed and have no independent reproduction; the preprint was released on September 14, arXiv ID 2609.15659.

Quellen:arxiv.org

Forschung

Humans Start Speaking a Median 151 ms Early; Current Voice Systems Struggle to Match

Kurzfassung
Verifiziert 2026-09-17 15:36 GMT+8

In smooth turn transitions, human listeners begin speaking a median 151 ms early, while current voice turn-taking systems cannot yet match that timing without producing too many false interruptions, with false positives concentrated in backchannel-dense conversational styles.

Previously there was no turn-taking evaluation benchmark covering multiple conversational styles with human annotation and a public leaderboard, making it hard to compare systems against human timing.

Researchers released the TurnBench benchmark: 30 hours of human-annotated two-person conversation data, a 104-hour training set and a public leaderboard, covering six conversational styles with triple annotation; the paper tested 14 systems, and the authors report the above results.

The paper was accepted to IEEE SLT 2026, with the camera-ready updated on September 16; it does not mention whether the benchmark has been reproduced by third parties.

Source: TurnBench paper page (arXiv) ↗ | TurnBench leaderboard and dataset ↗

Quellen:arxiv.org

Forschung
Nächste Leseseite →