LesenForschungRadarAnlageframework
Anmelden / Registrieren
Anmelden / Registrieren
LesenForschungRadarAnlageframework
Lesearchiv →

Lesen

2026-09-1854 Beiträge

Single-pass latent steering halves VLM driving inference latency

Kurzfassung
Verifiziert 2026-09-18 13:30 GMT+8

Inference latency for vision-language model end-to-end driving can be cut by about 50% versus standard two-pass CFG guidance, while instruction following and driving scores improve rather than degrade — a result self-reported by Meibo Hu and three co-authors.

Previously such models followed instructions weakly and typically relied on standard two-pass CFG guidance to compensate, at the cost of running inference twice. The authors propose Latent-Centroid Steering (LCS), which replaces this with single-pass latent-space steering that projects toward a precomputed instruction centroid.

By the authors' own measurement, on the Bench2Drive closed-loop and nuScenes open-loop benchmarks, instruction following and driving scores both beat the two-pass CFG baseline, with inference latency reduced by about 50%. The code repository is public.

The result comes from an arXiv preprint (submitted July 31, updated to v3 on September 16, page marked IROS 2026) and has not yet been reproduced by others.

Quellen:arxiv.org

Forschung

FRAUDSkill freezes weights, tunes external skills only, self-reported 73.50% Macro-F1 on TeleAntiFraud

Kurzfassung
2026-09-24 12:00 GMT+8

FRAUDSkill keeps the underlying audio language model's weights frozen and optimizes only external skill procedures, routing policies and decision rules; the authors self-report 73.50% Macro-F1 on the TeleAntiFraud benchmark, 31.96% higher than the shared frozen-model baseline, with invalid outputs down to 1.94%.

Previously, fine-tuning the model itself was required, which is costly, and the shared frozen-model baseline performed markedly worse with more invalid outputs.

The result was self-reported by Chengxian Hu and 11 co-authors — first-party data with no third-party replication yet.

Boundary: results are limited to the TeleAntiFraud benchmark; the preprint was submitted to arXiv on September 16 (updated to v2 on the 17th, id 2609.18766), with an anonymous code link provided by the authors.

Quellen:arxiv.org

Forschung

At ~50 cm, Zivid 2M+ 60 self-reported best across all tests; ZED 2i worst on phantom

Kurzfassung
Verifiziert 2026-09-18 13:12 GMT+8

Using touch-probe sampling as the reference at about 50 cm, a comparison of four commercial depth sensors imaging pig bone, pig abdomen and a silicone kidney phantom found Zivid 2M+ 60 best on all objects and metrics, with ZED 2i second on real tissue and worst on the phantom; the test covered stereo, structured light and time-of-flight technologies.

Previously there was no such cross-technology accuracy comparison of medical depth sensors referenced to touch-probe sampling, leaving buyers unable to judge how sensors differ on real tissue versus phantoms.

The result is first-party self-reported, measured at a working distance of about 50 cm with touch-probe sampling as the reference baseline, covering four commercial devices and three test objects.

Boundary: the results are not independently reproduced; the data come from an arXiv preprint (first submitted June 11, updated to v2 on September 17, accepted at CURAC 2026, arXiv:2606.13028).

Source: arXiv:2606.13028 abstract page ↗

Quellen:arxiv.org

Forschung

RECAL claims up to 105x fewer false alarms

Kurzfassung
Verifiziert 2026-09-18 13:05 GMT+8

The unsupervised framework RECAL cut average false-positive rate by roughly 105x, 4x and 41x versus the baseline reporting the lowest FPR on three DARPA E3 datasets, with F1 of 99.99%, 99.93% and 99.99%.

Prior provenance-based intrusion detection (PIDS) methods skewed toward high-frequency relations and were prone to false positives and misses; the authors say relation frequencies in CADETS differ by about 140,000x. RECAL first does relation-balanced masked graph learning, then calibrates reconstruction error against each relation's benign error distribution.

Those figures are self-reported by Lijie Zheng, Mauro Conti and colleagues, without peer review or independent reproduction, and cannot be taken as real-world performance; the preprint was submitted to arXiv on September 15 and updated to v2 on the 17th, as arXiv:2609.16462.

Quellen:arxiv.org

Forschung

Statistical mechanics explains double descent: more parameters act like stronger weight regularization, unverified

Kurzfassung
Verifiziert 2026-09-18 13:02 GMT+8

Under Congzhou M Sha's statistical mechanics account, double descent works like this: the training trajectory is a particle diffusing for finite time in the training-loss energy landscape, which produces an effective weight decay; by equipartition of energy, adding parameters at fixed training loss lowers the temperature and moves toward stationary paths, whose L2 norm can only decrease as parameters increase — which the author says is equivalent to stronger weight regularization.

Previously the phenomenon lacked this kind of unified theoretical explanation; this is a single-author, self-reported claim.

The claim has not been peer-reviewed or independently verified; the preprint was submitted to arXiv on September 16 and updated to v2 on September 17 (arXiv:2609.19076).

Source: arXiv:2609.19076 abstract page ↗

Quellen:arxiv.org

Forschung

Visuotactile representation leads baseline by 11 points on sim benchmark

Kurzfassung
Verifiziert 2026-09-18 12:40 GMT+8

With the VT-MUSE visuotactile manipulation representation framework, all simulation benchmark tasks lead the strongest baseline it evaluated by 11 percentage points, with clear gains also in real-robot experiments, as self-reported by the authors.

Prior methods mostly encoded vision and touch independently before fusing them, ignoring contact timing. This framework first jointly adapts encoders via cross-modal temporal alignment and masked-view consistency, then combines a conditional variational latent model with tactile history, feeding a lightweight Transformer policy through gated cross-attention.

These results are self-reported by the authors; both the simulation benchmark and real-robot experiments were self-tested by the authors, with no third-party reproduction mentioned. arXiv preprint, submitted August 21, updated to v2 on September 17 (arXiv:2608.21290).

Quellen:arxiv.org

Forschung

VLA adaptation cost can be diagnosed then allocated: 0.04% of parameters matches full fine-tuning on xArm-7

Kurzfassung
Verifiziert 2026-09-18 12:36 GMT+8

The adaptation cost of VLA models is structurally distributed — appearance changes concentrate in the vision encoder, instruction changes in the language backbone — so adaptation cost can be diagnosed first and then allocated by budget, matching full fine-tuning on a physical xArm-7 with 0.04% trainable parameters.

Previously, practitioners fine-tuned VLA models in full without knowing where the cost concentrates, paying unnecessary training overhead.

Authors self-report: their pipeline first uses 10 unlabeled goal observations to diagnose the cost of each region (median Spearman 0.91), then allocates rank-variable LoRA adapters by budget, matching full fine-tuning on a physical xArm-7 with 0.04% trainable parameters. Results are self-reported by Shahram Najam Syed, Jeffrey Ichnowski and two other authors.

Not yet reproduced by third parties; the preprint was submitted to arXiv on September 16 (v2 updated on the 17th, arXiv:2609.18084).

Quellen:arxiv.org

Forschung

Shaping force profiles via generated contact-sound loudness, Franka robot completes four contact tasks zero-shot

Kurzfassung
Verifiziert 2026-09-18 12:24 GMT+8

Readers can now know that actions extracted from generated video carry only kinematics and lack force information, making contact tasks prone to failure; Guanhua Ji, Nadia Figueroa and four other authors combine generated video with audio, shaping force profiles from the loudness of generated contact sounds and closing the loop on a Franka robot, filling that gap.

Previously, actions extracted from generated video carried only kinematics and no force information, so contact tasks were prone to failure.

The authors self-report that four tasks, including whiteboard wiping and peeling, succeed zero-shot while a purely kinematic baseline fails; the method can also serve as a data engine for training policies.

Boundary: results are self-reported and not yet independently reproduced; the preprint was submitted to arXiv on September 16 (v2 updated September 17, arXiv:2609.19137), with a project page at dreamingcontactsound.github.io.

Quellen:arxiv.org

Forschung

EfficientTDMPC claims new sample-efficiency best on HumanoidBench and DMControl

Kurzfassung
2026-09-29 12:00 GMT+8

EfficientTDMPC makes three changes to TD-MPC-family model-based reinforcement learning: aggregating multi-horizon planning objectives across different rollout depths, adding a state-action value ensemble for MuZero-style methods, and penalizing uncertain return estimates with pessimistic reanalyze when generating policy targets. The authors self-report a new sample-efficiency best on HumanoidBench and DeepMind Control Suite.

Before these three changes, the TD-MPC family gave readers no way to judge whether such modifications could yield sample-efficiency gains.

The best is a self-reported first-party result with no third-party reproduction yet; it comes from the arXiv preprint by Thomas Evers and four co-authors (first submitted May 15, 2026, updated to v4 on September 17, 2026, arXiv:2605.16692).

Source: arXiv abstract page: EfficientTDMPC (v4, 2026-09-17) ↗

Quellen:arxiv.org

Forschung

Self-reported RIR framework lets long-horizon LLM agents roll back, repair, and keep reflective memory

Kurzfassung
2026-09-23 12:00 GMT+8

Readers can now learn a new way for long-horizon LLM agents to recover from errors: the RIR (Rollback-Induced Reflection) framework models error recovery as a rollback-boundary control problem, rolling back to a selected prior state to repair a corrupted environment while retaining reflective memory distilled from discarded trajectories, so the agent avoids repeating mistakes.

Previously, agents lacked such memory-carrying rollback: once an environment was corrupted it was hard to repair, and retries could repeat the same mistakes.

The authors self-report consistent task-performance improvements across three long-horizon benchmarks and multiple LLM backbones; the abstract gives no concrete numbers, the results are self-tested by the authors, and independent verification is still pending.

Boundary: the result covers no concrete figures or additional benchmarks; it was submitted by five authors including Yi Yu to arXiv as a preprint on September 16 (v2 updated September 17, id 2609.18304), and no independent team has yet reproduced it.

Quellen:arxiv.org

Forschung

Self-reported theory claims extending Shannon coding theorems to the action level via isoteleia maps

Kurzfassung
2026-10-01 12:00 GMT+8

Readers can now learn of a self-reported mathematical theory of pragmatic information: an isoteleia mapping treats different semantic paths leading to the same optimal action as equivalent, and the authors claim to have proved counterparts of Shannon's source, channel and rate-distortion theorems at the action level, aimed at task-oriented communication, networked control and embodied AI.

Previously, Shannon's source, channel and rate-distortion theorems stayed at the level of information transmission and, by the authors' account, did not cover semantic equivalence at the action level.

The theory is self-reported by Kai Niu and Ping Zhang; the text runs 151 pages, with v3 updated on September 17 and the first version submitted to arXiv on September 10 (arXiv:2609.10986).

These are the authors' theoretical claims, without peer review or independent verification, and no third-party replication yet.

Quellen:arxiv.org

Forschung

CROP filters distillation tokens by counterfactual task relevance, gaining 1.92 and 2.96 points

Kurzfassung
2026-09-22 12:00 GMT+8

The selective online distillation method CROP concentrates supervision on response tokens relevant to the current input's semantic content: it uses rewrite-calibrated counterfactual sensitivity to estimate each token's task relevance, and the authors self-report aggregate performance gains of 1.92 and 2.96 points over the strongest non-CROP selector across two teacher-student distillation settings.

Previously, distillation supervision covered tokens indiscriminately, without selecting by task relevance; matched controls show CROP's selected points beat random and lowest-relevance selection.

The measurement is self-reported by the authors (first party), on two teacher-student distillation settings against the strongest non-CROP selector baseline, with gains of 1.92 and 2.96 points.

Results are author self-reported with no third-party replication yet; the paper was first submitted August 13 and updated to v3 on September 16 (arXiv:2608.13387).

Source: arXiv:2608.13387 abstract page ↗

Quellen:arxiv.org

Forschung
Nächste Leseseite →