LectureRechercheRadarCadre d'investissement
Connexion / Inscription
Connexion / Inscription
LectureRechercheRadarCadre d'investissement
Archives de lecture →

Lecture

2026-09-1854 publications

Uncertainty-guided test-time optimization boosts depth-only 3D segmentation

Prise rapide
Vérifié 2026-09-18 14:11 GMT+8

Indoor robots can now do open-vocabulary 3D segmentation from depth alone in privacy-preserving settings where RGB is banned: UTTO uses predictive uncertainty as a signal to test-time optimize a frozen open-vocabulary 3D segmentation backbone, and the authors report consistent gains across multiple depth-only backbones on ScanNet and Matterport3D.

Previously such settings had to rely on RGB or retrain the backbone, and segmentation quality suffered when RGB was banned.

The result is self-reported by Huang and two other authors, with a privacy-recoverability analysis and a real-robot case study; no independent reproduction has yet appeared, and the paper was first submitted on July 1, 2026 and updated to v2 on September 17 (arXiv:2607.00978).

Source: arXiv:2607.00978 abstract page (v2, updated 2026-09-17) ↗

Sources :arxiv.org

Recherche

Near-domain negatives drop fraud-detection classifiers from perfect Macro-F1 to 0.65–0.68

Prise rapide
Vérifié 2026-09-18 14:00 GMT+8

TeleAntiFraud 2.0 freezes a monthly benchmark of 900 Chinese phone-call audios (600 fraud, 300 near-domain non-fraud), letting readers test how fraud-detection classifiers really handle near-domain negatives. Previously, on unrelated or ordinary negatives, classifiers reached perfect Macro-F1, masking their failure on same-context samples.

The authors self-report that in controlled text experiments, three classifiers dropped to Macro-F1 of 0.65–0.68 once near-domain negatives were swapped in; full-audio and ASR+LLM evaluations also revealed class-prior shortcuts, prediction collapse, and snapshot sensitivity. Results are author self-reported.

Boundary: results are self-reported with no third-party replication; submitted to arXiv as a preprint by Huiyuan Liu, Zhiming Ma, and twelve other authors on September 16 (v2 updated September 17, id 2609.18748).

Sources: arXiv abstract page (2609.18748) ↗, authors' data and code repository ↗

Sources :arxiv.org

Recherche

On-device model guides home wound photos, self-reported usability good

Prise rapide
Vérifié 2026-09-18 13:55 GMT+8

Elderly patients with chronic wounds can now photograph and record their wounds at home: an on-device lightweight model segments the wound and guides the capture, while doctors monitor remotely and conduct video consultations — this is what the WoundAIssist app provides.

Previously such patients had to travel to hospital for doctors to examine wounds directly. The authors self-report that a usability study involving patients and dermatologists showed good ease of use, but the abstract gives no quantitative metrics or clinical outcomes.

Vanessa Borst and six other authors updated the paper to v2 on arXiv on September 17 (first submitted June 2025), with the page noting publication in an ACM Transactions on Computing for Healthcare 2026 special issue.

Source: arXiv abstract page (v2, updated 2026-09-17) ↗

Sources :arxiv.org

Recherche

CSWAM self-reported to lift RoboTwin 2.0 Clean-to-Randomized success from 10.16% to 45.18%

Prise rapide
2026-09-22 12:00 GMT+8

Readers can now know: CSWAM, a method that self-reports lifting RoboTwin 2.0 Clean-to-Randomized transfer success from 10.16% to 45.18%, while retaining efficient action-only inference. Previously, FastWAM-style world action models lacked causal semantic conditioning on sparse observation histories, leaving cross-randomization transfer success at roughly 10%.

The measurements were self-reported by eight authors: adding a V-JEPA 2.1-based causal semantic expert to FastWAM-style world action models, conditioning action denoising on sparse observation histories, and after embodied pretraining, RoboTwin 2.0 Clean-to-Randomized success rose from 10.16% to 45.18%; on two real-robot tasks across three OOD difficulty levels, average success rose from 27.5% to 70.0%. These are author self-reported simulation and self-test results.

Boundary: results are not independently reproduced; the work was submitted to arXiv on September 16 (v2 updated September 17, arXiv:2609.18462).

Sources :arxiv.org

Recherche

Single-pass latent steering halves VLM driving inference latency

Prise rapide
Vérifié 2026-09-18 13:30 GMT+8

Inference latency for vision-language model end-to-end driving can be cut by about 50% versus standard two-pass CFG guidance, while instruction following and driving scores improve rather than degrade — a result self-reported by Meibo Hu and three co-authors.

Previously such models followed instructions weakly and typically relied on standard two-pass CFG guidance to compensate, at the cost of running inference twice. The authors propose Latent-Centroid Steering (LCS), which replaces this with single-pass latent-space steering that projects toward a precomputed instruction centroid.

By the authors' own measurement, on the Bench2Drive closed-loop and nuScenes open-loop benchmarks, instruction following and driving scores both beat the two-pass CFG baseline, with inference latency reduced by about 50%. The code repository is public.

The result comes from an arXiv preprint (submitted July 31, updated to v3 on September 16, page marked IROS 2026) and has not yet been reproduced by others.

Sources :arxiv.org

Recherche

FRAUDSkill freezes weights, tunes external skills only, self-reported 73.50% Macro-F1 on TeleAntiFraud

Prise rapide
2026-09-24 12:00 GMT+8

FRAUDSkill keeps the underlying audio language model's weights frozen and optimizes only external skill procedures, routing policies and decision rules; the authors self-report 73.50% Macro-F1 on the TeleAntiFraud benchmark, 31.96% higher than the shared frozen-model baseline, with invalid outputs down to 1.94%.

Previously, fine-tuning the model itself was required, which is costly, and the shared frozen-model baseline performed markedly worse with more invalid outputs.

The result was self-reported by Chengxian Hu and 11 co-authors — first-party data with no third-party replication yet.

Boundary: results are limited to the TeleAntiFraud benchmark; the preprint was submitted to arXiv on September 16 (updated to v2 on the 17th, id 2609.18766), with an anonymous code link provided by the authors.

Sources :arxiv.org

Recherche

At ~50 cm, Zivid 2M+ 60 self-reported best across all tests; ZED 2i worst on phantom

Prise rapide
Vérifié 2026-09-18 13:12 GMT+8

Using touch-probe sampling as the reference at about 50 cm, a comparison of four commercial depth sensors imaging pig bone, pig abdomen and a silicone kidney phantom found Zivid 2M+ 60 best on all objects and metrics, with ZED 2i second on real tissue and worst on the phantom; the test covered stereo, structured light and time-of-flight technologies.

Previously there was no such cross-technology accuracy comparison of medical depth sensors referenced to touch-probe sampling, leaving buyers unable to judge how sensors differ on real tissue versus phantoms.

The result is first-party self-reported, measured at a working distance of about 50 cm with touch-probe sampling as the reference baseline, covering four commercial devices and three test objects.

Boundary: the results are not independently reproduced; the data come from an arXiv preprint (first submitted June 11, updated to v2 on September 17, accepted at CURAC 2026, arXiv:2606.13028).

Source: arXiv:2606.13028 abstract page ↗

Sources :arxiv.org

Recherche

RECAL claims up to 105x fewer false alarms

Prise rapide
Vérifié 2026-09-18 13:05 GMT+8

The unsupervised framework RECAL cut average false-positive rate by roughly 105x, 4x and 41x versus the baseline reporting the lowest FPR on three DARPA E3 datasets, with F1 of 99.99%, 99.93% and 99.99%.

Prior provenance-based intrusion detection (PIDS) methods skewed toward high-frequency relations and were prone to false positives and misses; the authors say relation frequencies in CADETS differ by about 140,000x. RECAL first does relation-balanced masked graph learning, then calibrates reconstruction error against each relation's benign error distribution.

Those figures are self-reported by Lijie Zheng, Mauro Conti and colleagues, without peer review or independent reproduction, and cannot be taken as real-world performance; the preprint was submitted to arXiv on September 15 and updated to v2 on the 17th, as arXiv:2609.16462.

Sources :arxiv.org

Recherche

Statistical mechanics explains double descent: more parameters act like stronger weight regularization, unverified

Prise rapide
Vérifié 2026-09-18 13:02 GMT+8

Under Congzhou M Sha's statistical mechanics account, double descent works like this: the training trajectory is a particle diffusing for finite time in the training-loss energy landscape, which produces an effective weight decay; by equipartition of energy, adding parameters at fixed training loss lowers the temperature and moves toward stationary paths, whose L2 norm can only decrease as parameters increase — which the author says is equivalent to stronger weight regularization.

Previously the phenomenon lacked this kind of unified theoretical explanation; this is a single-author, self-reported claim.

The claim has not been peer-reviewed or independently verified; the preprint was submitted to arXiv on September 16 and updated to v2 on September 17 (arXiv:2609.19076).

Source: arXiv:2609.19076 abstract page ↗

Sources :arxiv.org

Recherche

Visuotactile representation leads baseline by 11 points on sim benchmark

Prise rapide
Vérifié 2026-09-18 12:40 GMT+8

With the VT-MUSE visuotactile manipulation representation framework, all simulation benchmark tasks lead the strongest baseline it evaluated by 11 percentage points, with clear gains also in real-robot experiments, as self-reported by the authors.

Prior methods mostly encoded vision and touch independently before fusing them, ignoring contact timing. This framework first jointly adapts encoders via cross-modal temporal alignment and masked-view consistency, then combines a conditional variational latent model with tactile history, feeding a lightweight Transformer policy through gated cross-attention.

These results are self-reported by the authors; both the simulation benchmark and real-robot experiments were self-tested by the authors, with no third-party reproduction mentioned. arXiv preprint, submitted August 21, updated to v2 on September 17 (arXiv:2608.21290).

Sources :arxiv.org

Recherche

VLA adaptation cost can be diagnosed then allocated: 0.04% of parameters matches full fine-tuning on xArm-7

Prise rapide
Vérifié 2026-09-18 12:36 GMT+8

The adaptation cost of VLA models is structurally distributed — appearance changes concentrate in the vision encoder, instruction changes in the language backbone — so adaptation cost can be diagnosed first and then allocated by budget, matching full fine-tuning on a physical xArm-7 with 0.04% trainable parameters.

Previously, practitioners fine-tuned VLA models in full without knowing where the cost concentrates, paying unnecessary training overhead.

Authors self-report: their pipeline first uses 10 unlabeled goal observations to diagnose the cost of each region (median Spearman 0.91), then allocates rank-variable LoRA adapters by budget, matching full fine-tuning on a physical xArm-7 with 0.04% trainable parameters. Results are self-reported by Shahram Najam Syed, Jeffrey Ichnowski and two other authors.

Not yet reproduced by third parties; the preprint was submitted to arXiv on September 16 (v2 updated on the 17th, arXiv:2609.18084).

Sources :arxiv.org

Recherche

Shaping force profiles via generated contact-sound loudness, Franka robot completes four contact tasks zero-shot

Prise rapide
Vérifié 2026-09-18 12:24 GMT+8

Readers can now know that actions extracted from generated video carry only kinematics and lack force information, making contact tasks prone to failure; Guanhua Ji, Nadia Figueroa and four other authors combine generated video with audio, shaping force profiles from the loudness of generated contact sounds and closing the loop on a Franka robot, filling that gap.

Previously, actions extracted from generated video carried only kinematics and no force information, so contact tasks were prone to failure.

The authors self-report that four tasks, including whiteboard wiping and peeling, succeed zero-shot while a purely kinematic baseline fails; the method can also serve as a data engine for training policies.

Boundary: results are self-reported and not yet independently reproduced; the preprint was submitted to arXiv on September 16 (v2 updated September 17, arXiv:2609.19137), with a project page at dreamingcontactsound.github.io.

Sources :arxiv.org

Recherche
Page de lecture suivante →