LectureRechercheRadarCadre d'investissement
Connexion / Inscription
Connexion / Inscription
LectureRechercheRadarCadre d'investissement
Archives de lecture →

Lecture

2026-09-1854 publications

EfficientTDMPC claims new sample-efficiency best on HumanoidBench and DMControl

Prise rapide
2026-09-29 12:00 GMT+8

EfficientTDMPC makes three changes to TD-MPC-family model-based reinforcement learning: aggregating multi-horizon planning objectives across different rollout depths, adding a state-action value ensemble for MuZero-style methods, and penalizing uncertain return estimates with pessimistic reanalyze when generating policy targets. The authors self-report a new sample-efficiency best on HumanoidBench and DeepMind Control Suite.

Before these three changes, the TD-MPC family gave readers no way to judge whether such modifications could yield sample-efficiency gains.

The best is a self-reported first-party result with no third-party reproduction yet; it comes from the arXiv preprint by Thomas Evers and four co-authors (first submitted May 15, 2026, updated to v4 on September 17, 2026, arXiv:2605.16692).

Source: arXiv abstract page: EfficientTDMPC (v4, 2026-09-17) ↗

Sources :arxiv.org

Recherche

Self-reported RIR framework lets long-horizon LLM agents roll back, repair, and keep reflective memory

Prise rapide
2026-09-23 12:00 GMT+8

Readers can now learn a new way for long-horizon LLM agents to recover from errors: the RIR (Rollback-Induced Reflection) framework models error recovery as a rollback-boundary control problem, rolling back to a selected prior state to repair a corrupted environment while retaining reflective memory distilled from discarded trajectories, so the agent avoids repeating mistakes.

Previously, agents lacked such memory-carrying rollback: once an environment was corrupted it was hard to repair, and retries could repeat the same mistakes.

The authors self-report consistent task-performance improvements across three long-horizon benchmarks and multiple LLM backbones; the abstract gives no concrete numbers, the results are self-tested by the authors, and independent verification is still pending.

Boundary: the result covers no concrete figures or additional benchmarks; it was submitted by five authors including Yi Yu to arXiv as a preprint on September 16 (v2 updated September 17, id 2609.18304), and no independent team has yet reproduced it.

Sources :arxiv.org

Recherche

Self-reported theory claims extending Shannon coding theorems to the action level via isoteleia maps

Prise rapide
2026-10-01 12:00 GMT+8

Readers can now learn of a self-reported mathematical theory of pragmatic information: an isoteleia mapping treats different semantic paths leading to the same optimal action as equivalent, and the authors claim to have proved counterparts of Shannon's source, channel and rate-distortion theorems at the action level, aimed at task-oriented communication, networked control and embodied AI.

Previously, Shannon's source, channel and rate-distortion theorems stayed at the level of information transmission and, by the authors' account, did not cover semantic equivalence at the action level.

The theory is self-reported by Kai Niu and Ping Zhang; the text runs 151 pages, with v3 updated on September 17 and the first version submitted to arXiv on September 10 (arXiv:2609.10986).

These are the authors' theoretical claims, without peer review or independent verification, and no third-party replication yet.

Sources :arxiv.org

Recherche

CROP filters distillation tokens by counterfactual task relevance, gaining 1.92 and 2.96 points

Prise rapide
2026-09-22 12:00 GMT+8

The selective online distillation method CROP concentrates supervision on response tokens relevant to the current input's semantic content: it uses rewrite-calibrated counterfactual sensitivity to estimate each token's task relevance, and the authors self-report aggregate performance gains of 1.92 and 2.96 points over the strongest non-CROP selector across two teacher-student distillation settings.

Previously, distillation supervision covered tokens indiscriminately, without selecting by task relevance; matched controls show CROP's selected points beat random and lowest-relevance selection.

The measurement is self-reported by the authors (first party), on two teacher-student distillation settings against the strongest non-CROP selector baseline, with gains of 1.92 and 2.96 points.

Results are author self-reported with no third-party replication yet; the paper was first submitted August 13 and updated to v3 on September 16 (arXiv:2608.13387).

Source: arXiv:2608.13387 abstract page ↗

Sources :arxiv.org

Recherche

4-bit quantization largely preserves LLM personality structure

Prise rapide
Vérifié 2026-09-18 08:47 GMT+8

Model builders can now know that 4-bit quantization largely preserves the coarse-grained personality structure of large models, while 2-bit quantization breaks fine-grained prompt consistency and cross-precision consistency, and that personality decisions mainly form in the upper layers.

Previously, when choosing quantization for open-source large models, there was no layered MBTI personality analysis basis for whether personality structure changes with precision.

Yao Fu and five other authors performed a layered MBTI personality analysis on open-source large models, covering GPTQ, AWQ 4-bit and AQLM 2-bit quantization; the results are all self-reported by the authors.

The results have not been independently verified and serve only as a reference for quantization selection in personality-sensitive chat applications; preprint v2 was updated on September 15, first submitted on August 26, the page self-labels "in progress", arXiv:2608.25977.

Sources :arxiv.org

Recherche

Meta self-reports Bumblebee beating HSTU and DLRM on industrial c-NE

Prise rapide
Vérifié 2026-09-18 08:35 GMT+8

Bumblebee, a recommendation architecture from Meta's team, achieves an author-reported average c-NE of 0.7895 on industrial data with 155.2 million unique users per 1-hour timestep and 81 billion samples in total, below HSTU and DLRM.

Previously, industrial recommenders used HSTU (0.7964) and DLRM (0.8034) as baselines, with sequence modeling and feature crossing arranged sequentially, leaving the gains of interleaving untapped.

Author-reported: Bumblebee interleaves sequence modeling and feature crossing into stackable blocks, averaging c-NE 0.7895 versus HSTU's 0.7964 and DLRM's 0.8034; interleaving the same components improves NE by 0.20% over sequential arrangement, at the cost of roughly 7-10% slower training throughput.

The results have not been independently reproduced; the data comes from an arXiv preprint (2607.24804, updated to v3 on September 15).

Sources:

  • arXiv abstract page: Bumblebee (2607.24804) ↗
  • Full-text HTML v3 (with experiments and ablation data) ↗

Sources :arxiv.org

Recherche

Agent skill attack trajectory library opens 3,014 cases

Prise rapide
2026-09-23 12:00 GMT+8

You can now search 3,014 agent skill attack cases, 6,589 trajectories, 233 affected skills and 8 risk categories, drawn from the vetted, desensitised public attack trajectory library SkillAtlas.

Previously, agent skill security reports were mostly private, and existing static, dynamic and benchmark-style evaluations rarely preserved publicly verifiable evidence.

The authors self-report that 42.5% of successful attacks succeeded only after a failed first attempt, and that trajectory-based labels raise pre-execution defence accuracy to 0.770; all these figures are self-reported by the authors.

The preprint was submitted to arXiv on 11 September by Yuxin Tian, Liang Pang and three other authors (arXiv:2609.13353), has been accepted to the EMNLP 2026 REALM workshop as a non-archival paper, and has not yet been reproduced by others.

Sources :arxiv.org

Recherche

Training-free navigation framework tops 76% success

Prise rapide
Vérifié 2026-09-18 08:15 GMT+8

The training-free embodied navigation framework HarnessVLN reaches success rates of 60.8%, 53.9%, 76.0% and 59.3% on the R2R, RxR, HM3D-v2 and HM3D-OVON benchmarks, surpassing prior training-free methods — a new reference point for what training-free approaches can achieve.

Prior training-free methods required separately building perception, memory and execution pipelines, which were hard to coordinate. HarnessVLN uses a unified Agent Harness to coordinate perception, memory and execution tools, and the authors say it has been deployed on a humanoid robot.

These figures are self-reported by Yang Chen and co-authors and have not been independently reviewed.

The preprint was submitted to arXiv on September 14 and updated to v2 on the 16th; the results have not yet been independently reproduced.

Sources :arxiv.org

Recherche

AXON's dual-path gated attention beats dense baselines on six EEG tasks

Prise rapide
Vérifié 2026-09-18 07:40 GMT+8

Readers can now know that AXON exceeds dense attention baselines on average balanced accuracy across six EEG downstream tasks, under both linear probing and full fine-tuning.

Prior EEG modeling relied on dense attention, without separating a temporal path along electrodes from a spatial path across electrodes. AXON proposes adaptive anisotropic attention: attention is split into a temporal path along electrodes and a spatial path across electrodes, with a gate predicting a two-way convex combination per token.

The result is self-reported by Mahir Jain and six other authors; audio-spectrum experiments suggest transferability beyond EEG.

No independent verification yet; the preprint was submitted September 8 and updated to v3 on the 16th, arXiv:2609.08788.

Sources :arxiv.org

Recherche

DiffAdapterVLA embeds trajectory tokens into driving VLM late layers, low latency and few parameters on NAVSIM

Prise rapide
2026-09-23 12:00 GMT+8

DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, letting trajectory states co-evolve with driving conditions during the backbone's forward computation while training only lightweight adapter modules; the authors self-report high closed-loop planning quality, low latency and few trainable parameters on the NAVSIM benchmark.

Previous approaches did not let trajectory states co-evolve with driving conditions in the backbone's forward computation, making it hard to balance planning quality, latency and parameter count.

The measurement was done first-party by Changxin Lu and seven co-authors: on the NAVSIM benchmark they self-report high closed-loop planning quality, low latency and few trainable parameters.

Boundary: results are author self-reported with no third-party reproduction yet; the preprint was submitted to arXiv on September 14 (v2 updated on the 16th, arXiv:2609.15322).

Sources :arxiv.org

Recherche

CUA-Net self-reports 91.52% accuracy classifying uterine malformations on external 3D ultrasound set

Prise rapide
Vérifié 2026-09-18 07:24 GMT+8

CUA-Net can automatically classify congenital uterine malformations in 3D ultrasound without coronal plane reconstruction, self-reporting 93.88% accuracy on the internal test set and 91.52% on the external test set.

Previously this classification relied on manual reading by sonographers, requiring coronal plane reconstruction, with junior physicians performing below the model.

Self-tested by Yueyue Xu, Yuhao Huang and 12 others: built on 3D ResNet-18, self-reported as outperforming junior sonographers on all metrics and matching senior sonographers on most.

The results are not peer-reviewed or independently verified; it is an arXiv preprint (arXiv:2609.15225), submitted September 14 and revised as v2 on September 15.

Sources :arxiv.org

Recherche

Human-written narrative descriptions boost zero-shot hidden-narrative detection; few-shot examples often hurt

Prise rapide
Vérifié 2026-09-18 07:20 GMT+8

Providing large language models with human-written narrative descriptions significantly improves their zero-shot accuracy at detecting hidden narratives in social messaging; combined with majority-vote ensembling, the authors claim performance comparable to supervised systems.

The prior approach relied on auto-generated descriptions or few-shot examples, but these often reduced accuracy due to subtle framing shifts.

The measurement was run on the Dipromats and SemEval datasets, and the results are self-reported by the three authors, not independently verified.

Boundary: no third-party replication yet; the preprint was submitted September 15 and revised as v2 on September 16 (arXiv:2609.17310).

Sources :arxiv.org

Recherche
Page de lecture suivante →