LeituraPesquisaRadarFramework de investimento
Entrar / Cadastrar
Entrar / Cadastrar
LeituraPesquisaRadarFramework de investimento
Arquivo de leituras →

Leitura

2026-09-1854 posts

HumanEgo self-reported: 30 minutes of human video per task trains manipulation policies at 92.5% success

Visão rápida
Verificado 2026-09-18 01:46 GMT+8

With just 30 minutes of egocentric human video per task, you can train a manipulation policy that needs no robot data at all, achieving 92.5% average success on four real tasks — self-reported by Zhi Wang and six co-authors in HumanEgo. Previously, such policies relied on robot-collected data, which is costly and hard to transfer zero-shot to new robots, cameras, and environments.

The authors' self-reported measurements show: 30 minutes of human video per task, 92.5% average success across four real tasks, 41% higher than teleoperation of equal duration; the method converts egocentric human video into entity-level hand-object representations and trains a flow-matching policy; code and dataset are public.

Boundary: results await independent replication; the preprint was first posted May 24 and updated to v3 on September 14 (arXiv:2605.24934).

Sources: arXiv preprint ↗ | Code repository ↗ | Project page ↗

Fontes:arxiv.org

Pesquisa

RL-trained LoRA rewriting policy aligns SFT data distribution, non-downstream degradation drops across three backbones

Visão rápida
Verificado 2026-09-18 01:36 GMT+8

A lightweight LoRA rewriting policy trained with reinforcement learning can rewrite SFT data, optimizing QA-style distribution alignment and semantic diversity under a task-consistency hard gate; the authors self-report that across three instruction-tuned backbones, upstream-downstream gains are roughly on par with standard SFT, and non-downstream benchmark degradation is reduced in all evaluation settings.

Previously, SFT data rewriting lacked this kind of formalization, making it hard to control rewriting quality and downstream cost at the same time.

The result is self-reported by the authors: on three instruction-tuned backbones, upstream-downstream gains are roughly on par with standard SFT, and non-downstream benchmark degradation is reduced in all evaluation settings; cross-domain reuse of the rewriting policy is only preliminary evidence.

Boundary: the results have not been independently reproduced; the work is the arXiv preprint 2602.11220, first submitted February 11 and updated to v2 on September 16. Source: arXiv abstract page ↗

Fontes:arxiv.org

Pesquisa

Shared selective persistent memory lifts agent task completion to 96%, while storing full history drops it to 71%

Visão rápida
Verificado 2026-09-18 01:23 GMT+8

An agent LLM system using shared selective persistent memory reached 96% task completion across three enterprise deployment scenarios, compared with 79% with no memory and 71% when persisting the full history — results published by Apple's machine learning research team on September 16.

The prior approaches were persisting the full history or using no memory at all; the former actually hurt task completion, and session-level reasoning traces consume substantial context.

The method keeps only four reusable context types — task specifications, data schemas, tool configurations, and output constraints — discards session-level reasoning traces, and supports cross-user sharing. The team self-reports that summary-driven data representation cuts token cost by roughly 97x versus injecting raw data. All figures are vendor self-reported.

Boundary: results are limited to three enterprise deployment scenarios and no third-party replication yet; the preprint was submitted to arXiv on July 10 (arXiv:2607.09493) and updated to v2 on September 15.

Fontes:machinelearning.apple.com

Pesquisa

One steady-state snapshot suffices to recover particle interaction kernels, no trajectories, self-reported

Visão rápida
Verificado 2026-09-18 01:08 GMT+8

A single steady-state snapshot of collective behavior can now identify an interacting particle system, with no trajectory observations at all.

Identifying such systems previously relied on trajectory data; Baoli Hao, Mauro Maggioni and Ming Zhong instead regularize this ill-posed inverse problem using empirical distributions of configurations under different unobserved initial conditions.

The authors self-report stable and accurate recovery of the interaction kernels across multiple representative steady-state and quasi-steady-state models.

Boundary: the results are self-reported by the authors and not yet independently reproduced; the work is an arXiv preprint (arXiv:2609.12004), submitted September 10 and updated to v2 on September 16.

Fontes:arxiv.org

Pesquisa

DiffAdapterVLA injects trajectory tokens into late VLM layers, self-reported low-latency closed-loop planning on NAVSIM

Visão rápida
2026-09-23 12:00 GMT+8

DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, letting trajectory states co-evolve with driving conditions depth-by-depth inside the backbone, refining trajectories recursively with lightweight layer-wise adapters and dropping the separate planner. Previously, trajectory generation in driving VLMs was decoupled from driving-condition evolution and often relied on a separate planner, adding latency and complexity.

The authors self-report high-quality, low-latency closed-loop planning with a small number of trainable parameters on the NAVSIM benchmark.

Boundary: results are self-reported with no third-party replication; the work is by a team of eight including Changxin Lu, submitted to arXiv as a preprint on September 14 (v2 updated on the 16th, arXiv:2609.15322).

Fontes:arxiv.org

Pesquisa

Surgical video QA benchmark self-reports 14,256 pairs and 14.61% gain

Visão rápida
Verificado 2026-09-18 00:59 GMT+8

Surgical video AI now has a question-answering benchmark covering five surgical task types, SurgCoTBench, which the authors self-report contains 14,256 QA pairs, with retrieval-augmented multi-agent reasoning accuracy exceeding supervised models by 14.61%.

Previously the field lacked a unified surgical video question-answering benchmark, making it hard to compare methods across five surgical tasks.

Chang Han Low et al. self-reported the above benchmark and the 14.61% accuracy gain, updating arXiv v3 on September 16, with the entry marked as citing IEEE RA-L 2026; code is open on GitHub and the dataset has been released.

No independent verification yet.

Sources:

  • arXiv abstract page (v3, revised 2026-09-16) ↗
  • SurgRAW code repository (GitHub) ↗

Fontes:arxiv.org

Pesquisa

Você está em dia nesta visualização