LeituraPesquisaRadarFramework de investimento
Entrar / Cadastrar
Entrar / Cadastrar
LeituraPesquisaRadarFramework de investimento
Arquivo de leituras →

Leitura

2026-09-1854 posts

Apple's DSAS dynamically scales activation steering by input

Visão rápida
Verificado 2026-09-18 23:01 GMT+8

Readers can now know that Apple Machine Learning Research's Dynamically Scaled Activation Steering (DSAS) computes scaling factors dynamically per input and layer, strengthening intervention only when harmful behavior is detected, thereby improving the Pareto frontier between toxicity mitigation and utility preservation.

Previously, activation steering was typically applied at fixed strength, making it hard to balance intervention effectiveness and model utility. DSAS decouples "when to steer" from "how to steer"; the authors say the method is independent of the specific steering method, can be stacked with existing methods, and has been applied to text-to-image diffusion models with minimal computational overhead.

These are first-party self-reported results with no independent verification yet, and the code is said to be provided on GitHub. The paper was published in TMLR in September, arXiv ID 2512.03661.

Fontes:machinelearning.apple.com

Pesquisa

Appearance descriptions still carry graded gender associations; 16 LLMs compress them

Visão rápida
Verificado 2026-09-18 16:03 GMT+8

Readers can now verify that seemingly "objective" appearance descriptions still carry structured, graded gender associations in human interpretation: the GAPA dataset built by Yingjia Wan, Lin Lin and Elisa Kreiss covers 316 common appearance attributes, with 304 US annotators providing 14,706 gender-association ratings.

Previously, substituting descriptions for gender labels was treated as neutral communication, at the cost of hiding the strength and graded structure of the associations these descriptions themselves carry.

The authors report that 16 LLMs only partially reproduce the human ratings: the distributions are compressed, alignment with male-associated items is weaker, and abstentions cluster asymmetrically on the non-binary category; the dataset, code and proxy prediction model have been released by the authors.

The result does not cover non-US annotator populations and has not yet been independently reproduced; the preprint is arXiv 2609.16366, with the arXiv page noting publication at COLM 2026.

Fontes:arxiv.org

Pesquisa

Self-evolving agents' skill pools pollute past a critical size; VaG gating claims 72% pass@1 on Terminal-Bench 2

Visão rápida
Verificado 2026-09-18 15:48 GMT+8

When self-evolving agents distill skills from execution trajectories, new skills start to degrade performance once the skill pool passes a critical size — the core finding self-reported by Linfang Shang and six co-authors, whose Verifier-as-Gatekeeper (VaG) method claims a round-by-round rise to 72% pass@1 on Terminal-Bench 2.

Under unconditional skill accumulation, defective skills enter the decision context and become references for later distillation, forming a cross-round contamination chain; deleting the source skill afterwards recovers only a small fraction of the loss.

The authors therefore propose Verifier-as-Gatekeeper (VaG) gating, which filters skills one by one with three types of reviewers plus marginal-gain screening, and reports a skill pool about one-fifth the size of unconditional accumulation. These are first-party results.

Boundary: the results have not been reproduced by third parties; the preprint was submitted to arXiv on August 6 and updated as v2 on September 17 (2608.05810).

Fontes:arxiv.org

Pesquisa

MoRE shares expert pools across layers: self-reported lower perplexity than standard MoE at 114M–1.15B

Visão rápida
Verificado 2026-09-18 15:03 GMT+8

MoRE lets adjacent layer groups share a single expert pool while each layer keeps its own router, distinguished by a lightweight learnable depth embedding; the authors self-report that at three scales from 114M to 1.15B parameters, under equal compute and parameter budgets, perplexity is lower than standard MoE and weight-sharing architectures, requiring only minor changes to existing MoE implementations.

Previously, standard MoE kept each layer's experts separate, while weight-sharing architectures struggled to distinguish layers; MoRE aims to combine parameter efficiency with layer distinction via a shared pool plus depth embeddings.

The perplexity comparisons are self-reported by Eric S. Qiu, Kilian Q. Weinberger and seven authors in total, with no third-party benchmark named, and the scales are small.

The results have not been independently reproduced; the preprint was submitted to arXiv on September 16, updated to v2 on the 17th, with the page noting acceptance to COLM 2026 (arXiv:2609.18176).

Fontes:arxiv.org

Pesquisa

G2G freezes the base, trains ~32M params, self-reports SOTA on four datasets

Visão rápida
Verificado 2026-09-18 15:00 GMT+8

With the MapAnything base frozen, G2G trains only three modules — a resampler, a cross-group bridge and a multi-frame pose head — about 32M parameters, under 6% of the full model, to estimate relative 6-DoF pose between two groups of images, using only relative pose supervision.

Previously, the usual approach was to fine-tune the whole model, which costs more training and risks eroding the base's general capabilities.

The authors self-report that on four datasets — indoor/outdoor simulation, real cross-season, and zero-shot sim-to-real — cross-sequence relocalization and multi-camera rig odometry accuracy reach what they claim is state of the art; code and weights are open-sourced.

The results are not independently reproduced; the work is by Yufei Wei and seven others, as an arXiv preprint (2606.08284) first posted June 6 and updated to v3 on September 17.

Sources: arXiv abstract page ↗, authors' code repository ↗

Fontes:arxiv.org

Pesquisa

Fathom adapts K-cache bits per query, 1.67x faster decode step

Visão rápida
2026-09-21 12:00 GMT+8

In the million-token setting, each query can now adaptively decide how many bits of the 4-bit K cache to read: the author self-reports that Qwen3-8B decode-step GPU time is 1.67x faster than 136-bit sparse scanning, with the KV cache and index residing in host memory.

Previously, sparse scanning read the K cache at a fixed bit width and could not adjust the read amount per query.

The comparison is the author's own measurement: timing used synthetic KV, and in a real coding-agent session 92 bits already matched the step consistency of the most accurate scan; there is no speedup when the index is in GPU memory.

Single-author arXiv preprint (2609.17652), submitted September 15, updated to v2 on the 17th, not independently verified.

Sources: arXiv preprint ↗, author's code repository ↗

Fontes:arxiv.org

Pesquisa

ONNX SmolVLA: Latency Halved but Spatial Success Falls to 41%

Visão rápida
Verificado 2026-09-18 14:31 GMT+8

Deploying SmolVLA with ONNX Runtime on an RTX 2060 can cut p99 latency from 1181 ms to 601/532 ms, but in LIBERO simulation Spatial success drops from 70% to 41%, while Object stays at 89%.

Previously, anyone trying to cut this model's inference latency lacked latency-versus-success data for an ONNX deployment path, making it easy to misjudge capability retention after compression; an audit also found the INT8 export was actually an FP32 graph.

Single-author preprint self-report: evaluated in LIBERO simulation, Spatial success 70%→41%, Object stays at 89%, p99 latency 1181 ms→601/532 ms; the author says a language width of 24 restores it to 75% at roughly half the baseline latency, but width does not explain all of the difference.

The result is limited to an RTX 2060 and LIBERO simulation and has not yet been reproduced by others; the preprint was submitted on September 12 and updated on the 16th.

Fontes:arxiv.org

Pesquisa

M2Tok's multi-head multi-codebook action tokenization reportedly cuts reconstruction error and lifts VLA success

Visão rápida
Verificado 2026-09-18 14:26 GMT+8

M2Tok, an action tokenizer that splits latent action features across heads with an independent codebook per head, reportedly reduces reconstruction error and improves VLA task success rates, with code released open source.

Earlier discrete action tokenizers were limited by a single codebook's quantization expressiveness, incurring higher reconstruction loss and constraining downstream VLA performance.

Chunpu Xu, Yao Mu and nine authors in total self-report that M2Tok's reconstruction loss is significantly lower than existing discrete action tokenization methods, and that VLAs built on it achieve higher success rates on RoboTwin, Simpler-Env and 3 zero-shot real-world tasks.

These are self-reported results with no third-party replication yet; the work was submitted to arXiv on September 16 (id 2609.18259, labeled ECCV 2026).

Fontes:arxiv.org

Pesquisa

Uncertainty-guided test-time optimization boosts depth-only 3D segmentation

Visão rápida
Verificado 2026-09-18 14:11 GMT+8

Indoor robots can now do open-vocabulary 3D segmentation from depth alone in privacy-preserving settings where RGB is banned: UTTO uses predictive uncertainty as a signal to test-time optimize a frozen open-vocabulary 3D segmentation backbone, and the authors report consistent gains across multiple depth-only backbones on ScanNet and Matterport3D.

Previously such settings had to rely on RGB or retrain the backbone, and segmentation quality suffered when RGB was banned.

The result is self-reported by Huang and two other authors, with a privacy-recoverability analysis and a real-robot case study; no independent reproduction has yet appeared, and the paper was first submitted on July 1, 2026 and updated to v2 on September 17 (arXiv:2607.00978).

Source: arXiv:2607.00978 abstract page (v2, updated 2026-09-17) ↗

Fontes:arxiv.org

Pesquisa

Near-domain negatives drop fraud-detection classifiers from perfect Macro-F1 to 0.65–0.68

Visão rápida
Verificado 2026-09-18 14:00 GMT+8

TeleAntiFraud 2.0 freezes a monthly benchmark of 900 Chinese phone-call audios (600 fraud, 300 near-domain non-fraud), letting readers test how fraud-detection classifiers really handle near-domain negatives. Previously, on unrelated or ordinary negatives, classifiers reached perfect Macro-F1, masking their failure on same-context samples.

The authors self-report that in controlled text experiments, three classifiers dropped to Macro-F1 of 0.65–0.68 once near-domain negatives were swapped in; full-audio and ASR+LLM evaluations also revealed class-prior shortcuts, prediction collapse, and snapshot sensitivity. Results are author self-reported.

Boundary: results are self-reported with no third-party replication; submitted to arXiv as a preprint by Huiyuan Liu, Zhiming Ma, and twelve other authors on September 16 (v2 updated September 17, id 2609.18748).

Sources: arXiv abstract page (2609.18748) ↗, authors' data and code repository ↗

Fontes:arxiv.org

Pesquisa

On-device model guides home wound photos, self-reported usability good

Visão rápida
Verificado 2026-09-18 13:55 GMT+8

Elderly patients with chronic wounds can now photograph and record their wounds at home: an on-device lightweight model segments the wound and guides the capture, while doctors monitor remotely and conduct video consultations — this is what the WoundAIssist app provides.

Previously such patients had to travel to hospital for doctors to examine wounds directly. The authors self-report that a usability study involving patients and dermatologists showed good ease of use, but the abstract gives no quantitative metrics or clinical outcomes.

Vanessa Borst and six other authors updated the paper to v2 on arXiv on September 17 (first submitted June 2025), with the page noting publication in an ACM Transactions on Computing for Healthcare 2026 special issue.

Source: arXiv abstract page (v2, updated 2026-09-17) ↗

Fontes:arxiv.org

Pesquisa

CSWAM self-reported to lift RoboTwin 2.0 Clean-to-Randomized success from 10.16% to 45.18%

Visão rápida
2026-09-22 12:00 GMT+8

Readers can now know: CSWAM, a method that self-reports lifting RoboTwin 2.0 Clean-to-Randomized transfer success from 10.16% to 45.18%, while retaining efficient action-only inference. Previously, FastWAM-style world action models lacked causal semantic conditioning on sparse observation histories, leaving cross-randomization transfer success at roughly 10%.

The measurements were self-reported by eight authors: adding a V-JEPA 2.1-based causal semantic expert to FastWAM-style world action models, conditioning action denoising on sparse observation histories, and after embodied pretraining, RoboTwin 2.0 Clean-to-Randomized success rose from 10.16% to 45.18%; on two real-robot tasks across three OOD difficulty levels, average success rose from 27.5% to 70.0%. These are author self-reported simulation and self-test results.

Boundary: results are not independently reproduced; the work was submitted to arXiv on September 16 (v2 updated September 17, arXiv:2609.18462).

Fontes:arxiv.org

Pesquisa
Próxima página de leitura →