LeituraPesquisaRadarFramework de investimento
Entrar / Cadastrar
Entrar / Cadastrar
LeituraPesquisaRadarFramework de investimento
Arquivo de leituras →

Leitura

2026-09-1854 posts

SyncVoice claims zero-shot video dubbing SOTA in one model covering Chinese and English

Visão rápida
Verificado 2026-09-18 05:57 GMT+8

A model called SyncVoice can time synthesized speech to visual cues such as mouth movements in video, and its authors claim state-of-the-art zero-shot dubbing on the LRS3 dataset, with a unified Chinese-English model trained on about 600 hours of Chinese and 1,190 hours of English audiovisual data.

Previously, pretrained TTS generated speech from text alone without sensing visual cues, making it hard to align dubbing with lip movements, and Chinese and English dubbing typically required separate handling.

These are self-reported results by authors from Xiamen University, Xiaomi's MiLM Plus team and collaborators, with no independent verification yet.

Boundary: the results have not been reproduced by third parties; the work is an arXiv preprint (id 2512.05126), first submitted on November 23, 2025 and revised to v2 on September 15, 2026.

Fontes:arxiv.org

Pesquisa

Multi-judge voting tops three languages in hallucination detection

Visão rápida
Verificado 2026-09-18 05:45 GMT+8

Using multiple fine-tuned vision-language models as independent judges and fusing hallucination span predictions by character-level majority voting ranks first in three of four languages in SHROOM-Visions 2026 hallucination span detection, and places in the top three across all languages and metrics.

Previously, single models predicted hallucination spans directly and disagreement between models went unused; the authors state that model disagreement tracks human annotation disagreement.

The results are team self-reported, from Toqeer Ehsan et al.'s arXiv preprint (submitted September 15, updated to v2 on the 16th, accepted to UncertaiNLP 2026@EMNLP), with no task-organizer leaderboard corroboration yet.

Original: arXiv:2609.17327 ↗

Fontes:arxiv.org

Pesquisa

Context segmentation lets gemma-4 E4B solve 18.52% of picoCTF tasks standard agents miss

Visão rápida
Verificado 2026-09-18 05:08 GMT+8

On picoCTF tasks, the memory-constrained gemma-4 E4B, using a two-level "context segmentation" framework, solved 18.52% of tasks that standard agent execution failed to complete, while acting as an "intelligent search": rewards on par with brute-force retries and better token efficiency.

Previously, small-model agents on long-horizon CTF tasks suffered context bloat from accumulated tool outputs, making long-horizon tasks hard to complete.

The result was self-reported by Sebastiano Nordio and Michele Lotto, measured on the picoCTF dataset with the gemma-4 model (authors' claim, first-party testing).

It has no independent verification yet; the preprint was submitted September 11, updated to v3 on September 15 (arXiv:2609.12839), and accepted to the non-archival ESORICS 2026 RAISE workshop.

Fontes:arxiv.org

Pesquisa

Corridor-conditioned risk world model reaches intrusion AP 0.8567

Visão rápida
Verificado 2026-09-18 05:01 GMT+8

Readers can now know that CorrRisk-WM, a corridor-conditioned risk world model for planning proposed by Tingyu Guo and Reza Langari, self-reports intrusion AP 0.8567, first-entry 1-metre near-miss AP 0.8671, and the lowest measured open-loop collision rate of 4.88% across 29,176 scenarios drawn from 100 Waymo validation shards.

Prior risk world models lacked corridor conditioning and candidate-conditioned geometric interaction; ablations show that removing dynamic environment modelling or candidate-conditioned geometric interaction drops mean intrusion AP from 0.8590 to 0.7624 and 0.7252.

The intrusion AP 0.8567, near-miss AP 0.8671 and 4.88% open-loop collision rate are all author self-reports, not yet reproduced by third parties; the paper was submitted on 15 September, updated to v2 on the 16th, arXiv id 2609.16724.

Fontes:arxiv.org

Pesquisa

Sparse stations plus satellites yield 30m weather, error down 11–28%

Visão rápida
Verificado 2026-09-18 03:59 GMT+8

You can now obtain 30m-resolution temperature, dew point and wind fields across the contiguous US without resolving atmospheric dynamics: fusing sparse weather stations, high-resolution Earth observation and coarse-resolution atmospheric dynamics cuts error by 11–28% versus the strongest baseline.

Previously, obtaining near-surface weather fields at this scale required numerical models that resolve atmospheric dynamics.

Giezendanner, Wang and eight co-authors self-report: in spatiotemporal holdout tests, error falls 11–28% versus the strongest baseline, and at the median grid cell close to half the temperature variance is explained.

The results have not been peer-reviewed or independently verified; the arXiv preprint was first submitted 26 February 2026 and updated to v2 on 15 September.

Fontes:arxiv.org

Pesquisa

1.8M human-agent co-authored code edits self-reported to fine-tune better than human commit data

Visão rápida
Verificado 2026-09-18 03:52 GMT+8

Readers can now download and use a corpus of roughly 1.8 million code edits (371GB) co-authored by Claude Code, OpenAI Codex, and Cursor Agent with humans, whose natural language descriptions are nearly 10x longer than prior human commit datasets.

Previously, such agent-human collaborative edit data was lacking, and training relied on purely human commit corpora, which the authors say limited fine-tuning results.

A Northeastern University team self-reports: the corpus was collected from public GitHub records from early April to October 2025, and models fine-tuned on it outperform models trained on purely human commit corpora on most of the HumanEvalFix, CanItEdit, and similar benchmarks; the corpus is released on HuggingFace (nuprl/AgentPack).

Results are author self-reported, with no third-party reproduction yet; the work first appeared September 26, 2025 and was updated to v3 on September 16, 2026 (arXiv:2509.21891).

Fontes:arxiv.org

Pesquisa

Future-supervised reranker self-reports wins on all 12 confirmation tasks over SARAF

Visão rápida
Verificado 2026-09-18 03:30 GMT+8

A lightweight MLP reranker that trains on realized futures as supervision while using only past information at inference self-reports, by Yong-Hoon Choi and two co-authors, improved Pattern retrieval on six benchmarks and beat the protocol-matched SARAF rule on all 12 confirmation tasks.

Previously, similarity retrieval used historical samples directly without reranking; the last-value anchored L2 rule still wins in some domains, and historical relevance is domain-dependent.

These are self-reported results, not independently verified; the paper by Yong-Hoon Choi and two co-authors was first posted August 24 and updated to v2 on September 16 (arXiv:2608.23221).

Source: arXiv:2608.23221 abstract page (v2, 2026-09-16) ↗

Fontes:arxiv.org

Pesquisa

DBM's Statistical Fusion Edge Comes From Cross-Block Conditioning

Visão rápida
Verificado 2026-09-18 03:23 GMT+8

In statistical data fusion (two samples sharing covariates, disjoint outcome blocks), the advantage of a fine-tuned deep Boltzmann machine (DBM) comes almost none from generative pretraining and instead from cross-block conditioning when predicting the other outcome block: the author reports this contribution is positive across two datasets and 40 experimental cells, with magnitudes of +0.19 and +0.36 percentage points.

Previously, attributing the DBM's advantage to generative pretraining would misidentify its source; the author reports that against a tuned baseline with the same conditioning, the DBM is best in 37 of 40 cells.

The author self-reports that imputers capable of cross-block conditioning mostly lose accuracy because of it, while the DBM benefits in every cell.

The results have not been independently verified; single-author preprint, submitted September 14, updated to v2 on the 16th, arXiv:2609.14934.

Fontes:arxiv.org

Pesquisa

ACT encoder ablation's 35%→2% drop not reproduced in reruns of original code; latent zeroed at inference

Visão rápida
Verificado 2026-09-18 03:16 GMT+8

Rerunning the CVAE encoder ablation of Action Chunking Transformers on two simulation tasks in the original code, the paper's 35% to 2% success-rate drop did not reproduce, so readers now know that ablation result is not robust; small gains or losses remain uncertain, training duration and checkpoint selection can reverse which policy wins, and the cause of the drop is unidentified.

Previously one could only judge the encoder's role from the original paper's ablation numbers, while at inference ACT already zeroes the latent and does not use it; the author's timing shows skipping the encoder improves training throughput, and the code and evaluation tools are open-sourced.

The above is a self-reported result by Bo Kang as sole author, arXiv:2609.16745 (submitted September 15, updated to v2 on September 16), with no third-party reproduction yet, and the cause of the drop remains unidentified.

Source: arXiv:2609.16745 abstract page ↗; full-text HTML v2 ↗

Fontes:arxiv.org

Pesquisa

Codec metadata speeds VLM inference up to 3.3x

Visão rápida
Verificado 2026-09-18 02:52 GMT+8

Video codec metadata can serve directly as a runtime signal for streaming VLM inference: pruning image patches before visual encoding and selectively refreshing the KV cache across windows by frame type, with no model-specific training or offline profiling, lets concurrent streams reach up to 3.3x the baseline and cuts executed FLOPs by up to 93%.

Previously, speeding up streaming VLM inference usually required model-specific training or offline profiling, which was costly and hard to transfer to new models and workloads.

Yulin Zou and eight other authors self-report: across three VLMs and four video workloads, their vLLM-based implementation reaches up to 3.3x the baseline in concurrent streams, up to 5.3x faster average time-to-first-token, up to 93% fewer executed FLOPs, and at most a 4.64 percentage point drop in task quality. These are the authors' self-reported benchmark results.

The results do not cover other models or workloads, and no third party has reproduced them; the preprint was first submitted on April 7 and updated to v4 on September 15, arXiv:2604.06036.

Fontes:arxiv.org

Pesquisa

PICKT self-reported: difficulty features most informative for very hard items; text and knowledge-graph features estimate unseen items

Visão rápida
Verificado 2026-09-18 02:47 GMT+8

New items in intelligent tutoring systems lack response histories, degrading diagnostic reliability; after PICKT integrates multiple feature types, experiments show difficulty features are most informative for very hard items with extremely low accuracy, and fusing text and knowledge-graph features can estimate representations of unseen items via semantically or structurally similar items seen in training; the authors suggest prioritizing feature annotation according to educational-service needs. Results are self-reported by Wonbeen Lee and three co-authors.

Boundary: results are author self-reported with no third-party replication; the work first appeared in December 2025 and was updated as v2 on September 15, 2026 (arXiv:2512.07179).

Source: arXiv:2512.07179 abstract page ↗

Fontes:arxiv.org

Pesquisa

Skeleton-only emotion recognition hits 37.23% on hidden test, near human 39%

Visão rápida
Verificado 2026-09-18 02:26 GMT+8

Skeleton motion alone can now recognize 12 categories of performed emotion unseen by the performers: an 11-model ensemble by Naoto Nishida and Yoshio Ishiguro of the University of Tokyo scored 37.23% Macro-F1 on the MMAC Challenge 2026 hidden test set, close to the human 39% cited in the text, and the authors say it won the Best Performance Award.

The prior approach was 10-fold leave-performer-out cross-validation under the same protocol, yielding 36.80%, with a reproduced baseline of only 25.73%.

The 37.23% is author-reported, measured on the MMAC Challenge 2026 hidden test set; masking and counterfactual audits show the models rely on body-region motion evidence. The emotions are performed, and the results have not been independently reproduced; the data is the September 15 v2 preprint (arXiv:2609.02510), code at nawta/diema-challenge, author page nawta.github.io/mmac2026.

Fontes:arxiv.org

Pesquisa
Próxima página de leitura →