LecturaInvestigaciónRadarMarco de inversión
Iniciar sesión / Registrarse
Iniciar sesión / Registrarse
LecturaInvestigaciónRadarMarco de inversión
Archivo de lecturas →

Lectura

2026-09-2456 publicaciones

Dual-track memory self-reports top EverMemBench score, unreplicated

Resumen rápido
2026-09-23 12:00 GMT+8

Multi-party dialogue memory has a new method signal: SpeakerMem-R1 self-reports a 62.33% submission to EverMind-AI's public EverMemBench leaderboard, the highest reported, claiming to tackle both bottlenecks of who said what and reconstructing states across members and time. The method stores speaker-labeled verbatim messages and derived states in dual tracks, merging evidence by entity, event and time at query time.

Before this, no approach handled speaker attribution and state reconstruction together; the authors self-report binary accuracies of only 47.9%, 69.2% and 61.9% on GroupMemBench, SocialMemBench and EverMemBench.

In the authors' self-reported 305-question controlled test, RL lifted the SFT writer's mean accuracy from 57.38% to 68.20%; the leaderboard-best 62.33% is also the authors' own submission.

All numbers are first-party evaluations, not independently reproduced, and the sub-48% GroupMemBench score shows the problem is far from solved; the preprint material is dated 2026-09-22. Treat as a method signal to track.

Fuentes:arxiv.org

Investigación

CoAx locates backup heads at 0.941 AUC, still unreplicated

Resumen rápido
2026-09-23 12:00 GMT+8

Readers doing circuit analysis or safety evaluation now have a new idea to use: conditional co-ablation (CoAx), which after ablating a primary component set ranks candidates by growth in ablation effect to find backups the intervention activates — the authors report 0.941 ROC-AUC recovering known backup heads on GPT-2's IOI circuit.

Previously, recovering backup heads masked by primary components relied on intact-state attribution, whose strongest baseline reached only 0.815 ROC-AUC on GPT-2-small's IOI circuit and could not directly complete an incomplete circuit.

The authors' reported measurement: CoAx recovers the documented backup heads on that IOI circuit at 0.941 ROC-AUC, and adding the selected heads to the incomplete circuit cuts incompleteness from 0.75 to 0.21; they also report it beats matched random completions on 8 non-GPT-2 models across 6 architecture families.

All numbers are first-party evaluations, backup-head recovery is mainly validated on the known-answer IOI setting, and nothing has been independently reproduced; preprint arXiv 2607.01940 (v3, 23 Sep 2026).

Fuentes:arxiv.org

Investigación

VERPO's token-level credit assignment avoids GRPO training collapse

Resumen rápido
2026-09-22 12:00 GMT+8

VERPO decomposes teacher guidance into token-level credit assignment, avoiding the optimization collapse caused by GRPO's sequence-level advantage estimation, and self-reports the highest multi-task average across five scientific reasoning and tool-use tasks.

Previously GRPO estimated advantages from whole-sequence scores, and the authors say this sequence-level granularity triggers optimization collapse.

The authors self-report larger gains on smaller models; this is a first-party benchmark with no independent reproduction.

The method includes a Fisher movement-cost controller and an FEC projection; version 4 was updated on September 23, and third-party validation is still pending.

Fuentes:arxiv.org

Investigación

Apple's probe guidance steers diffusion language models without an extra forward pass

Resumen rápido
2026-09-23 08:00 GMT+8

Readers can now learn of a way to guide diffusion language models without adding inference cost: probe guidance, introduced by Apple's machine learning research team, builds a guidance signal from the frozen internal states of an existing diffusion model, eliminating the extra inference forward pass that autoguidance requires, and the team says that applied to a 1.7B diffusion language model it consistently improves multiple-choice question answering benchmarks.

Previously, autoguidance required an extra forward pass and had lacked an account of why it works.

The team reports the method sets a new state of the art on unconditional generation for continuous diffusion language models. These results are first-party benchmarks.

The work also reports a mechanism finding: the weak model in autoguidance must come from a low-entropy region of training. If it holds, this may be more useful for diffusion language model research than any single benchmark number. The results have not been independently reproduced and come from a paper submitted on 2026-09-23.

Fuentes:machinelearning.apple.com

Investigación

Radical Numerics says its genome model can continue an aptamer optimization trajectory

Resumen rápido
2026-09-23 21:27 GMT+8

Radical Numerics CEO Eric Nguyen said in the show notes of the September 23 Latent Space episode that his genomic language model, in an aptamer experiment, recapitulated some held-out high-scoring samples after seeing only low-scoring RNA sequences.

Per the episode's show notes, the experiment held out the best-performing aptamers, showed the model only lower-scoring sequences in a rising series, and the model continued the trajectory on its own. Nguyen calls this chain-of-thought "thinking in DNA."

This is a first-party claim: the dataset size, scoring method and full results are not published with the notes, and it should not be read as independently verified performance. Watch for a checkable paper or benchmark.

Fuentes:latent.space

Investigación

Ringg self-reports 7M monthly AI-handled calls; 90% cost cut is vendor-claimed

Resumen rápido
2026-09-23 20:00 GMT+8

OpenAI published a customer story on September 23 saying that Ringg, an Indian voice support agent platform, now handles more than 7 million connected calls per month, with up to 65% of routine requests resolved without human involvement.

The story says Ringg cut model costs by roughly 90% after migrating selected real-time workloads from GPT-4.1 to GPT-5.6, and that at Policybazaar 67% of connected calls are handled without humans, with average response time falling from 8–12 minutes to under 60 seconds.

All resolution rates, cost savings and the 4.8 CSAT are self-reported by Ringg and OpenAI with no independent verification, and the 65% figure is an upper bound, so typical performance may be lower.

Fuentes:openai.com

Investigación

Maoxiang lead leaves ByteDance to bet on AI generative entertainment

Resumen rápido
2026-09-24 07:00 GMT+8

Liang Chenqi, who led the AI interactive-content product Maoxiang, left ByteDance in the summer of 2025 to found the AI entertainment startup Dongnian Yinxian.

Per the show notes of LateTalk episode 182, he spent seven years at ByteDance and built Maoxiang within its AI product unit Flow from 2023; his new company is betting on a blend of generative creation, expression and light gaming, and trains its own models.

His thesis: once people no longer work for a living, casual creation becomes a mainstream, high-frequency form of entertainment. No AI entertainment product has yet become an undisputed hit, and the startup has disclosed no funding or user figures; the above are his own statements.

Fuentes:podcast.latepost.com

Investigación

Models can be trained jointly without data leaving the device, but it was only validated in simulation

Material
Verificado 2026-10-03 18:32 GMT+8

Lectura a largo plazo · 《Communication-Efficient Learning of Deep Networks from Decentralized Data》(2016)

Some data is scattered across a large number of phones, and privacy does not allow it to be centralized on a server. The approach in the FedAvg paper is: each phone trains locally for a few rounds on its own data, and only the model changes are sent up to be averaged, without transmitting raw data, and the number of communication round trips is ten to a hundred times fewer than centralized training. This type of approach is collectively called federated learning.

Today, when vendors advertise that "the model can be improved without uploading data," what runs underneath is still this local training plus averaging, and later improvements are patches on top of it. When you hear this kind of advertising, first ask whether the results were measured in simulation or on real phones; the original paper itself only went as far as simulation.

If devices go offline, if data changes, or if someone deliberately submits bad updates, it handles none of these. If the question is whether this kind of training leaks privacy, or how well it works on real devices, it was only validated in simulation, so don't use it to draw conclusions.

Communication-Efficient Learning of Deep Networks from Decentralized Data (2016) | Next review 2027-09-20

Fuentes:arxiv.org

Investigación

Estás al día en esta vista