LecturaInvestigaciónRadarMarco de inversión
Iniciar sesión / Registrarse
Iniciar sesión / Registrarse
LecturaInvestigaciónRadarMarco de inversión
Archivo de lecturas →

Lectura

2026-09-1717 publicaciones

Humans Start Speaking a Median 151 ms Early; Current Voice Systems Struggle to Match

Resumen rápido
Verificado 2026-09-17 15:36 GMT+8

In smooth turn transitions, human listeners begin speaking a median 151 ms early, while current voice turn-taking systems cannot yet match that timing without producing too many false interruptions, with false positives concentrated in backchannel-dense conversational styles.

Previously there was no turn-taking evaluation benchmark covering multiple conversational styles with human annotation and a public leaderboard, making it hard to compare systems against human timing.

Researchers released the TurnBench benchmark: 30 hours of human-annotated two-person conversation data, a 104-hour training set and a public leaderboard, covering six conversational styles with triple annotation; the paper tested 14 systems, and the authors report the above results.

The paper was accepted to IEEE SLT 2026, with the camera-ready updated on September 16; it does not mention whether the benchmark has been reproduced by third parties.

Source: TurnBench paper page (arXiv) ↗ | TurnBench leaderboard and dataset ↗

Fuentes:arxiv.org

Investigación

Value induction breeds sycophancy: Apple's self-reported finding

Resumen rápido
Verificado 2026-09-17 15:10 GMT+8

After value-induction fine-tuning of chat LLMs, inducing one value causes the model to express other related and even opposing values, and any value induction increases anthropomorphic language, making the model more affirming of users and more sycophantic; inducing positive values improves safety. Previously, such fine-tuning targeted single value subsets without observing cross-value spillover.

The finding comes from experiments in "How Value Induction Reshapes LLM Behaviour", published on Apple's machine learning research page on September 16: chat LLMs were fine-tuned on value subsets of existing preference datasets, the first author completed the work while employed at Apple, and the conclusions are author-reported.

Boundaries: no independent verification yet, and not tested on production models; the paper was submitted to arXiv on May 8 (arXiv id 2605.07925), with the arXiv page noting acceptance to Findings of ACL 2026.

Fuentes:machinelearning.apple.com

Investigación

EgoPHI estimates hand-object contact and 3D force from one RGB image

Resumen rápido
Verificado 2026-09-17 13:39 GMT+8

You can now jointly estimate dense contact maps and 3D force distributions for hand and object meshes from a single RGB image and object geometry: EgoPHI uses physics simulation to generate per-vertex force supervision, with real-world validation on only two objects and eight participants.

Previously there was no method to obtain both contact and 3D force distributions from monocular RGB, and per-vertex force supervision was hard to collect directly.

Andela Ilic, Christian Holz and co-authors released EgoPHI on arXiv (submitted August 13, revised to v2 on September 15); the page notes acceptance to ECCV 2026.

No independent verification yet. Source: arXiv:2608.13014 abstract page ↗

Fuentes:arxiv.org

Investigación

2026-09-122 publicaciones

Dynamic power allocation recovers 63% of throttled performance gap

Verificado 2026-09-12 11:00 GMT+8

Under a 30% total power cut, allocating power dynamically along each task's power-performance curve recovered an average of about 1,500 tokens/s/task, or 63% of the gap between even allocation and the theoretical optimum.

Previous throttled scheduling mostly used even allocation, leaving this gap because it did not exploit differences in tasks' power-performance curves.

The result comes from 131 H200 training runs, 24 validation runs and 34 matched H100 tasks, self-reported by the preprint's authors.

Experiments used at most 32 GPUs, concentrated on H100/H200; the preprint was submitted on September 10, has not been peer-reviewed, and has no third-party replication; they do not extrapolate to 10,000-GPU clusters, inference workloads or grid benefits.

Fuentes:arxiv.org

Investigación

Real citations still fail to prove novelty

Verificado 2026-09-12 11:00 GMT+8

Even when citations are real and their conclusions point the right way, more than 70% of positive evidence still fails to logically support a paper's claim of novelty. NovGauge tested novelty judgments across 18 models using 619 paper pairs and 50 multi-paper sets; the authors report hallucination rates between 0% and 39%, with best Verified F1 around 43% to 72%.

Previously, novelty judgments relied on whether citations were real and whether their conclusions pointed the right way, but that standard misses a large share of positive evidence that does not hold up logically.

The measurement is author-reported: 619 paper pairs, 50 multi-paper sets, 18 models, hallucination rates 0% to 39%, best Verified F1 around 43% to 72%.

This is an author-built preprint benchmark that has not yet been independently reproduced; it is suited to testing whether citations support conclusions, and cannot be extrapolated to the accuracy of all AI literature reviews.

Fuentes:arxiv.org

Investigación

2026-09-011 publicación

Astronomy alert field extraction hits 99.8% agreement with audit

Verificado 2026-09-01 11:24 GMT+8

Field extraction from astronomy alerts can now reach about 99.80% agreement with human audit: the Astro-COLIBRI team processed 1,775 astronomy alerts, and 25,827 of 25,880 explicit field judgments matched an internal human audit, with errors concentrated in observation times and facility attribution.

Previously this kind of alert field curation relied on manual work, which was costly and hard to scale. The method chains a constrained schema, model extraction and deterministic rule checks together.

The agreement rate is about 99.80%, measured by the team's own internal human audit (first-party self-test). Boundary: it covers only one scientific domain, has not been peer-reviewed or externally reproduced, and cannot be extrapolated to the reliability of general enterprise agents; the preprint was released on August 24 (arXiv:2608.23270).


Tonight's earnings and macro calendar

Fuentes:arxiv.org

Investigación

2026-08-311 publicación

SKILL.state cuts hundred-step agent tokens to about one-sixteenth

Verificado 2026-08-31 10:30 GMT+8

SKILL.state cuts token use on a hundred-step warehouse task to about one-sixteenth of keeping the full history, while still reaching 0.94 accuracy; public agent benchmarks also show lower consumption.

Previously agents typically kept the full history, so tokens grew linearly with steps. SKILL.state compresses history into an explicit state, supporting the "state-first" engineering hypothesis, but not enough to prove general agent reliability.

The result is reported in an arXiv preprint, and the method relies on predefined sufficient patterns.

It does not yet cover multi-agent concurrent writes, and it has not been independently reproduced; next steps require code, reproduction and production tasks.

Source check Full paper and limitations ↗

  • Dell reports on September 1: FMP consensus expects revenue of $44.93 billion and EPS of $4.91. Watch AI server backlog conversion, infrastructure segment gross margin and working capital together. Company events page ↗
  • Credo reports on September 1: FMP consensus expects revenue of $473 million and EPS of $1.17. The key is not the high-speed interconnect demand narrative but customer concentration, product mix and incremental gross margin. Company events page ↗
  • Palo Alto Networks reports on September 1: FMP consensus expects revenue of $3.350 billion and EPS of $0.98. Check whether platformization brings remaining performance obligations and cash flow, not just AI security wording. Company events page ↗
  • Two midnight checks: whether Anthropic's actual billing switches to standard pricing, and whether Azure v2.0 calls return a retirement error; both should be read back from accounts and logs.

Data boundary FMP real-time stock quotes and press release retrieval are limited by plan; this page does not use real-time stock prices it did not obtain, nor does it present company targets, vendor tests or consensus expectations as realized results.

Fuentes:arxiv.org

Investigación

2026-08-301 publicación

Nearly 4,000 employees: gains from formal AI training do not persist

Verificado 2026-08-30 10:19 GMT+8

After a single round of formal AI training, employees' mature-usage behavior improved only in the month the training was completed, with no sustained improvement in later months — the conclusion of an analysis of nearly 4,000 back-office employees at a large professional services firm across eight months of 2025 AI usage data (158,496 conversations and 713,564 prompts). Companies have commonly followed a deploy path of "hand out accounts, run one training, wait for productivity," but behavioral evidence for it was lacking.

The researchers combined whether goals were clear, whether tasks covered multiple use types, and whether stepwise or verification strategies were used into a "mature usage" metric. Senior employees and the strategy, digital innovation, and project management teams scored higher; longer initial prompts and more iterations also correlated with these metrics. High usage frequency did not reliably correspond to more mature behavior. The authors accordingly propose that domain experience, task structure, and continuous feedback may matter more than a one-off course.

But the result does not prove training ineffective: trainees were not randomly assigned, and the study did not observe actual output quality or productivity. The sample comes from a single firm's back-office roles, and the main data window is still 2025 chat-style tools, so it cannot be directly extrapolated to today's agentic workflows; the study measures usage behavior, not final performance, and long prompts and multi-turn exchanges may simply reflect harder tasks. Next steps require randomized or staged training trials that place task outcomes, review errors, time spent, and sustained behavior on the same timeline. Compared with counting active users, firms can more verifiably check whether employees can define goals, provide context, check results, and embed these behaviors into concrete workflows.

The study was released as a working paper (arXiv:2608.27364, v1) and has not yet been replicated by independent teams.

Fuentes:arxiv.org

Investigación

2026-08-251 publicación

Isolated review context saves 33% tokens

Verificado 2026-08-25 08:00 GMT+8

Put a search agent's final review in a separate context and it uses 33% fewer tokens while keeping comparable performance across four tests.

Previously search execution and final review were often carried by the same context, so the model was easily swayed by its own earlier decisions and the review could not be independent.

The authors claim a model that took part struggles to judge its own prior decisions objectively; with isolated contexts performance across the four tests was comparable and tokens fell 33%, measured by the authors themselves.

The result has not yet been reproduced or compared item by item by other teams, and cross-model reproduction and component-ablation experiments are still needed; it reports only an agent phenomenon and cannot be directly extrapolated to the same mechanism in human investors. The paper was submitted on 24 August, is marked EMNLP 2026, and its version and experimental claims are at arXiv:2608.23045.

Fuentes:arxiv.org

Investigación

Estás al día en esta vista