LectureRechercheRadarCadre d'investissement
Connexion / Inscription
Connexion / Inscription
LectureRechercheRadarCadre d'investissement
Archives de lecture →

Lecture

2026-10-0155 publications

BMW to cut a fifth of management, says AI will speed decisions

Sujet · 宝马重组Matériel
2026-10-01 08:06 GMT+8

BMW announced a restructuring plan on September 30, cutting one fifth of its departments and corresponding management positions by mid-2027, with around 8,000 jobs in Germany affected.

According to Reuters, new CEO Milan Nedeljković said AI will become an important tool for improving organizational efficiency and speeding up decision-making, insisting "this is not a cost-cutting program."

The backdrop: BMW's third profit warning in just over three years on weak China performance. Its core automotive margin stands at 2.3% against a 3%-5% target for 2028, and the stock has lost more than a third of its value over the past year.

Sources :ithome.com

Recherche

Meta's AI assistant hit 5 million US downloads in 22 days, lifted by its own ad spend

Sujet · Meta助手MusePrise rapide
2026-10-01 07:45 GMT+8

According to Sensor Tower, Meta's agentic AI assistant Muse passed 5 million US downloads on September 30, 22 days after launch, topping the App Store free chart for 12 consecutive days.

That is far faster than ChatGPT at 56 days, Grok at 103 days and Claude at 492 days. Sensor Tower counts only 23 US apps ever crossing 5 million downloads within 22 days; before Muse, the only non-game ones were Disney+, HBO Max and Threads.

The growth has a clear paid component: Meta ramped up advertising from September 16, lifting downloads 73% that day, and over the past two weeks Muse took up to 50% of Meta's own ad inventory, shrinking Facebook's and WhatsApp's share. The 5 million figure is a projection by Sensor Tower analyst Kara Lee; the firm's original report is not linked in the source.

Sources :ithome.com

Recherche

Entity indexing preserves raw dialogue for agent long-term memory

Sujet · EnSIMem实体记忆Prise rapide
2026-09-24 12:00 GMT+8

Agent long-term memory no longer has to rely on lossy summaries: EnSIMem builds indexes of the form [entity][entity_type][property: value], each preserving source turns and temporal information, with the code open-sourced on GitHub.

Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it hard for an agent to identify the right entity, property and evidence. EnSIMem organizes interactions offline into theme-coherent episodes and, online, decomposes each request into evidence requirements and answers from the index.

Xuanyu Meng and five coauthors self-report that the system achieves high answer accuracy while keeping contexts compact; the accompanying repository shows the experiments run on the LoCoMo and LongMemEval long-term conversational memory benchmarks, with ablation scripts included, but the abstract names no baseline or number. The preprint was submitted on September 23, and the performance claim awaits the paper body and independent reproduction.

Sources :arxiv.org

Recherche

Repairing an agent's failed experience is not the same as reusing it

Prise rapide
2026-09-29 12:00 GMT+8

The value of repairing an agent's failed experience is distinct from the value of reusing it: ThinkingBox's Full mode shows a 44-percentage-point correction gain, yet 12 of its 15-point edge over Skill comes from worse uncorrected performance, not better corrected memory.

Looking only at correction gains can mistake repair value for reuse value, overstating the effect of memory updates.

Experiments by Yanfei Zhang and Xu Lin, submitted to arXiv on September 28, covered 3,300 runs under eleven conditions; the authors conclude memory updates require both a previous-version reference and a fresh-start reference.

All experiments ran on the authors' own ThinkingBox and APEX frameworks, with no third-party replication.

Sources :arxiv.org

Recherche

Researchers propose monitoring closed models via open-weight internals

Sujet · 模型解释学探针Prise rapide
2026-10-01 04:43 GMT+8

It may be possible to monitor deception and other misalignment in closed models through the internals of an open stand-in, without ever touching the closed model's weights.

A team led by conti published model hermeneutics on LessWrong on September 30: the closed model under study is the author, an open-weight reader reads its outputs, and linear probes are trained inside the reader. In the authors' own tests, a 27B reader probing a 397B author showed an average AUROC gap of only 0.004.

The team reports probes detecting reward hacking, sycophancy and deception in outputs from three frontier models including Claude Opus 4.6, Gemini 3.1 Pro Preview and GPT-5.4, beating direct reader interrogation in roughly half the settings. Distillation lifted the reader's deception-probing AUROC from 0.69 to 0.90.

This is roughly 1.5 months of early work; closed-model internals have no ground truth, so faithfulness is measured only against indirect open-weight baselines, and real audit use awaits independent verification.

Sources :lesswrong.com

Recherche

GPT-6.1 Sol's no-CoT runs close 80% of the gap to Astra

Sujet · GPT-6.1 SolPrise rapide
2026-09-30 11:57 GMT+8

GPT-6.1 Sol closes 80% of the no-CoT gap between GPT-6 Sol and GPT-6 Astra, according to researcher Rauno Arike's own runs across 27 tasks.

Previously, models in the no-CoT tier lagged Astra clearly, and there was little independent testing of that tier.

The runs reuse the task suite from an earlier Astra no-CoT evaluation, one sample per question; 6.1 Sol sits closer to Astra than to 6 Sol on 24 of 27 tasks. The 50% time-horizon estimate rises from 4 minutes for 6 Sol to 35 minutes, though most benchmarks saturate and the intervals are wide.

The author hypothesizes that 6.1 Sol is a looped transformer, possibly Astra Minor released under a different name, based on circumstantial evidence such as an Azure registry path; OpenAI has not confirmed this. The results are one researcher's own runs, and the Astra and GPT-5.5 scores were taken from the earlier post rather than re-run.

Sources :lesswrong.com

Recherche

Corrigibility fund adds $48,000 in prizes for recent papers

Sujet · 可纠正性研究基金Prise rapide
2026-10-01 00:08 GMT+8

The Corrigibility Research Fund announced an additional $48,000 on September 30, rewarding roughly two dozen researchers across about a dozen teams working on corrigibility — making AI systems accept human correction.

Fund manager Max Harms handed out $27,000 in a first round in July and plans to disburse more than $60,000 in December. The largest single award, $14,000, went to Rubi Hudson's paper on a corrigibility transformation.

The selection is one person's judgment; Harms himself calls the awards ad hoc and says the purse sizes should not be taken too seriously, and he deliberately passed over established figures such as Yudkowsky and Christiano.

Sources :lesswrong.com

Recherche

Manus ships 2.0, with efficiency gains measured by its own tests

Sujet · Manus 2.0Prise rapide
2026-09-29 09:35 GMT+8

Manus released version 2.0 on September 28, alongside Cue, a group-chat app for multiple agents — its first major update since parent Butterfly Effect restored independent operations on September 1.

According to the company's blog, 2.0 runs on its in-house Cascade framework. In one test configuration, compared with the previous system, token consumption fell 23.2%, task completion time shortened 28.2%, and running costs dropped 32%. These figures are vendor-measured, and the tested tasks are not specified.

Cue gives each agent its own email address, phone number and computer environment, and supports team collaboration; it remains in invite-only early access. This release is the overseas version, and the company says it is building a team for a China-market product.

Sources :zhidx.com

Recherche

UserTesting rebrands as Auros, pivoting to human feedback for AI training

Prise rapide
2026-10-01 18:19 GMT+8

UserTesting has rebranded as Auros, extending its community of 7.6 million testers into human feedback for AI training and model evaluation.

CEO Eric Johnson told diginomica that AI training is not user testing, and the company will lean on "human intelligence," supplying humans to evaluate models and agents. An ARQ benchmark measuring trust and affinity in AI experiences is due to debut next week.

The rename was announced this week with no exact date given; the expansion is the company's own account, with revenue and the size of its AI business undisclosed.

Sources :diginomica.com

Recherche

Modal's multi-node GPU clusters go generally available, billed by the second

Prise rapide
Vérifié 2026-10-01 20:34 GMT+8

Modal announced on October 1 that Modal Clusters, its multi-node GPU cluster product, is now generally available, requested through a single decorator.

Clusters connect over RDMA, which the company says reaches up to 6.4 Tbps per node and is auto-configured for PyTorch and NCCL; a gang scheduler allocates all nodes at once from Modal's shared capacity pool, and usage is billed by the second. The company says the product was battle-tested for a year and a half before this release.

For teams without their own machine room, this means distributed training and inference can rent multi-node capacity in short bursts instead of hourly reservations. The bandwidth and acquisition-speed figures come from Modal's own blog and have no third-party measurements yet.

Sources :modal.com

Recherche

A few dozen samples can silently rewrite a small model's behavior

Sujet · 预训练投毒样本数Prise rapide
2026-10-01 10:45 GMT+8

About 56 wrong-answer samples can flip a 20-million-parameter model to a different answering route on half the questions while accuracy on normal tests stays perfect — the author's own reported result, provided a few hundred samples first teach the model to bypass a shortcut.

Such behavior rewriting was previously thought to require large amounts of data: Souly et al. found in 2025 that pre-training poisoning takes about 250 documents regardless of model or data size. In this experiment, scaling the training set from 32,000 to 320,000 samples left the required special samples at a few hundred (about 378 at 96,000), meaning the absolute count matters, not the proportion; the task was a manually constructed verb-question set.

In the September 30 LessWrong post the author offers an explanation: ordinary samples have a capped strengthening effect on the long route, and only samples where the shortcut is blocked or wrong truly reinforce it. The paper is under review; whether the finding generalizes to real large models remains unverified.

Sources :lesswrong.com

Recherche

Most open models say they're being tested, but don't change behavior

Sujet · 模型评估意识Prise rapide
2026-10-01 11:07 GMT+8

Most open reasoning models say "I might be being tested" in their chain of thought, yet do not change their behavior because of it — fears that safety evals are being distorted may be overstated.

The concern had been that models noticing they are being evaluated would alter their answers and skew safety evals; a first-party study tested whether that concern holds.

Sixteen open reasoning models were tested with the Fortress safety prompt dataset. Rates of verbalized eval awareness ranged from very rare to about a third of rollouts and were largely unrelated to model size; within the same question, rollouts containing such statements refused more often, in 14 of 16 models. The causal tests point the other way: inserting "this might be a test" raised refusal about as much as any sentence about being manipulated, and after removing the model's own eval-awareness sentences, only Nemotron 3 Super and Qwen3 32B changed behavior. The authors note this does not strictly prove causality, and the findings may not extend to frontier closed models.

Sources :lesswrong.com

Recherche
Page de lecture suivante →