読書リサーチレーダー投資フレームワーク
ログイン / 新規登録
ログイン / 新規登録
読書リサーチレーダー投資フレームワーク
アーカイブ →

読書

2026-10-0155 投稿

BMW to cut a fifth of management, says AI will speed decisions

トピック · 宝马重组重大
2026-10-01 08:06 GMT+8

BMW announced a restructuring plan on September 30, cutting one fifth of its departments and corresponding management positions by mid-2027, with around 8,000 jobs in Germany affected.

According to Reuters, new CEO Milan Nedeljković said AI will become an important tool for improving organizational efficiency and speeding up decision-making, insisting "this is not a cost-cutting program."

The backdrop: BMW's third profit warning in just over three years on weak China performance. Its core automotive margin stands at 2.3% against a 3%-5% target for 2028, and the stock has lost more than a third of its value over the past year.

ソース:ithome.com

リサーチ

Meta's AI assistant hit 5 million US downloads in 22 days, lifted by its own ad spend

トピック · Meta助手Museクイックテイク
2026-10-01 07:45 GMT+8

According to Sensor Tower, Meta's agentic AI assistant Muse passed 5 million US downloads on September 30, 22 days after launch, topping the App Store free chart for 12 consecutive days.

That is far faster than ChatGPT at 56 days, Grok at 103 days and Claude at 492 days. Sensor Tower counts only 23 US apps ever crossing 5 million downloads within 22 days; before Muse, the only non-game ones were Disney+, HBO Max and Threads.

The growth has a clear paid component: Meta ramped up advertising from September 16, lifting downloads 73% that day, and over the past two weeks Muse took up to 50% of Meta's own ad inventory, shrinking Facebook's and WhatsApp's share. The 5 million figure is a projection by Sensor Tower analyst Kara Lee; the firm's original report is not linked in the source.

ソース:ithome.com

リサーチ

Entity indexing preserves raw dialogue for agent long-term memory

トピック · EnSIMem实体记忆クイックテイク
2026-09-24 12:00 GMT+8

Agent long-term memory no longer has to rely on lossy summaries: EnSIMem builds indexes of the form [entity][entity_type][property: value], each preserving source turns and temporal information, with the code open-sourced on GitHub.

Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it hard for an agent to identify the right entity, property and evidence. EnSIMem organizes interactions offline into theme-coherent episodes and, online, decomposes each request into evidence requirements and answers from the index.

Xuanyu Meng and five coauthors self-report that the system achieves high answer accuracy while keeping contexts compact; the accompanying repository shows the experiments run on the LoCoMo and LongMemEval long-term conversational memory benchmarks, with ablation scripts included, but the abstract names no baseline or number. The preprint was submitted on September 23, and the performance claim awaits the paper body and independent reproduction.

ソース:arxiv.org

リサーチ

Repairing an agent's failed experience is not the same as reusing it

クイックテイク
2026-09-29 12:00 GMT+8

The value of repairing an agent's failed experience is distinct from the value of reusing it: ThinkingBox's Full mode shows a 44-percentage-point correction gain, yet 12 of its 15-point edge over Skill comes from worse uncorrected performance, not better corrected memory.

Looking only at correction gains can mistake repair value for reuse value, overstating the effect of memory updates.

Experiments by Yanfei Zhang and Xu Lin, submitted to arXiv on September 28, covered 3,300 runs under eleven conditions; the authors conclude memory updates require both a previous-version reference and a fresh-start reference.

All experiments ran on the authors' own ThinkingBox and APEX frameworks, with no third-party replication.

ソース:arxiv.org

リサーチ

Researchers propose monitoring closed models via open-weight internals

トピック · 模型解释学探针クイックテイク
2026-10-01 04:43 GMT+8

It may be possible to monitor deception and other misalignment in closed models through the internals of an open stand-in, without ever touching the closed model's weights.

A team led by conti published model hermeneutics on LessWrong on September 30: the closed model under study is the author, an open-weight reader reads its outputs, and linear probes are trained inside the reader. In the authors' own tests, a 27B reader probing a 397B author showed an average AUROC gap of only 0.004.

The team reports probes detecting reward hacking, sycophancy and deception in outputs from three frontier models including Claude Opus 4.6, Gemini 3.1 Pro Preview and GPT-5.4, beating direct reader interrogation in roughly half the settings. Distillation lifted the reader's deception-probing AUROC from 0.69 to 0.90.

This is roughly 1.5 months of early work; closed-model internals have no ground truth, so faithfulness is measured only against indirect open-weight baselines, and real audit use awaits independent verification.

ソース:lesswrong.com

リサーチ

GPT-6.1 Sol's no-CoT runs close 80% of the gap to Astra

トピック · GPT-6.1 Solクイックテイク
2026-09-30 11:57 GMT+8

GPT-6.1 Sol closes 80% of the no-CoT gap between GPT-6 Sol and GPT-6 Astra, according to researcher Rauno Arike's own runs across 27 tasks.

Previously, models in the no-CoT tier lagged Astra clearly, and there was little independent testing of that tier.

The runs reuse the task suite from an earlier Astra no-CoT evaluation, one sample per question; 6.1 Sol sits closer to Astra than to 6 Sol on 24 of 27 tasks. The 50% time-horizon estimate rises from 4 minutes for 6 Sol to 35 minutes, though most benchmarks saturate and the intervals are wide.

The author hypothesizes that 6.1 Sol is a looped transformer, possibly Astra Minor released under a different name, based on circumstantial evidence such as an Azure registry path; OpenAI has not confirmed this. The results are one researcher's own runs, and the Astra and GPT-5.5 scores were taken from the earlier post rather than re-run.

ソース:lesswrong.com

リサーチ

Corrigibility fund adds $48,000 in prizes for recent papers

トピック · 可纠正性研究基金クイックテイク
2026-10-01 00:08 GMT+8

The Corrigibility Research Fund announced an additional $48,000 on September 30, rewarding roughly two dozen researchers across about a dozen teams working on corrigibility — making AI systems accept human correction.

Fund manager Max Harms handed out $27,000 in a first round in July and plans to disburse more than $60,000 in December. The largest single award, $14,000, went to Rubi Hudson's paper on a corrigibility transformation.

The selection is one person's judgment; Harms himself calls the awards ad hoc and says the purse sizes should not be taken too seriously, and he deliberately passed over established figures such as Yudkowsky and Christiano.

ソース:lesswrong.com

リサーチ

Manus ships 2.0, with efficiency gains measured by its own tests

トピック · Manus 2.0クイックテイク
2026-09-29 09:35 GMT+8

Manus released version 2.0 on September 28, alongside Cue, a group-chat app for multiple agents — its first major update since parent Butterfly Effect restored independent operations on September 1.

According to the company's blog, 2.0 runs on its in-house Cascade framework. In one test configuration, compared with the previous system, token consumption fell 23.2%, task completion time shortened 28.2%, and running costs dropped 32%. These figures are vendor-measured, and the tested tasks are not specified.

Cue gives each agent its own email address, phone number and computer environment, and supports team collaboration; it remains in invite-only early access. This release is the overseas version, and the company says it is building a team for a China-market product.

ソース:zhidx.com

リサーチ

UserTesting rebrands as Auros, pivoting to human feedback for AI training

クイックテイク
2026-10-01 18:19 GMT+8

UserTesting has rebranded as Auros, extending its community of 7.6 million testers into human feedback for AI training and model evaluation.

CEO Eric Johnson told diginomica that AI training is not user testing, and the company will lean on "human intelligence," supplying humans to evaluate models and agents. An ARQ benchmark measuring trust and affinity in AI experiences is due to debut next week.

The rename was announced this week with no exact date given; the expansion is the company's own account, with revenue and the size of its AI business undisclosed.

ソース:diginomica.com

リサーチ

Modal's multi-node GPU clusters go generally available, billed by the second

クイックテイク
検証済み 2026-10-01 20:34 GMT+8

Modal announced on October 1 that Modal Clusters, its multi-node GPU cluster product, is now generally available, requested through a single decorator.

Clusters connect over RDMA, which the company says reaches up to 6.4 Tbps per node and is auto-configured for PyTorch and NCCL; a gang scheduler allocates all nodes at once from Modal's shared capacity pool, and usage is billed by the second. The company says the product was battle-tested for a year and a half before this release.

For teams without their own machine room, this means distributed training and inference can rent multi-node capacity in short bursts instead of hourly reservations. The bandwidth and acquisition-speed figures come from Modal's own blog and have no third-party measurements yet.

ソース:modal.com

リサーチ

A few dozen samples can silently rewrite a small model's behavior

トピック · 预训练投毒样本数クイックテイク
2026-10-01 10:45 GMT+8

About 56 wrong-answer samples can flip a 20-million-parameter model to a different answering route on half the questions while accuracy on normal tests stays perfect — the author's own reported result, provided a few hundred samples first teach the model to bypass a shortcut.

Such behavior rewriting was previously thought to require large amounts of data: Souly et al. found in 2025 that pre-training poisoning takes about 250 documents regardless of model or data size. In this experiment, scaling the training set from 32,000 to 320,000 samples left the required special samples at a few hundred (about 378 at 96,000), meaning the absolute count matters, not the proportion; the task was a manually constructed verb-question set.

In the September 30 LessWrong post the author offers an explanation: ordinary samples have a capped strengthening effect on the long route, and only samples where the shortcut is blocked or wrong truly reinforce it. The paper is under review; whether the finding generalizes to real large models remains unverified.

ソース:lesswrong.com

リサーチ

Most open models say they're being tested, but don't change behavior

トピック · 模型评估意识クイックテイク
2026-10-01 11:07 GMT+8

Most open reasoning models say "I might be being tested" in their chain of thought, yet do not change their behavior because of it — fears that safety evals are being distorted may be overstated.

The concern had been that models noticing they are being evaluated would alter their answers and skew safety evals; a first-party study tested whether that concern holds.

Sixteen open reasoning models were tested with the Fortress safety prompt dataset. Rates of verbalized eval awareness ranged from very rare to about a third of rollouts and were largely unrelated to model size; within the same question, rollouts containing such statements refused more often, in 14 of 16 models. The causal tests point the other way: inserting "this might be a test" raised refusal about as much as any sentence about being manipulated, and after removing the model's own eval-awareness sentences, only Nemotron 3 Super and Qwen3 32B changed behavior. The authors note this does not strictly prove causality, and the findings may not extend to frontier closed models.

ソース:lesswrong.com

リサーチ
次の読書ページ →