読書リサーチレーダー投資フレームワーク
ログイン / 新規登録
ログイン / 新規登録
読書リサーチレーダー投資フレームワーク
アーカイブ →

読書

2026-10-0258 投稿

Third-party tests put Snapdragon X2 Elite ahead of Intel's new chip in most measures

クイックテイク
検証済み 2026-10-02 10:42 GMT+8

Testing firm Signal65 published a report on October 1 saying Qualcomm's 18-core Snapdragon X2 Elite (X2E-88-100) beats Intel's Core Ultra X7 358H in CPU, AI and most battery-life tests.

In Geekbench 7 multi-core, the Snapdragon system scored 24,016 against Intel's 18,360, a 31% lead; in UL Procyon AI Computer Vision it scored 4,360 versus 2,302. Local video playback ran 16.1 hours against 14.4.

Office battery life went the other way, with Intel lasting 17.4 hours to Snapdragon's 14.3. The three test machines come from Lenovo and Dell and run different Windows versions, and the report does not say who commissioned the testing.

ソース:signal65.com

リサーチ

Bloomberg: Anthropic targets week of Nov 9 for IPO launch, timing fluid

重大
2026-10-02 07:43 GMT+8

Bloomberg reported on October 1, citing anonymous sources, that Anthropic plans to begin formally marketing its IPO in the week of November 9, listing before Thanksgiving; SiliconANGLE, relaying the report, notes internal discussions on the exact timeline are ongoing and could change.

According to the prospectus seen by Reuters, Anthropic recorded $4.6 billion in fiscal 2025 revenue, up more than 12 times year over year, alongside an operating loss above $8 billion and a net loss above $42 billion. It plans to spend $518 billion on AI infrastructure in the coming years.

The prospectus devotes several pages to what it calls "existential risks to humanity," citing incidents of self-preserving behavior and information concealment by its models. A delay could create a shockwave across other AI stocks.

ソース:siliconangle.com

リサーチ

ServiceNow pipeline turns model failures into agent training data automatically

トピック · 自动训练数据生成クイックテイク
2026-10-02 12:01 GMT+8

ServiceNow CoreAI released AutoSynthData on October 2, a pipeline that uses a target model's failures and a stronger teacher's successes to generate training tasks for enterprise agents.

The pipeline first runs diagnostic tasks in the target environment to find capabilities the model repeatedly fails, then generates new system specifications, user prompts and verifiers, executing each candidate in the environment before it enters training. As the model improves, task difficulty follows the remaining gaps.

The team ran experiments in its own EnterpriseOps Gym environment, but the loaded text cuts off before the results section, so no verifiable improvement numbers are available and the gains remain a vendor claim.

ソース:huggingface.co

リサーチ

Greylock co-leads Series A in Parallax, betting on 3D-printed gas turbines for data centers

トピック · Parallax燃气轮机クイックテイク
2026-10-01 23:30 GMT+8

Greylock announced on October 1 that it co-led the Series A in Parallax, with Founders Fund, Lux, General Catalyst and others participating. The funding amount was not disclosed.

Parallax's first product is a 10MW gas turbine designed for data centers, using 3D printing to control its own supply chain. The premise: turbine orders are booked out past 2030, and data centers cannot wait.

The 100GW US power shortfall cited in the post is the investor's retelling of the founder's thesis; the announcement offers no independent basis for it. The next test is whether the company actually delivers turbines.

ソース:greylock.com

リサーチ

New self-distillation method modestly improves computer-use agent training

トピック · 智能体自我改进クイックテイク
2026-10-01 12:00 GMT+8

ComputerSD converts real-time feedback from executed actions into training signals for computer-use agents; its authors report it beats training with outcome rewards only on the OSWorld-Verified benchmark.

The numbers: 1.9 percentage points higher on the Qwen3-VL-8B-Thinking backbone and 4.1 points higher on the specialized EvoCUA-8B, all measured by the paper's own authors.

The gains are modest and the benchmark was run by the authors themselves; whether the method delivers comparable benefits in real use awaits independent replication.

ソース:arxiv.org

リサーチ

Key-position confound inflates no-CoT reasoning depth on benchmark

クイックテイク
2026-10-02 22:39 GMT+8

Moving the initial state to the end of the prompt cuts GPT-6.1 Sol's no-chain-of-thought reasoning depth by a median 16%.

Neel Nanda's no-cot-bench had shown Astra, suspected to be a looped transformer, solving problems at roughly 8.6x the odds of Fable 5.1. Two researchers identify a confound: the key usually sits at the start of the prompt, so a model can compute while reading, meaning the benchmark measures both serial depth and the ability to spread computation across tokens.

They built key-last versions of six serial state-tracking tasks, with an unrelated-question control to rule out parsing difficulty. The median depth loss across 12 tasks was 16%, ranging from 9% to 45%. The authors tested only GPT-6.1 Sol, saying they lacked the compute to rerun other models.

ソース:lesswrong.com

リサーチ

50,000 agent error pairs released, proposed fixes lift pass rate to 51.1%

トピック · 智能体错误数据集クイックテイク
2026-10-01 12:00 GMT+8

Agent failure traces can now be reused without replaying rollouts: the Agent Error Dataset contains 50,228 error-diagnosis pairs, and in the authors' own tests proposed corrections raised replay verifier pass rates from 18.4% to 51.1%.

Previously, diagnosing failures required replaying the original tasks, which was costly and hard to reuse; the dataset retains source traces and execution metadata so failures can be re-diagnosed without replaying.

Kunlun Zhu and five co-authors released the dataset on arXiv on September 30; the pairs come from 9,961 tasks across 33 environments and 23 policy models, and the authors' own tests on 3,062 matched replay pairs produced the pass-rate result, while fine-tuning Qwen3-8B lifted exact-step agreement with teacher labels from 47.2% to 63.6%.

The replay comparison covers only environments that support replay, the results are the authors' own, and no third party has yet reproduced them.

ソース:arxiv.org

リサーチ

Agents self-improve without reward signals, matching top harness for $4

トピック · 智能体自我改进クイックテイク
2026-09-30 12:00 GMT+8

You can now improve agents without reward signals, using only records of past self-modification attempts: SelfSearch raised success over the initial agent in all six model-benchmark settings, and with DeepSeek V4 Flash, $4.03 of search cost produced a harness solving 82.0% of Terminal-Bench 2.1.

Improving agents previously meant repeated downstream evaluation, which is costly. SelfSearch has the agent read the reasoning, tool actions and outcomes of earlier modification attempts as its basis for improving, removing that cost.

Author-reported numbers include a gain of up to 11.2 percentage points on Terminal-Bench 2.1, and on SWE-bench Multilingual a 5.0-point success gain with 38.5% lower execution cost. The $4.03 harness matches Codex, the top scorer in a public nine-harness comparison.

The abstract does not say whether the 82.0% comparison reuses the original settings of the public nine-harness comparison or the authors' own evaluation setup; the paper by Jungwoo Yang, Injin Kong and Yohan Jo was submitted to arXiv on September 29, is author-reported and awaits replication.

ソース:arxiv.org

リサーチ

One round of skill-graph self-evolution lifts retrieval reward from 52.4% to 59.4%

トピック · SE-GoS技能图自进化クイックテイク
2026-10-02 12:00 GMT+8

Maintaining a skill retrieval graph from execution traces alone, one round of self-evolution lifts agent skill-library retrieval reward from 52.4% to 59.4% — with no model training, no retrieval-algorithm changes, and no separate model judging which skills relate.

Previous approaches relied on fixed indexing such as full-library loading or vector retrieval, which cannot automatically adjust the graph's connections and node descriptions from execution records.

The SE-GoS authors' own tests show one evolution round lifting average reward on SkillsBench from 52.4% to 59.4%, above full-library loading and vector retrieval baselines; on a held-out split it never saw, from 52.9% to 58.3%, at about two-thirds the input tokens of loading the full library. Repeating the round adds nothing; the abstract does not explain why.

The results are self-reported by the authors with no independent replication yet; the method is by Dawei Fu and four coauthors, updated on arXiv on October 1.

ソース:arxiv.org

リサーチ

Off-the-shelf role vectors curb sycophancy to 68-98% of purpose-built

トピック · 人格向量抑制谄媚クイックテイク
2026-10-02 12:00 GMT+8

Off-the-shelf role vectors such as Skeptic and Judge, built without any sycophancy data, can curb sycophancy to 68-98% of what purpose-built methods achieve.

Previously, curbing sycophancy required a CAA vector built specifically from sycophancy data; persona vectors are directions extracted from a model's activations, and steering along one raises or lowers a behavioral tendency.

Measured by Kelkar and five co-authors: on Gemma 2 27B, role vectors achieved on average 68% of CAA's sycophancy reduction, and 98% on Qwen 3 32B; steered Qwen still answered all 16 factual claims correctly. The benchmark is a philosophy-statements set with task-tuned coefficients.

All results are author-run, transfer to production models is unverified; the preprint was revised to v4 on September 30, is a spotlight at an ICML 2026 workshop, and has no independent replication yet.

ソース:arxiv.org

リサーチ

Confidence-guided decoding sends masked diffusion models down a shortcut that magnifies addition errors by an order of magnitude

トピック · 掩码扩散置信度捷径クイックテイク
2026-10-02 12:00 GMT+8

Masked diffusion models that order generation by confidence fall into a "confidence shortcut" that neglects long-range dependencies; the authors' own tests say confidence-aligned training objectives can raise addition error rates by an order of magnitude.

It was previously assumed that masked diffusion models (MDMs), which generate text by progressively unmasking tokens in any order, could in principle reveal intermediate steps along logical dependencies. Authors Dueun Kim and Albert No report that standard decoding simply prioritizes high-confidence tokens, so in multi-digit addition the models predict higher-order digits without tracking carry chains.

The authors' own controlled pretraining shows confidence-guided ordering often selects suboptimal sequences, and confidence-aligned training objectives can raise addition error rates by an order of magnitude. The experimental code is open source; the findings come from the authors' own setup and await independent reproduction.

ソース:arxiv.org

リサーチ

Delayed-compression memory CoEM lifts long-context reasoning F1 by over 10 points in self-tests

トピック · CoEM延迟压缩记忆クイックテイク
2026-09-30 12:00 GMT+8

Readers can now learn of CoEM, a memory-management method for long-context reasoning: potentially useful source excerpts are kept verbatim in a pending set, and only after later context clarifies their relevance does the method commit them to compressed memory or discard them, avoiding the loss of key details from premature compression; the authors' own tests show F1 gains of more than 10 points.

The prior practice was to compress memory early, which could discard key details before an excerpt's relevance became clear, causing information loss in long-context reasoning.

The decision policy is trained with reinforcement learning, and a frozen verifier accepts memory facts only when supported by retained excerpts. On 6,400-document inputs, the authors' self-tests show CoEM beats the strongest memory baseline by 10.4-11.4 F1 points on Qwen3.5-9B.

The experiments cover a single model, Qwen3.5-9B, and a single input length of 6,400 documents; performance on other models and shorter or longer contexts is unknown, and no independent replication exists yet. The code is open on GitHub for verification.

ソース:arxiv.org

リサーチ
次の読書ページ →