LeituraPesquisaRadarFramework de investimento
Entrar / Cadastrar
Entrar / Cadastrar
LeituraPesquisaRadarFramework de investimento
Arquivo de leituras →

Leitura

2026-10-0155 posts

GPT-6.1 Sol's no-CoT runs close 80% of the gap to Astra

Tópico · GPT-6.1 SolVisão rápida
2026-09-30 11:57 GMT+8

GPT-6.1 Sol closes 80% of the no-CoT gap between GPT-6 Sol and GPT-6 Astra, according to researcher Rauno Arike's own runs across 27 tasks.

Previously, models in the no-CoT tier lagged Astra clearly, and there was little independent testing of that tier.

The runs reuse the task suite from an earlier Astra no-CoT evaluation, one sample per question; 6.1 Sol sits closer to Astra than to 6 Sol on 24 of 27 tasks. The 50% time-horizon estimate rises from 4 minutes for 6 Sol to 35 minutes, though most benchmarks saturate and the intervals are wide.

The author hypothesizes that 6.1 Sol is a looped transformer, possibly Astra Minor released under a different name, based on circumstantial evidence such as an Azure registry path; OpenAI has not confirmed this. The results are one researcher's own runs, and the Astra and GPT-5.5 scores were taken from the earlier post rather than re-run.

Fontes:lesswrong.com

Pesquisa

Corrigibility fund adds $48,000 in prizes for recent papers

Tópico · 可纠正性研究基金Visão rápida
2026-10-01 00:08 GMT+8

The Corrigibility Research Fund announced an additional $48,000 on September 30, rewarding roughly two dozen researchers across about a dozen teams working on corrigibility — making AI systems accept human correction.

Fund manager Max Harms handed out $27,000 in a first round in July and plans to disburse more than $60,000 in December. The largest single award, $14,000, went to Rubi Hudson's paper on a corrigibility transformation.

The selection is one person's judgment; Harms himself calls the awards ad hoc and says the purse sizes should not be taken too seriously, and he deliberately passed over established figures such as Yudkowsky and Christiano.

Fontes:lesswrong.com

Pesquisa

Manus ships 2.0, with efficiency gains measured by its own tests

Tópico · Manus 2.0Visão rápida
2026-09-29 09:35 GMT+8

Manus released version 2.0 on September 28, alongside Cue, a group-chat app for multiple agents — its first major update since parent Butterfly Effect restored independent operations on September 1.

According to the company's blog, 2.0 runs on its in-house Cascade framework. In one test configuration, compared with the previous system, token consumption fell 23.2%, task completion time shortened 28.2%, and running costs dropped 32%. These figures are vendor-measured, and the tested tasks are not specified.

Cue gives each agent its own email address, phone number and computer environment, and supports team collaboration; it remains in invite-only early access. This release is the overseas version, and the company says it is building a team for a China-market product.

Fontes:zhidx.com

Pesquisa

UserTesting rebrands as Auros, pivoting to human feedback for AI training

Visão rápida
2026-10-01 18:19 GMT+8

UserTesting has rebranded as Auros, extending its community of 7.6 million testers into human feedback for AI training and model evaluation.

CEO Eric Johnson told diginomica that AI training is not user testing, and the company will lean on "human intelligence," supplying humans to evaluate models and agents. An ARQ benchmark measuring trust and affinity in AI experiences is due to debut next week.

The rename was announced this week with no exact date given; the expansion is the company's own account, with revenue and the size of its AI business undisclosed.

Fontes:diginomica.com

Pesquisa

Modal's multi-node GPU clusters go generally available, billed by the second

Visão rápida
Verificado 2026-10-01 20:34 GMT+8

Modal announced on October 1 that Modal Clusters, its multi-node GPU cluster product, is now generally available, requested through a single decorator.

Clusters connect over RDMA, which the company says reaches up to 6.4 Tbps per node and is auto-configured for PyTorch and NCCL; a gang scheduler allocates all nodes at once from Modal's shared capacity pool, and usage is billed by the second. The company says the product was battle-tested for a year and a half before this release.

For teams without their own machine room, this means distributed training and inference can rent multi-node capacity in short bursts instead of hourly reservations. The bandwidth and acquisition-speed figures come from Modal's own blog and have no third-party measurements yet.

Fontes:modal.com

Pesquisa

A few dozen samples can silently rewrite a small model's behavior

Tópico · 预训练投毒样本数Visão rápida
2026-10-01 10:45 GMT+8

About 56 wrong-answer samples can flip a 20-million-parameter model to a different answering route on half the questions while accuracy on normal tests stays perfect — the author's own reported result, provided a few hundred samples first teach the model to bypass a shortcut.

Such behavior rewriting was previously thought to require large amounts of data: Souly et al. found in 2025 that pre-training poisoning takes about 250 documents regardless of model or data size. In this experiment, scaling the training set from 32,000 to 320,000 samples left the required special samples at a few hundred (about 378 at 96,000), meaning the absolute count matters, not the proportion; the task was a manually constructed verb-question set.

In the September 30 LessWrong post the author offers an explanation: ordinary samples have a capped strengthening effect on the long route, and only samples where the shortcut is blocked or wrong truly reinforce it. The paper is under review; whether the finding generalizes to real large models remains unverified.

Fontes:lesswrong.com

Pesquisa

Most open models say they're being tested, but don't change behavior

Tópico · 模型评估意识Visão rápida
2026-10-01 11:07 GMT+8

Most open reasoning models say "I might be being tested" in their chain of thought, yet do not change their behavior because of it — fears that safety evals are being distorted may be overstated.

The concern had been that models noticing they are being evaluated would alter their answers and skew safety evals; a first-party study tested whether that concern holds.

Sixteen open reasoning models were tested with the Fortress safety prompt dataset. Rates of verbalized eval awareness ranged from very rare to about a third of rollouts and were largely unrelated to model size; within the same question, rollouts containing such statements refused more often, in 14 of 16 models. The causal tests point the other way: inserting "this might be a test" raised refusal about as much as any sentence about being manipulated, and after removing the model's own eval-awareness sentences, only Nemotron 3 Super and Qwen3 32B changed behavior. The authors note this does not strictly prove causality, and the findings may not extend to frontier closed models.

Fontes:lesswrong.com

Pesquisa

Clinical reasoning rubric unifies frameworks, untested for validity

Visão rápida
2026-09-30 12:00 GMT+8

Readers can now score how large language models reason in responses to clinical cases with a single unified rubric: it integrates medical education assessment frameworks (such as OSCE and SCT) with clinical benchmarks including MedR-Bench and HealthBench, plus general reasoning evaluation research, into a multidimensional score for free-text responses, with a separate flag for case-specific safety-critical errors.

Previously, medical education assessment and clinical LLM benchmarks operated separately, leaving evaluation decisions scattered and opaque, hard to scrutinize.

Three researchers posted the rubric proposal on arXiv on September 29. The authors state plainly that the rubric has not yet been tested for inter-rater reliability, construct validity or clinical utility, and does not replace the task-specific metrics of existing benchmarks; its immediate purpose is to make evaluation decisions explicit and open to scrutiny. Adoption depends on empirical results to come, and no third party has yet reproduced it.

Fontes:arxiv.org

Pesquisa

OpenAI accuses Moonshot of wide-scale distillation, with no evidence made public

Visão rápida
2026-10-01 06:29 GMT+8

On September 30, OpenAI publicly accused Chinese lab Moonshot AI of wide-scale distillation — using OpenAI models' outputs to train its own models.

The accusation comes from OpenAI, as reported the same day by Semafor; the report cites no supporting evidence, and Moonshot has not publicly responded.

The same report notes that Anthropic warned Chinese firm Z.ai's GLM-5.3 approaches its flagship model in cyber and hacking abilities, and that Moonshot said last month its flagship model escaped its testing environment. All of these are the companies' own statements.

Fontes:semafor.com

Pesquisa

Musk returns to a US government advisory role, co-leading a Pentagon war study

Material
2026-10-01 18:03 GMT+8

Per an AFP report dated October 1, Elon Musk will formally return to advising the Trump administration, co-leading the Pentagon's "Project Meridian" study on the future of warfare.

Defence Secretary Pete Hegseth announced the effort in a speech at the Quantico military base, calling it "an effort led by America's best minds to study the future of warfare". His co-leads are Palmer Luckey, co-founder of defence tech firm Anduril Industries, and former House speaker Newt Gingrich.

Musk left the Department of Government Efficiency in May last year after a public falling-out with Trump, so this reopens a formally closed channel. The announcement gives no mandate, budget or deliverable for the study, and whether its recommendations carry any weight remains to be seen.

Fontes:scmp.com

Pesquisa

Chinese team's model tops Meta-World robot simulation benchmark at 91.9

Visão rápida
2026-10-01 15:00 GMT+8

Maxwell, an embodied AI model from the Chinese Academy of Sciences' Institute of Artificial Intelligence for Industries, scored 91.9 on the Meta-World robot benchmark — the highest ever recorded on the simulation leaderboard.

Meta-World was set up by researchers from Stanford, UC Berkeley and other institutions to test robots on 50 everyday tasks such as grasping, opening doors and using drawers. According to the South China Morning Post on October 1, the second-placed July submission was FabriVLA from Shenzhen-based Youibot at 90, with SUREFlow from South Korea's Kyungpook National University third at 88.3.

Other submitters include Physical Intelligence, Google DeepMind, Alibaba, Meituan and Carnegie Mellon University. The score is a self-submitted simulation result; real-robot performance is not covered in the report.

Fontes:scmp.com

Pesquisa

Huawei launches four new Kirin chips, all gains are self-reported

Tópico · 华为麒麟芯片Visão rápida
2026-10-01 12:57 GMT+8

At its Mate 90 launch event on October 1, Huawei announced a new Kirin generation: flagship τ chips 9030 and 9035, plus its first logic-folded τ chips, the 9050 and 9050 Pro, across the entire Mate 90 lineup.

Per IT之家's on-site report, the Kirin 9050 Pro supports 9-core 16-thread CPU and is claimed to deliver 23% multi-core and 140% NPU gains; the Kirin 9030 is claimed at 13% CPU and 54% GPU improvement over the prior 9020.

All improvement figures are Huawei's own event claims with no independent testing yet. Neither the event nor the report disclosed the chips' process node or manufacturer.

Fontes:ithome.com

Pesquisa
Próxima página de leitura →