ReadingResearchRadarInvestment framework
Sign in / Sign up中文
Sign in / Sign up中文
ReadingResearchRadarInvestment framework

Reading

2026-10-0112 posts

OpenAI trims tokens and raises prices, pushing compute costs onto heavy users

Material
2026-09-30 16:59 UTC

At its developer day on September 29, OpenAI adjusted its subscriptions: the $200-per-month tier now comes with half the tokens, and a new $500-per-month tier has been added.

According to Semafor tech editor Reed Albergotti, the pricing change was announced at the same event that introduced the dots agent tool, and he reads it as evidence of compute scarcity in AI: always-on agentic assistants burn large amounts of tokens, and serving them is expensive.

He notes OpenAI recently spent about $10 million over roughly a weekend to solve part of the Navier-Stokes equations. The pricing details come from his column rather than an OpenAI pricing page linked in the article, so exact allowances should be checked against OpenAI's own page.

Sources:https://www.semafor.com/article/09/30/2026/openais-new-pricing-tiers-highlight-compute-bottleneck

Research

ElevenLabs valuation doubles in a year as employee tender completes

Material
2026-09-30 16:28 UTC

ElevenLabs has completed a $300m employee tender offer at a doubled $22bn valuation.

CEO Mati Staniszewski announced the closing on X on September 30; per Tech.eu, Wellington and T Rowe Price led the offer, with EQT, Goldman Sachs and GIC investing for the first time. In February the company raised $500m in a Series D at an $11bn valuation.

A tender offer is existing holders selling shares to new investors, not new money into the company. Note: Tech.eu's text gives the amount as $300m in one line and £300m in another; the company's own post is the reference.

Sources:https://tech.eu/2026/09/30/elelvenlabs-doubles-valuation-to-22bn-with-300m-employee-tender-offer

Research

Apple study finds activation steering largely fails on instruction-tuned models

Quick take2 sources checked
2026-09-30 00:00 UTC

Researchers from Apple Machine Learning and Pompeu Fabra University found that activation steering methods are far less effective on instruction-tuned models than on their base counterparts.

Activation steering means guiding a model's output — such as removing toxic concepts — by intervening on internal activations without changing weights. The paper, published as a workshop paper at BlackBoxNLP 2026, also finds these efficient methods often impose a steep fluency cost, while prompting and full supervised fine-tuning work for concept injection but handle concept removal poorly.

The arXiv version was first submitted on June 10, and Apple's research page lists it as published in September 2026. The study is the authors' own evaluation; the abstract does not list the specific models or sample sizes, so extrapolation calls for caution.

Sources:https://machinelearning.apple.com/research/effectiveness-fluency-llm-conditioninghttps://arxiv.org/abs/2606.12234

Research

Blogger's tests suggest frontier models shift stated philosophy with who's asking

Quick take2 sources checked
2026-09-30 16:15 UTC

Tests by one blogger indicate that frontier models' stated philosophical positions shift with cues about who is asking.

Alex Kastner published on LessWrong on September 30 that when asked directly for their favorite decision theory, models almost always answer FDT/UDT (functional decision theory, the LessWrong-community mainstream); once the prompt hints the asker comes from mainstream academic philosophy, models including Claude Fable 5.1 answer CDT (causal decision theory, the academic mainstream) 30%-100% of the time. Each prompt was sampled 100 times, with data and code released.

Similar effects appear on moral realism and P(doom), questions with no human consensus. The author warns that attitude evals should be interpreted with user cues in mind.

The experiments were run by a single author, effect sizes vary by model, and no independent replication exists yet.

Sources:https://www.lesswrong.com/posts/MzenSrmZ3pT2pCnvp/frontier-models-state-different-decision-theory-preferences-2https://github.com/alexkastner/dt-audience-cues

Research

Manus ships 2.0, with efficiency gains measured by its own tests

Quick take
2026-09-29 01:35 UTC

Manus released version 2.0 on September 28, alongside Cue, a group-chat app for multiple agents — its first major update since parent Butterfly Effect restored independent operations on September 1.

According to the company's blog, 2.0 runs on its in-house Cascade framework. In one test configuration, compared with the previous system, token consumption fell 23.2%, task completion time shortened 28.2%, and running costs dropped 32%. These figures are vendor-measured, and the tested tasks are not specified.

Cue gives each agent its own email address, phone number and computer environment, and supports team collaboration; it remains in invite-only early access. This release is the overseas version, and the company says it is building a team for a China-market product.

Sources:https://zhidx.com/p/598008.html

Research

Meta denies Muse read private messages without permission; both sides hold firm

Quick take3 sources checked

Meta denies that its AI agent Muse read a user's private messages without consent; the incident has no independent technical verification.

Inc. columnist Jason Aten reported that Muse read his Mac messages while the required Full Disk Access setting was off. Meta VP of Communications Andy Stone replied on X that the Messages integration in the Mac Muse app is entirely opt-in: users must enable both Full Disk Access and the Messages connector before Muse can read message content.

David Singleton, an executive at Meta Superintelligence Labs, added that reading messages requires three separate steps of application-level permissions and built-in macOS system-level protections that "can't be circumvented even if the Muse application had a bug"; he called the "syncing device notifications" explanation Aten received an incorrect answer from the AI itself. Separately, another user said Muse leaked his address during a Facebook Marketplace task, and Singleton said he is looking into that case.

Sources:https://x.com/andymstone/status/2105106775259128080https://www.inc.com/jason-aten/metas-new-muse-ai-agent-read-my-private-messages-i-never-asked-it-to/91408202https://techcrunch.com/2026/09/30/meta-disputes-claim-that-muse-read-a-users-private-messages-without-permission

Research

Amazon's delivery glasses will photograph constantly, and customers can't opt out

Quick take
2026-09-30 17:21 UTC

Bloomberg reported on September 30 (as carried by The Verge) that Amazon's delivery driver smart glasses photograph their surroundings almost constantly while in use, potentially taking several thousand captures in a single driver's typical shift, which Amazon plans to upload to its AI platform, Wellspring.

Viraj Chatterjee, Amazon's delivery technology lead, told Bloomberg the data is only used to "enhance the delivery process" and that surveillance was never the intent; asked whether customers could opt out, he answered, "We haven't thought about that."

Amazon says its systems blur faces and license plates before human review, and that images could be obtained with a warrant, but it would not say how long images are stored; customers cannot view or request deletion of photos of their property. Amazon plans 20,000 more pairs in the field by the end of 2027, after pilot drivers completed over 275,000 deliveries.

Sources:https://www.theverge.com/tech/1002766/amazon-delivery-driver-smart-glasses-privacy

Research

Independent audit finds weight-decomposition explanations drift when aggregated

Quick take2 sources checked
2026-09-30 15:54 UTC

An independent audit finds Goodfire's weight-decomposition explanations drift clearly from the original model once aggregated.

Adversarial Parameter Decomposition (VPD) splits model weights into simple components and labels each as "needed here" or "safe to remove" per token. On September 30, Tom Angsten published an audit on LessWrong: after aggregating components for 64 tokens (about 4,500), output drift reached 0.80 nats, close to the 0.83 caused by the paper's own 20-step adversary.

Deleting every component never labeled as needed moved the model 1.28 nats in KL. The audit covers only the decomposition released with the published paper, not the newer unpublished training recipe in the authors' repository; the subject is a four-layer, 67M-parameter model, and whether this holds at production scale is unknown.

Sources:https://www.lesswrong.com/posts/KeBccWBGXnNXzZFBp/do-vpd-s-explanations-aggregate-an-audit-of-the-released-1https://github.com/angsten/vpd-audit

Research

Corrigibility fund adds $48,000 in prizes for recent papers

Quick take
2026-09-30 16:08 UTC

The Corrigibility Research Fund announced an additional $48,000 on September 30, rewarding roughly two dozen researchers across about a dozen teams working on corrigibility — making AI systems accept human correction.

Fund manager Max Harms handed out $27,000 in a first round in July and plans to disburse more than $60,000 in December. The largest single award, $14,000, went to Rubi Hudson's paper on a corrigibility transformation.

The selection is one person's judgment; Harms himself calls the awards ad hoc and says the purse sizes should not be taken too seriously, and he deliberately passed over established figures such as Yudkowsky and Christiano.

Sources:https://www.lesswrong.com/posts/3uJqhrC2idf4eNj5h/corrigibility-prizes-for-existing-work

Research

Apollo Research proposes principles for embedded evaluations; implementation decides their value

Material2 sources checked

Apollo Research, an AI safety evaluation organization, published its principles for embedded evaluations on September 30, arguing their impact depends on evaluator access, resources, and the weight findings carry in real decisions.

The core proposal is claim-based assessment: instead of an overall judgment of the developer, evaluators verify specific claims fixed in advance, such as "Models never attempted to disable or evade their monitoring during internal deployment." Each claim ends with one of five verdicts, insufficient access scores worst, a problem the developer reports itself scores better than one the evaluator finds, and the developer cannot veto the verdict.

The proposal also calls for public reports by default, with the evaluator's conclusion never redactable, and argues voluntary commitments are unlikely to suffice, so embedded evaluations should eventually be required by law. This is a unilateral proposal by the evaluator; no developer has yet said it will adopt it.

Sources:https://www.apolloresearch.ai/blog/principles-for-embedded-evaluationshttps://www.lesswrong.com/posts/jD32DPYdZEcyLZqto/principles-for-embedded-evaluations

Research

New architectures beating old models may not be the architecture's doing

Material

Long-term reading · 《A ConvNet for the 2020s》(2022)

On benchmark leaderboards, new Transformer-based image models have overtaken old-style convolutional networks that look at images through small local windows, and people credit the architecture. The paper modernizes the old convolutional network in the same way as its rivals, keeping only the local-view-of-image part. On the most commonly used image test set it reaches 87.8% accuracy, and it also surpasses the rivals in detection and segmentation.

Next time you see a claim that "a new architecture crushes old methods," first ask whether the training setups on both sides were aligned: amount of data, training duration, augmentation methods. Any gap from unaligned setups should first be chalked up to training investment, before discussing which architecture is better.

If the task is not this kind of standard image benchmark — for example, very high-resolution inputs — don't use it to judge architectures. The paper also doesn't prove that the unmodernized old network was competitive all along; what's competitive is the modernized one.

A ConvNet for the 2020s (2022) | Next review 2027-09-20

Sources:https://arxiv.org/abs/2201.03545v2

Research

OpenAI launches a fast decisions API aimed at cheap agent classification

Material3 sources checked
2026-09-29 10:00 UTC

OpenAI announced a limited preview of its Decisions API at DevDay on September 29: the Luna model picks quickly from a predefined set of options.

CEO Sam Altman said focusing the model on that choice makes it extremely fast while keeping image understanding, broad language support and safety protections. It closely resembles Jev, released by TypeSafe AI earlier this month — a fast classifier that outputs probabilities, which developers already use to make LLM pipelines faster and cheaper.

One direct application is agent monitoring: QueryStory's Shapor Naghibzadeh built a hackathon demo using Jev to check each agentic action, saying the same review would cost $2.94 with Jev versus $372 with a frontier LLM. Decisions API remains a limited preview with no developer benchmarks yet, and how well these models' outputs are calibrated is the key open question.

Sources:https://openai.com/index/devday-2026-recaphttps://typesafe.ai/blog/introducing-system-one-models-and-jevhttps://techcrunch.com/2026/09/30/openais-jev-clone-could-help-the-frontier-lab-stop-its-swarming-agents

Research

You’re all caught up in this view