ReadingResearchRadarInvestment framework
Sign in / Sign up
Sign in / Sign up
ReadingResearchRadarInvestment framework
Reading archive →

Reading

2026-09-2456 posts

CoAx locates backup heads at 0.941 AUC, still unreplicated

Quick take
2026-09-23 12:00 GMT+8

Readers doing circuit analysis or safety evaluation now have a new idea to use: conditional co-ablation (CoAx), which after ablating a primary component set ranks candidates by growth in ablation effect to find backups the intervention activates — the authors report 0.941 ROC-AUC recovering known backup heads on GPT-2's IOI circuit.

Previously, recovering backup heads masked by primary components relied on intact-state attribution, whose strongest baseline reached only 0.815 ROC-AUC on GPT-2-small's IOI circuit and could not directly complete an incomplete circuit.

The authors' reported measurement: CoAx recovers the documented backup heads on that IOI circuit at 0.941 ROC-AUC, and adding the selected heads to the incomplete circuit cuts incompleteness from 0.75 to 0.21; they also report it beats matched random completions on 8 non-GPT-2 models across 6 architecture families.

All numbers are first-party evaluations, backup-head recovery is mainly validated on the known-answer IOI setting, and nothing has been independently reproduced; preprint arXiv 2607.01940 (v3, 23 Sep 2026).

Sources:arxiv.org

Research

VERPO's token-level credit assignment avoids GRPO training collapse

Quick take
2026-09-22 12:00 GMT+8

VERPO decomposes teacher guidance into token-level credit assignment, avoiding the optimization collapse caused by GRPO's sequence-level advantage estimation, and self-reports the highest multi-task average across five scientific reasoning and tool-use tasks.

Previously GRPO estimated advantages from whole-sequence scores, and the authors say this sequence-level granularity triggers optimization collapse.

The authors self-report larger gains on smaller models; this is a first-party benchmark with no independent reproduction.

The method includes a Fisher movement-cost controller and an FEC projection; version 4 was updated on September 23, and third-party validation is still pending.

Sources:arxiv.org

Research

Apple's probe guidance steers diffusion language models without an extra forward pass

Quick take
2026-09-23 08:00 GMT+8

Readers can now learn of a way to guide diffusion language models without adding inference cost: probe guidance, introduced by Apple's machine learning research team, builds a guidance signal from the frozen internal states of an existing diffusion model, eliminating the extra inference forward pass that autoguidance requires, and the team says that applied to a 1.7B diffusion language model it consistently improves multiple-choice question answering benchmarks.

Previously, autoguidance required an extra forward pass and had lacked an account of why it works.

The team reports the method sets a new state of the art on unconditional generation for continuous diffusion language models. These results are first-party benchmarks.

The work also reports a mechanism finding: the weak model in autoguidance must come from a low-entropy region of training. If it holds, this may be more useful for diffusion language model research than any single benchmark number. The results have not been independently reproduced and come from a paper submitted on 2026-09-23.

Sources:machinelearning.apple.com

Research

Radical Numerics says its genome model can continue an aptamer optimization trajectory

Quick take
2026-09-23 21:27 GMT+8

Radical Numerics CEO Eric Nguyen said in the show notes of the September 23 Latent Space episode that his genomic language model, in an aptamer experiment, recapitulated some held-out high-scoring samples after seeing only low-scoring RNA sequences.

Per the episode's show notes, the experiment held out the best-performing aptamers, showed the model only lower-scoring sequences in a rising series, and the model continued the trajectory on its own. Nguyen calls this chain-of-thought "thinking in DNA."

This is a first-party claim: the dataset size, scoring method and full results are not published with the notes, and it should not be read as independently verified performance. Watch for a checkable paper or benchmark.

Sources:latent.space

Research

Ringg self-reports 7M monthly AI-handled calls; 90% cost cut is vendor-claimed

Quick take
2026-09-23 20:00 GMT+8

OpenAI published a customer story on September 23 saying that Ringg, an Indian voice support agent platform, now handles more than 7 million connected calls per month, with up to 65% of routine requests resolved without human involvement.

The story says Ringg cut model costs by roughly 90% after migrating selected real-time workloads from GPT-4.1 to GPT-5.6, and that at Policybazaar 67% of connected calls are handled without humans, with average response time falling from 8–12 minutes to under 60 seconds.

All resolution rates, cost savings and the 4.8 CSAT are self-reported by Ringg and OpenAI with no independent verification, and the 65% figure is an upper bound, so typical performance may be lower.

Sources:openai.com

Research

Maoxiang lead leaves ByteDance to bet on AI generative entertainment

Quick take
2026-09-24 07:00 GMT+8

Liang Chenqi, who led the AI interactive-content product Maoxiang, left ByteDance in the summer of 2025 to found the AI entertainment startup Dongnian Yinxian.

Per the show notes of LateTalk episode 182, he spent seven years at ByteDance and built Maoxiang within its AI product unit Flow from 2023; his new company is betting on a blend of generative creation, expression and light gaming, and trains its own models.

His thesis: once people no longer work for a living, casual creation becomes a mainstream, high-frequency form of entertainment. No AI entertainment product has yet become an undisputed hit, and the startup has disclosed no funding or user figures; the above are his own statements.

Sources:podcast.latepost.com

Research

Models can be trained jointly without data leaving the device, but it was only validated in simulation

Material
Verified 2026-10-03 18:32 GMT+8

Long-term reading · 《Communication-Efficient Learning of Deep Networks from Decentralized Data》(2016)

Some data is scattered across a large number of phones, and privacy does not allow it to be centralized on a server. The approach in the FedAvg paper is: each phone trains locally for a few rounds on its own data, and only the model changes are sent up to be averaged, without transmitting raw data, and the number of communication round trips is ten to a hundred times fewer than centralized training. This type of approach is collectively called federated learning.

Today, when vendors advertise that "the model can be improved without uploading data," what runs underneath is still this local training plus averaging, and later improvements are patches on top of it. When you hear this kind of advertising, first ask whether the results were measured in simulation or on real phones; the original paper itself only went as far as simulation.

If devices go offline, if data changes, or if someone deliberately submits bad updates, it handles none of these. If the question is whether this kind of training leaks privacy, or how well it works on real devices, it was only validated in simulation, so don't use it to draw conclusions.

Communication-Efficient Learning of Deep Networks from Decentralized Data (2016) | Next review 2027-09-20

Sources:arxiv.org

Research

2026-09-2366 posts

Both frontier labs cut prices the same day; intelligence keeps getting cheaper per token

Material
2026-09-23 17:00 GMT+8

Anthropic and OpenAI released new model tiers about 90 minutes apart on September 22, and both cut prices sharply: frontier-level intelligence keeps falling in price per token.

Anthropic's Claude Opus 5.5 is priced at $4 input and $20 output per million tokens, with cache reads at $0.20; the company says typical workloads cost 40% less than Opus 5, and cache reads 60% less. OpenAI's GPT-6 Sol is 50% below GPT-5.6 pricing at $2/$10, and Luna at $0.10/$0.50 per million tokens.

All performance-leadership claims are vendor-reported: Anthropic itself concedes that benchmark margins have become a less reliable guide at this capability level, and OpenAI's AutomationBench comparisons use its own framing. Treat the price cuts as verified facts and the rankings as claims awaiting independent checks.

Sources:therundown.ai

Research

Opus 5.5 cuts prices 40% and becomes the default, but heavy users save little

Material
2026-09-23 14:41 GMT+8

Anthropic shipped Opus 5.5 and made it the default model, with list prices down 20% — but real costs at high effort barely moved.

Launched September 22, the model is claimed to match Fable 5.1 on most tasks and run 40% cheaper than Opus 5. Input/output pricing fell from $5/$25 to $4/$20 per million tokens, and it is now the default in Claude Code and the Claude app.

Artificial Analysis, however, measured that higher token usage at max effort would raise per-task cost about 80%; after the price cut and cheaper caching it lands at $5.98 per Intelligence Index task, essentially flat versus Opus 5's $5.86. The 40% saving holds only at default effort.

Performance figures are vendor and partner numbers, and the unconfirmed claim that the model is smaller remains speculation; test cost on your own workload and effort setting before switching.

Sources:latent.space

Research

Anthropic ships Opus 5.5 with cache-read prices cut 60%

Material
2026-09-23 06:58 GMT+8

Anthropic released Claude Opus 5.5 on September 22, cutting input to $4 per million tokens and output to $20, with cache reads down 60% to 20 cents; the faster serving mode runs $8 and $40.

The model is live on the Claude Developer Platform and on AWS, Google Cloud and Azure, with Sonnet 5.5 and Haiku 5.5 due in coming weeks. OpenAI released GPT-6 Sol and Luna at half price the same day, so the frontier price war moved at both labs at once.

Benchmark figures such as 66.4% on Terminal-Bench 4.0 are Anthropic's own, with no third-party verification yet; for teams running long-session agents, the cache price cut is the most certain cost change available today.

Sources:siliconangle.com

Research

Anthropic ships its new flagship 20% cheaper, performance claims still self-tested

Material
2026-09-23 00:30 GMT+8

Anthropic released Opus 5.5 on September 22, cutting output token pricing to $20 per million tokens from $25 for the previous model, a 20% drop.

The company calls it the strongest-performing model it has tested and says it outpaces its own larger Fable model on many benchmarks. Those results are all vendor-run; METR and other outside groups did pre-release safety evaluation only, and no independent performance verification exists yet.

Anthropic says the model is comparable to Mythos in biology and cybersecurity capabilities, so it carries the same usage safeguards as Fable. Sonnet 5.5 and Haiku 5.5 are promised in the coming weeks; whether the price cut extends to mid-tier models is what buyers should watch.

Sources:techcrunch.com

Research

OpenAI's new Sol and Luna ship at half price, reliability claims still self-tested

Material
2026-09-23 02:00 GMT+8

OpenAI released GPT-6 Sol and Luna on September 22, pricing the API at half the cost of the 5.6-series equivalents, which the company attributes to improvements in caching and inference.

Sol targets complex tasks like coding; Luna handles high-volume clerical work such as summarizing and extraction. Both are live in ChatGPT Work, Codex and the API, and Luna will also reach Free and Go users.

The reliability claim deserves a discount: the reported halving of mistakes comes from OpenAI's internal evaluation based on user-flagged conversations, not third-party testing. Anthropic shipped Opus 5.5 just 90 minutes earlier, and the pricing race between the two is clearly deliberate.

Sources:techcrunch.com

Research
Next reading page →