LectureRechercheRadarCadre d'investissement
Connexion / Inscription
Connexion / Inscription
LectureRechercheRadarCadre d'investissement
Archives de lecture →

Lecture

2026-09-2933 publications

Momentic launches Mo, an AI testing agent that needs no test scripts

Prise rapide
2026-09-29 00:00 GMT+8

Software testing company Momentic launched Mo on September 28, an AI agent that tests applications directly from a developer's instructions, with no test scripts to write or maintain.

Mo works like coding tools such as Claude Code: give it a URL and a testing brief, and it spins up a swarm of agents to operate the app, trying thousands of edge cases, then reports reproduction steps and video for each bug. It can also read a product requirements document or a Jira ticket for context.

Co-founder and CEO Wei-Wei Wu told SiliconANGLE that as long as scripts sit in a codebase, someone has to maintain them. The company says customers including Notion have been trying Mo; there is no independent evaluation yet.

Sources :siliconangle.com

Recherche

Security startup Rig raises $12M to police AI agents running under employee accounts

Prise rapide
2026-09-29 20:00 GMT+8

Israeli security startup Rig Security launched on September 29 with $12 million in seed funding.

The blind spot it targets is specific: coding assistants and autonomous agents rarely hold accounts of their own, so their actions are logged under the employee or service account that launched them. In the words of CEO Guy Kozliner, agents are now the most active identities in an enterprise, "and they are acting under human names" — security teams cannot tell person from agent, and stopping the agent means blocking the employee too.

The product uses machine learning to match one identity across identity systems, with the company itself claiming better than 96% accuracy; an endpoint sensor can stop an agent's risky action without interrupting the employee. The company says Fortune 200 customers in financial services, insurance and health care already run it in production — both that claim and the accuracy figure are the company's own. The round was co-led by Ten Eleven Ventures and Brightmind Partners, with CrowdStrike participating as a strategic investor.

Sources :siliconangle.com

Recherche

Four pharma giants join Apheris consortium to train AI on 10,000 antibodies

Prise rapide
2026-09-29 20:00 GMT+8

AbbVie, argenx, Lundbeck and Takeda have joined Berlin-based Apheris and Ginkgo Datapoints in a new Antibody Developability Consortium, training AI on roughly 10,000 antibody sequences to predict which drug candidates can be manufactured and reach the clinic.

Each member contributes proprietary sequences and trains models inside its own environment via Apheris's federated infrastructure, without raw sequences being exposed to other members. Ginkgo Datapoints fills remaining capacity from public sources and runs wet-lab characterisation; the initial dataset is due to members by early 2027.

The consortium has just kicked off and has released no model performance data. Charlotte Deane of Oxford and Peter Tessier of the University of Michigan provide independent scientific oversight.

Sources :tech.eu

Recherche

US Navy stands up drone command to integrate air, sea and undersea tactics

Prise rapide
2026-09-29 22:00 GMT+8

The US Navy formally established the Robotic and Autonomous Systems Warfighting Development Centre in Virginia on September 24, tasked with turning its aerial, surface and undersea drones into an integrated fighting force.

The centre will develop doctrine and tactics for unmanned systems, train sailors, and test whether the platforms keep operating under realistic conditions. Chief of Naval Operations Admiral Daryl Caudle said years of robotic-systems work had proceeded separately across domains, that the next step is integration, and that "a prototype is not combat power."

US commanders have previously linked this drone capability to disrupting a PLA attack on Taiwan. The centre's budget and staffing have not been disclosed; its integration progress remains to be seen.

Sources :scmp.com

Recherche

Musk says Cybercab won't run commercial rides in California until mid-2027

Prise rapide
2026-09-29 17:37 GMT+8

Elon Musk expects Tesla's Cybercab to start commercial operation in California only around mid-2027.

In an interview with China Global Television Network released on Saturday, September 26, Musk said the two-seat, steering-wheel-free vehicle "will soon be operating commercially in Florida, in Nevada, and a couple of other states," and that in California it will "probably" be running "by the middle of next year."

Tesla had earlier promised to sell the vehicle to regular buyers for $30,000 before the end of 2026 — the promise behind YouTuber MKBHD's head-shaving bet. The car still faces an NHTSA audit query into Tesla's self-certification that it meets federal safety standards, a hard gate before any wider rollout.

Sources :businessinsider.com

Recherche

Alignment is purpose-agnostic — and works as a censor's toolkit

Prise rapide
2026-09-29 04:07 GMT+8

The same alignment techniques built to make models safe can just as easily censor or distort information: alignment methods are purpose-agnostic, making a model serve someone's will with nothing in the methodology guaranteeing good intent.

Alignment has largely been assumed to be a safety measure; Sarah Ball and Phil Hackemann sort control into three levers — pretraining data filtering, post-training alignment, and inference-time intervention — with cost falling and ease of change rising down the stack, making lower layers easier to abuse unilaterally.

This is not hypothetical: China's cyberspace regulator requires providers to maintain refusal datasets, and Elon Musk publicly said he would "fix" Grok outputs he disagreed with, with reported behavior shifts traced to system-prompt changes. The analysis won an Outstanding Position Paper Award at ICML 2026; the authors do not call for stopping alignment, but back transparency, verifiable alignment, and model pluralism.

This is a position paper's argument plus public examples — author-reported views, with no third-party replication of the quantified risk.

Sources :lesswrong.com

Recherche

Pinker's open letter calls AI extinction fears overblown, backs sober safety engineering

Prise rapide
2026-09-28 23:32 GMT+8

Harvard psychologist Steven Pinker wrote in a September 26 open letter on Quillette that fears of AI wiping out humanity are overblown, and declined psychiatrist and influential tech blogger Scott Alexander's invitation to a public debate, calling such events a "spectator sport."

Pinker admits he underestimated how capable large language models would become and how recklessly AI companies would release their products. The dangers he takes seriously are bioterrorism, cyberattacks and unchecked AI agents; what works, he writes, is safety engineering built on independent oversight, liability and human control.

Alexander had put the odds of AI wiping out humanity at 20 percent on his blog in June, calling for better safety research and international agreements to slow development.

Sources :the-decoder.com

Recherche

Computation shifts monotonically deeper as skill setting rises

Prise rapide
2026-09-22 12:00 GMT+8

A frozen-weight chess model, Maia-3, shifts its computation monotonically deeper as the skill setting rises — with Elo dialed from 700 to 2500, the causal center of computation moved later for every piece type, most of all for knight forks.

The direction contradicts the plausible prediction, suggested by Princeton professor Tom Griffiths, that higher skill should compute key features earlier. Maia-3 is an 8-layer transformer that takes an Elo rating as input to mimic human players of different strengths, and author David Litman made the measurements without changing any weights.

The author notes that circuits found under one condition may not stay in place under another. The result comes from a preprint and the author's own analysis tooling.

Sources :arxiv.org

Recherche

To ask how confident a model is, have it answer several times and see how scattered the answers are

Matériel
Vérifié 2026-10-03 18:30 GMT+8

Lecture à long terme · 《Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning》(2015)

There is an old trick when training neural networks that make judgments automatically: at each step, randomly and temporarily turn off some units, known in jargon as Dropout. A 2015 paper explained mathematically that if you keep this randomness at test time and compute the same question several times, the disagreement among the answers is the model's uncertainty about itself.

Today models make decisions for people everywhere, and when they are wrong they still answer with certainty, so how much uncertainty they themselves have has become an unavoidable question. Later methods mostly require training several sets of models, while this one needs no retraining and only a few extra computations, and it is still the default starting point today. When vendors claim a model knows what it does not know, first ask whether it is done this way.

If this randomness was not turned on during network training, do not use it to judge confidence — there is no randomness to keep, and the method cannot be used. It has only been validated on small tasks such as predicting numerical values and handwritten digits. When the several answers scatter into several clusters, the confidence it gives is only a rough approximation, so do not treat it as a precise probability of error.

《Dropout as a Bayesian Approximation》(2015) | Next review 2027-09-20

Sources :arxiv.org

Recherche

Vous êtes à jour dans cette vue