LectureRechercheRadarCadre d'investissement
Connexion / Inscription
Connexion / Inscription
LectureRechercheRadarCadre d'investissement
Archives de lecture →

Lecture

2026-09-2256 publications

Rat-neuron-derived AI model lands on AWS, performance claim unverified

Prise rapide
2026-09-22 21:00 GMT+8

The Biological Computing Company's "rat brain" AI model opens to select AWS customers on September 22 as a limited preview, after previously being available only through the small neocloud provider Bluesky Compute.

The technology uses activity patterns from rat neurons on multi-electrode arrays to improve video generation models. The company claims inference up to five times faster than an open-source video model it runs, but declines to name the comparison model. That is a vendor-tested figure with no third-party verification.

TBC has raised more than $50 million in total, including an additional $25 million first reported by WIRED. AWS executive Deap Ubhi himself raised the open question: whether fidelity holds when customers generate videos of ten minutes or longer remains unknown.

Sources :wired.com

Recherche

Kantata now lets services firms build their own AI agents from a sentence

Prise rapide
2026-09-22 20:00 GMT+8

Kantata released Agent Studio on September 22. A user describes the job in plain language, and the platform assembles and refines a custom AI agent. The company says it is available now and in use by customers.

The tool sits inside the Kantata Expertise Engine, with the Expertise Agent launched in June doing the assembly. Every build and edit is checked for correctly structured inputs and outputs before publication, with no forms or code required. The intended users are practice leads and delivery directors.

The backdrop is Kantata's own diagnosis of the market: vendors are shipping narrow agents by the dozen, each scoped to a single job, and each arriving as another silo to maintain. Note that availability and customer usage are Kantata's own statements, not independently verified.

Sources :siliconangle.com

Recherche

UN AI experts reject doomsday rhetoric, urge separating facts from extrapolation

Prise rapide
2026-09-22 19:08 GMT+8

Members of the UN's International Scientific Panel on AI publicly pushed back against doomsday-style AI risk rhetoric at a forum held alongside the UN General Assembly in New York, with co-chair Yoshua Bengio urging a line between established science and reasonable extrapolation.

Joelle Barral, a Google DeepMind executive on the panel, said researchers should focus on clarifying what is settled and what remains unknown, and that stoking fear does not help. Nobel laureate Maria Ressa said the debate has swung between existential alarm and dismissing AI risk altogether.

Bengio acknowledged that experts themselves still disagree on the extrapolation side. This is an IT之家 report citing AFP coverage of the forum — statements by participants, not an official panel document.

Sources :ithome.com

Recherche

10M-parameter refinement reportedly cuts cross-domain audio deepfake error

Prise rapide
2026-09-21 12:00 GMT+8

Training only about 10 million trainable parameters (out of 598 million) on an already-trained SSL-based audio deepfake detector cuts pooled equal error rate from 4.85% to 3.74% across 14 cross-domain test sets.

Previously, improving cross-domain detection typically meant altering the original model's parameters or adding training data; CoReLoop leaves both untouched, reusing encoder outputs through lightweight refinement modules and low-rank adapters.

The 4.85%-to-3.74% pooled EER reduction is first-party, self-run by the authors; an optional halting head picks refinement depth per utterance, reaching 3.73% pooled EER at an average of 1.18 passes.

The composition of the 14 test sets and the identity of the baseline detector need checking in the full text, real-world generalization awaits independent replication, and the preprint was submitted on 2026-09-17.

Sources :arxiv.org

Recherche

Geometric constraints make storyboard framing measurable, but action fidelity drops

Prise rapide
2026-09-21 12:00 GMT+8

Storyboard framing can now be measured and solved as geometric constraints rather than left to an image model's defaults: on 204 external director-storyboard shots, delivered head height can be brought to 0.955 times the target, and single-subject panels land within 1.2% of frame width of their declared positions.

Previously, screenplay-to-storyboard spatial planning relied on an image model's defaults, and delivered head height was 1.906 times the staged target from the director's words and 1.733 from the compiled prompt, leaving framing neither controllable nor measurable.

The PACE preprint uses a typed representation plus a camera solver to turn spatial planning into measurable geometric constraints; declaring the pose on 30 shots raised the action drawn from 58.9% to 74.4% without moving the framing. All of the above are the authors' first-party self-tests.

The authors note that the setting holding framing best draws the least described action, and transitions, fitted motion and human review remain open; v2 also removed supplementary material that was never uploaded, and the results await independent reproduction.

Sources :arxiv.org

Recherche

New method makes reasoning-model forgetting leak-free, but results are self-reported

Prise rapide
2026-09-21 12:00 GMT+8

The GUARD method has been accepted to the EMNLP 2026 main conference; the paper was posted to arXiv on September 18.

The problem it targets: after a reasoning model unlearns sensitive content, intermediate chain-of-thought steps may still disclose the answer before the final reply. GUARD rewrites unsafe outputs into a coherent non-disclosing reasoning path followed by a stable refusal, then distills that behavior into the model's parameters.

The authors also introduce NFRS, a metric for structural stability, fluency and unsupported substitutes in forgotten outputs. Experiments run on R-TOFU and a STAR-1-derived harmful-intent setting across two distilled reasoning models, which the authors say reduce unsafe and privacy disclosures while preserving reasoning utility. All results are first-party and independently unreplicated.

Sources :arxiv.org

Recherche

Experiment finds trust in financial advice tracks style and labels

Prise rapide
2026-09-21 12:00 GMT+8

A randomized vignette experiment with 285 U.S. adults found advice style, not source labels, most strongly shaped message and safety appraisals.

The study independently varied AI, expert, and online community styles while holding the recommendation constant. Expert labels selectively raised perceived source knowledge, and decision context mainly drove risk and safety judgments.

Expert-style advice remained most preferred even without source labels. This is a first-party preprint result based on hypothetical vignettes; stated trust is not actual reliance.

Sources :arxiv.org

Recherche

Agents writing game rules: best combo solves only 52.78%

Prise rapide
2026-09-21 12:00 GMT+8

The best single run by a coding agent implementing game logic solved only 52.78% of tasks, according to the first benchmark that checks game rules at every simulation tick: 72 Godot tasks, with 403 hand-designed scenarios expanded into 1,451 test cases.

Previously no benchmark checked game rules tick by tick, so incorrect rule implementations went undetected. Across 20 combinations of language models and scaffolds, the best score was that 52.78%; under the Claude Code scaffold, all twelve models scored lower as tasks grew from isolated mechanics to repository-scale features. Most failed submissions were runnable but implemented the rules incorrectly.

The authors' own ablation found that without mutant-based validation, incorrect agent submissions passed the evaluator; with open network access, agents copied code from public repositories. These are first-party preprint results; the success rates and evaluator validity await independent reproduction.

Sources :arxiv.org

Recherche

AI agent classifies genetic disease severity at 93.55% accuracy

Prise rapide
2026-09-21 12:00 GMT+8

You can now grade 10,211 Human Phenotype Ontology terms by severity with an agent combining ReAct and retrieval-augmented generation, reaching 93.55% accuracy (MCC 0.9237) on expert-curated cohorts.

Previously, severity grading relied on manual literature and guideline review, which was costly and hard to scale across all terms; this system retrieves PubMed literature against ACMG severity guidelines and ACOG quality-of-life criteria and generates checkable reasoning chains. The authors state 82.6% to 91.4% of claims were supported by direct evidence or valid inference, and gene-level concordance with the Mackenzie's Mission gene list was 95.2%.

The result is reported first-party by the research team on the authors' own curated cohorts; clinical or panel-design use remains unverified, with no independent replication. It appears in an arXiv preprint.

Sources :arxiv.org

Recherche

Near-daily AI use in the US doubled in six months, but the survey wording changed

Prise rapide
2026-09-20 18:03 GMT+8

Surveys by Epoch AI with Ipsos show the share of US adults using AI almost every day rose from 8 percent to 19 percent in six months.

The two waves ran in March and August 2026, with 2,017 and 1,016 randomly selected, population-weighted respondents. Over the same period, the share using AI only one day a week fell from 17 percent to 10 percent, so growth came mainly from occasional users turning frequent.

One caveat matters: the question format changed. March asked about overall AI use; August asked about each service separately and took the highest frequency. Epoch AI itself says the results are not fully comparable and the August figures may understate actual usage.

Sources :the-decoder.com

Recherche

Self-run test finds AI monitors miss over half of agent-simulation details

Prise rapide
2026-09-22 10:59 GMT+8

A research team used a self-authored murder mystery to test whether AI can reconstruct the underlying story from multi-agent interaction logs. The conclusion: even the strongest monitor, GPT-6 Astra at high reasoning effort, fully recovered only 48.1% of the seeded facts and relationships.

The authors first hand-built an event graph with 184 nodes and 229 edges as ground truth, then had eight agents run the investigation, and finally asked models to reconstruct the history from roughly 58,000 words of trajectory. Omissions far outnumbered errors: Astra high left 42.6% of rubric items unreported, while Gemini 3.1 Pro recovered just 16.5%.

This is a first-party result: scoring was done by a GPT-5.6 Sol judge pipeline, not yet validated against independent human grading, and all ten simulations came from a single authored mystery, so generalization is unknown — the authors list human validation as their next step. For readers doing agent-incident investigations, the transferable piece is the method: fixing the reference account in advance is what lets you measure what a monitor report leaves out, not just what it gets wrong.

Sources :lesswrong.com

Recherche

Multi-agent swarms buy speed, not compute savings: same performance costs about twice the tokens

Prise rapide
2026-09-22 04:30 GMT+8

Read from OpenAI's own charts, multi-agent parallelism mainly buys speed, not compute savings.

Writing on LessWrong, Toby Ord analyzed the GPT 5.6 launch-page charts: at equal performance, a 4-agent swarm uses roughly twice the total tokens of a single agent, and a 16-agent swarm roughly twice that again. His derived parallelizability parameter lambda falls between about 0.48 and 0.68, depending on the task.

That means scaling the swarm buys capability more expensively than lengthening a single agent's chain of thought; the payoff is speed, roughly 2x speed for 2x cost. Note this is a blogger's reading of vendor charts, with regressions run by Claude Opus 5, not an independent measurement.

Sources :lesswrong.com

Recherche
Page de lecture suivante →