読書リサーチレーダー投資フレームワーク
ログイン / 新規登録
ログイン / 新規登録
読書リサーチレーダー投資フレームワーク
アーカイブ →

読書

2026-09-2456 投稿

African team trains small model from scratch, self-reports benchmark lead

クイックテイク
2026-09-24 05:25 GMT+8

Vambo AI has released MORENA, a 1.5B-parameter model trained from scratch for 12 African languages, with self-reported benchmark results ahead of 26 tested models.

The developers report 1.408 bpb on African language modelling, the best of 26 models tested; the closest rival, Lugha-Llama-8B, carries over five times the parameters and scored 1.423 bpb. All figures are developer-reported, relayed via TechRadar, with no independent replication yet.

The mechanism difference is a custom vocabulary: African text encodes with 1.39 times fewer tokens than Gemma 3, though it still costs about 6% more tokens per byte than English, a gap the team says it cannot fully explain. Training took roughly 22,000 A100 GPU hours, with compute support from UNDP and others.

ソース:techradar.com

リサーチ

Microsoft opens Defender security operations preview built around agents

クイックテイク
2026-09-24 06:10 GMT+8

Microsoft opened the public preview of its security operations center (ISOC) on September 23, folding the log-and-alert analysis (SIEM) features of its analytics platform Sentinel directly into Defender.

The preview is available to customers with Microsoft Defender Suite, Microsoft 365 E5 or E7 licenses; organizations already running an active Sentinel workspace are excluded for now. Case management and workbooks work without setup, while user and entity behavior analytics and third-party data ingestion require extra configuration, and ingestion charges may apply.

The design gives security agents the same signals and controls as a human analyst, letting them investigate and act on incidents, with humans still approving high-stakes actions. Microsoft has not disclosed pricing, and Defender data is kept for 30 days at no charge during the preview. The claim that attackers are putting agents to work comes from Microsoft's own executives, not independent verification.

ソース:siliconangle.com

リサーチ

Neo4j makes binary quantization the default, cutting vector search memory

クイックテイク
2026-09-23 19:06 GMT+8

Starting with Neo4j 2026.09, rescored binary quantization is the default for newly created vector indexes in Aura, Enterprise Edition and Community Edition.

In vendor-run benchmarks, Neo4j claims roughly five times better throughput and query latency than 2026.08, and about four times as many 768-dimensional Float32 embeddings fitting in the same memory budget. The approach searches compressed binary vectors first, then rescores candidates against full-precision vectors read from disk.

All figures are Neo4j's own first-party benchmarks; the methodology post has not been published and nothing is independently verified. The technique draws on RaBitQ from NTU and the BBQ implementation in Lucene.

ソース:neo4j.com

リサーチ

Vercel Connect upgrade: AI agents call external tools without handling credentials

クイックテイク
検証済み 2026-09-24 08:46 GMT+8

Vercel announced in its September 24 changelog that Connect now supports TanStack AI: agents built with TanStack AI can call OAuth-protected MCP servers, with no credentials for developers to store or rotate.

A new @vercel/connect/tanstack-ai subpath exports connectMCPTransport, which takes a TanStack transport config and attaches a Connect-backed auth provider. The provider is called before every MCP request, so the token is always fresh.

One detail worth noting is where authorization failures surface: if the user has not granted access, createMCPClient fails with a consent challenge before the model runs, and developers can catch it with getConsentChallenge and redirect to Connect's consent URL; otherwise the error would reach the model as an error string rather than the user as a redirect. This is a vendor-reported product capability from Vercel's own changelog, with no independent usage data yet.

ソース:vercel.com

リサーチ

Accenture moves the AI agent trust boundary down into the Oracle database

クイックテイク
2026-09-24 08:12 GMT+8

Accenture's Oracle practice constrains what AI agents can read at the database rather than the application layer, serving roughly 40 agents.

The team runs an MCP server connected to more than 925,000 documents across 27 Oracle product namespaces. Agents hold different authorization levels; policies are defined in SQL and enforced by the database on every query, regardless of what SQL the agent constructs.

The team says it has tested scale to 300 simultaneous queries, with most results returned in 15 seconds or less; that figure is self-reported and independently unverified. The case comes from an interview at an Oracle-sponsored event where theCUBE is a paid media partner, so readers should treat performance and effect claims as first-party statements.

ソース:siliconangle.com

リサーチ

Blogger's own test: Jev monitors chain-of-thought 6x faster but misses more harm

クイックテイク
2026-09-24 07:59 GMT+8

A LessWrong author's self-run benchmark reports that a Jev-format model monitoring chain-of-thought for harm averaged 542 ms per classification, about 6x faster than Claude Sonnet 5 at 3,348 ms, and cost $0.056 per 1,000 classifications versus $3.17 for Sonnet, roughly 566x cheaper.

This is a first-party benchmark the author ran on 2,200 traces from the ReasoningShield dataset; it has no independent reproduction. On exact-match accuracy Jev scored 71.5%, essentially tied with Sonnet's 71.4% and below GPT-5.6 Luna's 79.4%.

The key weakness is on the harmful side: the author himself notes Jev was worse than both Sonnet and Luna at identifying genuinely harmful traces, and that catching harm matters more than fewer false positives. The headline multipliers hold only for the harmless side.

ソース:lesswrong.com

リサーチ

Persona selection model author self-review: stop extrapolating risk from it

クイックテイク
2026-09-24 09:21 GMT+8

Sam Marks, who coined the persona selection model, has published a self-assessment: the hypothesis is over-applied and its predictions are narrow.

The model holds that during pre-training LLMs learn to simulate diverse human-like personas, and post-training elicits one 'Assistant' persona; it is often used to infer whether AIs will seek reward or how high takeover risk is.

In a September 24 LessWrong post he argues those inferences do not follow — reward-seeking, approval-seeking and alignment faking are all compatible with human-like personas. These are his views, not new experimental results.

His real update is that personas are more conditional than he expected: reward-seeker on scored tasks, 'good person' in chat. He also says there is no strong evidence yet that large amounts of RLVR (reinforcement learning with verifiable rewards) break the model.

ソース:lesswrong.com

リサーチ

Anthropic engineer says Claude's writing worsened because models learned to write for AI

クイックテイク
2026-09-23 23:14 GMT+8

Jackson Kernion, an Anthropic engineer working on Claude fine-tuning, attributes the regression in Claude's writing to the reinforcement-learning reward structure: some rewards optimize for comprehension by other AI models, and the more training leans on math and code, the more the model learns to write for AI rather than people, producing what humans experience as "overly-dense info dumps."

Per The Decoder's relay of his X posts, he compares the resulting "Claudeish" style to habits formed inside a closed communication group, and says Opus 4.6 remains Anthropic's last good writing model.

He says Opus 5.5 strikes a better balance but does not claim it surpasses the older model. These are his personal statements, with no evaluation data behind them.

ソース:the-decoder.com

リサーチ

Rabbit ships OS3, shifting from R1 hardware to cross-device agents

クイックテイク
2026-09-24 00:20 GMT+8

Rabbit unveiled OS3 on September 23, a personal AI agent that runs in the cloud and connects to users' local devices.

OS3 installs on Windows, macOS and Linux machines with a single command; the company said it can be installed on up to five machines, and the R1 works with it but is not required. The agent is model-agnostic but users must bring their own key for Anthropic, OpenAI or a local model, or connect via aggregators like OpenRouter, and it can swap models without losing context.

Rabbit stressed that OS3 acts only on user instructions, sensitive actions require locally requested system permissions, and permissions can be revoked at any time. All capabilities are company-reported with no independent evaluation; the field already includes Meta Muse, OpenClaw and other rivals, so real-world performance remains to be seen.

ソース:siliconangle.com

リサーチ

OPPO turns AI wearables into a 499-yuan lapel accessory

クイックテイク
2026-09-22 19:00 GMT+8

At its Find X10 launch on September 22, OPPO introduced the 499-yuan AI wearable "Xinliqiu" (Heart Sphere), a lapel-clipped device focused on all-day recording.

Per the launch presentation, it identifies important items and offers proactive reminders, with claimed 24-hour battery life; it runs on OPPO's in-house context-awareness and memory-evolution algorithms, with an open ecosystem.

All capabilities are vendor claims with no independent testing or sales data. It signals a major phone maker pushing AI hardware down to the 499-yuan price point; whether real demand materializes remains to be seen.

ソース:zhidx.com

リサーチ

Offloading robot inference to cloud lifts success and battery life

クイックテイク
2026-09-24 00:01 GMT+8

Offloading robot AI inference from onboard GPUs to edge or cloud improves task success rates, response times, and battery life, according to Microsoft Research's own tests published September 23.

Previously, robotics teams ran inference on onboard GPUs; in Microsoft's own tests, mapping and planning ran up to 383% slower than an A100, timely obstacle detection in navigation dropped 30%, and VLA model accuracy fell 50%; with large onboard GPUs like Jetson Thor, a Stretch-3 robot's battery drained up to 160% more.

The team also added Kubernetes-based containerized orchestration to its Physical AI Toolchain for scheduling inference across robots, edge, and cloud. All figures are Microsoft's first-party measurements, with hardware details in its technical report; there is no independent reproduction, and the post does not quantify the cost of network latency or connectivity loss.

ソース:microsoft.com

リサーチ

Amazon opens seller accounts to outside AI assistants

クイックテイク
2026-09-24 09:22 GMT+8

At its Accelerate seller conference on September 23, Amazon launched a plugin that lets sellers check inventory, adjust prices and update listings through Claude or Amazon Quick, without logging into Seller Central.

The plugin is in beta for US-based sellers only, takes about 60 seconds to connect, and will expand internationally in the coming weeks. Sellers can limit which data the plugin accesses and must approve each action before it runs.

The contrast is notable: this week Amazon blocked Meta's Muse shopping agent for not following its rules, while opening the door to Claude from Anthropic, in which Amazon is a major investor. Amazon says more assistant integrations are coming.

ソース:siliconangle.com

リサーチ
次の読書ページ →