독서리서치레이더투자 프레임워크
로그인 / 회원가입
로그인 / 회원가입
독서리서치레이더투자 프레임워크
읽기 아카이브 →

독서

2026-09-2933 게시물

Anthropic ships a new mid-tier Claude, claiming a big coding leap

중대
검증됨 2026-09-30 02:37 GMT+8

Anthropic released Claude Sonnet 5.5 on September 28, positioning it as the mid-tier model below Opus 5.5, with claimed speed gains of 30%+ and per-task costs up to 30% lower.

The headline number on Anthropic's page: a score of 70.6% on the Terminal-Bench 4.0 agentic coding evaluation, versus 10.3% for the previous Sonnet 5. Prices are unchanged — $2 per million input tokens and $10 per million output tokens. It is also the first Sonnet-tier model to ship with cyber safeguards and fallbacks.

One caveat: all scores are Anthropic's own tests, and because OpenAI has not published GPT-6 Sol results, the comparison tables actually use GPT-5.6 Sol instead, weakening cross-vendor comparability. The launch lands just before OpenAI's DevDay.

출처:anthropic.com

리서치

Claude Code launches Projects, splitting one conversation into parallel cloud tasks

간단한 요점
2026-09-29 09:48 GMT+8

Anthropic's official developer account announced on September 17 that Projects is rolling out in Claude Code on desktop and web. It is in beta for select users.

Per the announcement, a project is one conversation with Claude: it splits the work into threads itself, runs them as parallel cloud sessions, passes context between them, and keeps going when you leave. Developer Boris Cherny says he has stopped managing sessions and simply sends thoughts as they come.

The change moves task splitting and scheduling from the user to the agent itself. The announcement does not say how widely the beta will open or when it will reach general availability.

출처:latent.space

리서치

Meta launches a Muse agent for small businesses that reads social and ad data

간단한 요점
2026-09-29 18:04 GMT+8

Meta on September 29 launched Muse for Small Business, an agent that can read a company's Facebook and Instagram business accounts, analytics and ad accounts, then produce concrete task plans from that data.

The product also connects to 15 third-party services including Canva, Figma, Slack, Shopify, Stripe and QuickBooks, covering design, payments and collaboration. Per IT之家's report on Meta's announcement, the company plans further enterprise AI products.

A week ago Muse was reported to have leaked sellers' home addresses without consent and arranged buyer visits on its own; handing business and ad data to the agent makes the permission boundary the first thing to scrutinize.

출처:ithome.com

리서치

Oracle launches Fusion Claw, letting enterprise apps rewrite business records on their own

간단한 요점
2026-09-29 20:00 GMT+8

Oracle on September 29 introduced Fusion Claw, an agentic runtime that lets its Fusion enterprise applications carry out complex business tasks autonomously, alongside 25 new Claw-powered applications that bring its Fusion Agentic Applications portfolio to 75.

Claw uses large language models only for reasoning and planning, keeping calculations and transaction processing outside the model to limit cost; the initial release supports frontier models from Gemini and OpenAI, and users cannot yet plug in their own smaller open models. Oracle executives told SiliconANGLE the company previously had no solutions for advanced optimization.

To address the risk of software changing business records, Oracle offers an Enterprise Operating Envelope with permissions, risk thresholds and approval requirements, plus an "Outcome Receipt" documenting policies and transactions after each task. The company provided no measured cost savings or production performance evidence.

출처:siliconangle.com

리서치

GPT-6 Sol and Luna land on Snowflake in public preview

간단한 요점
2026-09-29 02:13 GMT+8

Snowflake announced on September 28 that OpenAI's GPT-6 Sol and GPT-6 Luna are now in public preview on Cortex AI.

Both models are callable today through Cortex AI Functions in SQL — for example, using the AI_COMPLETE function to analyze filing text directly — and through Snowflake's OpenAI-compatible endpoint for applications. All calls run within Snowflake's own security and governance perimeter.

Snowflake describes Sol as suited to multi-step professional work and Luna to responsive execution at scale, drawing that characterization from OpenAI. Integrations with CoCo, CoWork and Cortex Agents are announced as coming soon, not live.

출처:snowflake.com

리서치

OpenAI doubles its investment in the Lenfest journalism AI program

간단한 요점
2026-09-28 15:00 GMT+8

OpenAI has doubled its support for the Lenfest AI Collaborative and Fellowship Program: a new $5 million commitment, plus up to $5 million in software credits and engineering support.

The program, run by the Lenfest Institute for Journalism (a nonprofit funder of American local news) with OpenAI since 2024, places full-time AI engineers in 11 US news organizations. The Philadelphia Inquirer used it to build archive-search and monitoring tools; Chicago Public Media sped up Spanish-language publishing.

A new cohort of news organizations will be invited to join. The announcement is joint and describes results from the program's own perspective; several fellows are expected to stay on as full-time employees at their host organizations, a trackable signal of whether the model actually takes root.

출처:openai.com

리서치

Momentic launches Mo, an AI testing agent that needs no test scripts

간단한 요점
2026-09-29 00:00 GMT+8

Software testing company Momentic launched Mo on September 28, an AI agent that tests applications directly from a developer's instructions, with no test scripts to write or maintain.

Mo works like coding tools such as Claude Code: give it a URL and a testing brief, and it spins up a swarm of agents to operate the app, trying thousands of edge cases, then reports reproduction steps and video for each bug. It can also read a product requirements document or a Jira ticket for context.

Co-founder and CEO Wei-Wei Wu told SiliconANGLE that as long as scripts sit in a codebase, someone has to maintain them. The company says customers including Notion have been trying Mo; there is no independent evaluation yet.

출처:siliconangle.com

리서치

Security startup Rig raises $12M to police AI agents running under employee accounts

간단한 요점
2026-09-29 20:00 GMT+8

Israeli security startup Rig Security launched on September 29 with $12 million in seed funding.

The blind spot it targets is specific: coding assistants and autonomous agents rarely hold accounts of their own, so their actions are logged under the employee or service account that launched them. In the words of CEO Guy Kozliner, agents are now the most active identities in an enterprise, "and they are acting under human names" — security teams cannot tell person from agent, and stopping the agent means blocking the employee too.

The product uses machine learning to match one identity across identity systems, with the company itself claiming better than 96% accuracy; an endpoint sensor can stop an agent's risky action without interrupting the employee. The company says Fortune 200 customers in financial services, insurance and health care already run it in production — both that claim and the accuracy figure are the company's own. The round was co-led by Ten Eleven Ventures and Brightmind Partners, with CrowdStrike participating as a strategic investor.

출처:siliconangle.com

리서치

Four pharma giants join Apheris consortium to train AI on 10,000 antibodies

간단한 요점
2026-09-29 20:00 GMT+8

AbbVie, argenx, Lundbeck and Takeda have joined Berlin-based Apheris and Ginkgo Datapoints in a new Antibody Developability Consortium, training AI on roughly 10,000 antibody sequences to predict which drug candidates can be manufactured and reach the clinic.

Each member contributes proprietary sequences and trains models inside its own environment via Apheris's federated infrastructure, without raw sequences being exposed to other members. Ginkgo Datapoints fills remaining capacity from public sources and runs wet-lab characterisation; the initial dataset is due to members by early 2027.

The consortium has just kicked off and has released no model performance data. Charlotte Deane of Oxford and Peter Tessier of the University of Michigan provide independent scientific oversight.

출처:tech.eu

리서치

US Navy stands up drone command to integrate air, sea and undersea tactics

간단한 요점
2026-09-29 22:00 GMT+8

The US Navy formally established the Robotic and Autonomous Systems Warfighting Development Centre in Virginia on September 24, tasked with turning its aerial, surface and undersea drones into an integrated fighting force.

The centre will develop doctrine and tactics for unmanned systems, train sailors, and test whether the platforms keep operating under realistic conditions. Chief of Naval Operations Admiral Daryl Caudle said years of robotic-systems work had proceeded separately across domains, that the next step is integration, and that "a prototype is not combat power."

US commanders have previously linked this drone capability to disrupting a PLA attack on Taiwan. The centre's budget and staffing have not been disclosed; its integration progress remains to be seen.

출처:scmp.com

리서치

Musk says Cybercab won't run commercial rides in California until mid-2027

간단한 요점
2026-09-29 17:37 GMT+8

Elon Musk expects Tesla's Cybercab to start commercial operation in California only around mid-2027.

In an interview with China Global Television Network released on Saturday, September 26, Musk said the two-seat, steering-wheel-free vehicle "will soon be operating commercially in Florida, in Nevada, and a couple of other states," and that in California it will "probably" be running "by the middle of next year."

Tesla had earlier promised to sell the vehicle to regular buyers for $30,000 before the end of 2026 — the promise behind YouTuber MKBHD's head-shaving bet. The car still faces an NHTSA audit query into Tesla's self-certification that it meets federal safety standards, a hard gate before any wider rollout.

출처:businessinsider.com

리서치

Alignment is purpose-agnostic — and works as a censor's toolkit

간단한 요점
2026-09-29 04:07 GMT+8

The same alignment techniques built to make models safe can just as easily censor or distort information: alignment methods are purpose-agnostic, making a model serve someone's will with nothing in the methodology guaranteeing good intent.

Alignment has largely been assumed to be a safety measure; Sarah Ball and Phil Hackemann sort control into three levers — pretraining data filtering, post-training alignment, and inference-time intervention — with cost falling and ease of change rising down the stack, making lower layers easier to abuse unilaterally.

This is not hypothetical: China's cyberspace regulator requires providers to maintain refusal datasets, and Elon Musk publicly said he would "fix" Grok outputs he disagreed with, with reported behavior shifts traced to system-prompt changes. The analysis won an Outstanding Position Paper Award at ICML 2026; the authors do not call for stopping alignment, but back transparency, verifiable alignment, and model pluralism.

This is a position paper's argument plus public examples — author-reported views, with no third-party replication of the quantified risk.

출처:lesswrong.com

리서치
다음 읽기 페이지 →