LesenForschungRadarAnlageframework
Anmelden / Registrieren
Anmelden / Registrieren
LesenForschungRadarAnlageframework
Lesearchiv →

Lesen

2026-10-0733 Beiträge

AI Self-Planning Beats Fixed Algorithms

Wesentlich
2026-10-06 12:00 GMT+8

The AgentDiscover framework demonstrates that letting large language models (LLMs) plan their own search paths is more efficient than using fixed algorithms designed by humans.

Traditional AI research assistants typically act only as "proposers," while the search steps are controlled by human-designed algorithms like Monte Carlo tree search. As model capabilities grow, these manual constraints become bottlenecks. AgentDiscover promotes the model to a "planner," using its context as working memory to run experiments and record every attempt in a database. This database serves as long-term memory, allowing the model to freely use, combine, or replace the selection rules of classical algorithms.

Author-reported tests show the framework achieves higher scores at lower cost. In simulations of seven past AtCoder heuristic programming contests, its generated programs would have placed first among human competitors. On eleven mathematical and systems optimization tasks, it matched or exceeded baselines using the same underlying model.

This is an arXiv preprint (submitted October 4, 2026). Results are self-evaluated by the author team and have not yet been independently reproduced.

Quellen:arxiv.org

Forschung

Small Model Cuts Video Agent Costs

Kurzfassung
2026-10-06 12:00 GMT+8

The VideoResearchAgent framework enables Qwen3.5-4B to achieve 40.48% accuracy on the Video-BrowseComp benchmark, comparable to Gemini-3-Flash-Preview.

This result comes from a preprint submitted on October 4. Traditional video agents rely on slow, unreliable live web interactions, while fixed local simulations often induce retrieval-specific shortcuts that fail to generalize to the open web.

The team introduced the RDR-GRPO algorithm, which randomizes candidate rankings, distractors, and metadata during training to reduce overfitting. Compared to the untrained model, cumulative API token consumption dropped by 74.9%.

These are author-reported results based on a specific simulator environment; independent reproduction on real open-web video has not yet been verified.

Quellen:arxiv.org

Forschung

ForkPilot Cuts Tokens by 59%

Wesentlich
2026-10-06 12:00 GMT+8

The ForkPilot framework achieves up to 59.2% token reduction in long-horizon agent tasks while maintaining state-of-the-art performance.

Traditional agents often struggle with retrospective search due to delayed outcomes causing attribution complexity and stale estimates from evolving execution evidence. ForkPilot introduces Search Value Dynamics (SVD) and uses a two-stage self-evolving policy: learning offline from completed trajectories, then making real-time decisions and updating itself.

Evaluated across 6 benchmarks and 7 LLM backbones (including GPT-5.6 Sol and Opus 4.8) against 9 baselines, the study reports significant efficiency gains. Note that these results come from a preprint and await independent reproduction.

Quellen:arxiv.org

Forschung

MemTrace Boosts Long-Horizon Coding Agents

Wesentlich
2026-10-06 12:00 GMT+8

The MemTrace system improved the pass@1 metric on the DeepSWE benchmark by 21.2 percentage points.

Long-horizon coding agents often face context overflow from extended execution trajectories or invalidation of earlier evidence due to repository changes. MemTrace uses immutable memory traces and a dependency graph to validate evidence against the current state before retrieval, fetching only what is needed for the next action.

Under the Codex CLI environment, the system also increased the SWE-EVO Resolved Rate by 4.4 points and the SWE-Milestone Score by 17.8 points. These are author-reported results, not yet independently reproduced.

Quellen:arxiv.org

Forschung

61% UK Workers Lack AI Threat Training

Kurzfassung
2026-10-07 15:00 GMT+8

A new survey by UK training provider QA found that 61% of workers received no training on spotting AI-driven cyber threats in the last 12 months.

As AI enables convincing phishing emails and deepfake audio, traditional defenses like proofreading have become ineffective. The study also noted that 22% of organizations lack policies against such attacks, and younger employees (18-24) report lower confidence in identifying AI scams than older peers.

The findings, based on a sample of 1,000 people, underscore the need for foundational AI literacy to close security skill gaps.

Quellen:techradar.com

Forschung

Surface Laptop Ultra specs leak

Kurzfassung
2026-10-07 11:30 GMT+8

Tech outlet Windows Central leaked detailed specifications for the Microsoft Surface Laptop Ultra on October 6, revealing a top configuration with 128GB of unified memory.

The device features the NVIDIA RTX Spark Superchip, integrating a Grace CPU and Blackwell RTX GPU. The 20-core version delivers 1 PetaFLOP of AI compute, claimed to support running frontier models with over 120 billion parameters locally.

It includes a 15-inch PixelSense Ultra display with 2000 nits HDR peak brightness. A new system feature allows users to adjust GPU memory allocation within the unified memory pool to optimize AI workflows.

Pricing remains unknown, though Microsoft showcased the product prototype in June this year.

Quellen:ithome.com

Forschung

TaSQ Boosts 1-Bit Cache Throughput

Kurzfassung
2026-10-05 12:00 GMT+8

The TaSQ algorithm achieves 14x larger batch sizes and 1.87x higher peak throughput on an RTX 6000 Ada GPU.

The study addresses memory bottlenecks in long-context LLM inference by tailoring vector quantization target spaces. It maintains reasoning stability under extreme 1-bit compression using query-guided channel weighting.

Authors report significant performance gains over BF16 baselines via SGLang implementation. Results are self-reported preprint data without independent reproduction.

Quellen:arxiv.org

Forschung

Ex-Ramp Engineers Raise $25M for Melius

Kurzfassung
2026-10-07 06:34 GMT+8

AI creative generation platform Melius announced a total of $25 million in funding, comprising a $20 million Series A led by CRV and a $5 million seed round led by General Catalyst.

The founding team consists of former engineers from Ramp. They initially built a tool to help marketers optimize ad spend but abandoned the project after six months, deeming it unpromising. The team scrapped the entire codebase to pivot toward building an "agents lab" that generates creative assets and campaigns directly.

The company claims to have hit more than $1 million in annualized revenue within two months of coming out of stealth in July. This figure is self-reported and has not been independently audited. Melius will compete with rivals like Higgsfield and Krea in the AI marketing content space.

Quellen:techcrunch.com

Forschung

AI milestones follow as compute rises — but it breaks down when data runs out

Wesentlich
Verifiziert 2026-10-07 00:02 GMT+8

Langfristiges Lesen · 《AI and Compute》(2018)

《AI and Compute》 is a statistical report released by OpenAI in 2018, looking at how much compute it took to train an AI model. It tallied the compute used in historical model trainings and found that compute demand has long grown exponentially, doubling over time, and that several famous capability breakthroughs appeared right after a big leap in compute.

When you hear a vendor today say "we stacked ten times the compute and the model got stronger," the framework established by this report is what you use to judge whether that claim is credible: there is indeed a historical correspondence between compute investment and capability gains. Later analyses have only updated the numbers on top of its statistics; the framework itself hasn't been replaced.

If a field's data has been exhausted or has hit a physical limit, don't use it to make judgments — doubling compute again won't buy an equivalent improvement. The report itself is an observational statistic, not a law, and the trend could be interrupted by algorithmic progress at any time.

《AI and Compute》(2018) | Next review 2027-09-20

Quellen:openai.com

Forschung

2026-10-0656 Beiträge

OpenAI and Synopsys Partner on Chip Design AI

Thema · OpenAI新思芯片合作Strukturell
Verifiziert 2026-10-06 19:38 GMT+8

OpenAI and Synopsys announced a multi-year agreement on September 30 to jointly develop GPT-Synopsys, a specialized model designed to directly operate Electronic Design Automation (EDA) tools.

Historically, chip design has relied heavily on engineers manually running synthesis, place-and-route, and verification software, iterating repeatedly to balance Power, Performance, and Area (PPA). While Synopsys introduced DSO.ai in 2020 to optimize specific steps using reinforcement learning, the overall workflow remained largely human-driven.

According to the announcement, GPT-Synopsys aims to make frontier models "expert users" of EDA tools. Engineers will delegate high-level objectives, while agents run the tools, interpret results, and implement changes until a verified outcome is ready for review. The model will run on OpenAI-hosted infrastructure and integrate deeply with Synopsys' Autopilot platform.

Early technology engagements are underway with leading semiconductor customers, though no release date or pricing structure was disclosed. Whether this collaboration can truly replace senior engineering judgment in complex nodes remains to be seen.

Quellen:news.synopsys.com

Forschung

Mistral Launches 1T, 49B-Active Large 4 Preview

Wesentlich
2026-10-06 20:00 GMT+8

Mistral launches Large 4 preview; weights due, it says.

API is live on Mistral Studio; 1T total, 49B active.

Mistral says it is strong in coding, agents and vision; vendor says.

Quellen:mistral.ai

Forschung

AI Runaway May Implicate CEOs

Thema · AI智能体责任保险Wesentlich
2026-10-06 12:00 GMT+8

Insurers and lawyers are assessing more than 300 AI cases and preparing for rogue-AI claims against OpenAI and Anthropic executives.

Tim Rayner, UK head of underwriting and claims at Verisk, said OpenAI’s CEO is ultimately liable for the Hugging Face incident because of an absence of control in the business. If OpenAI holds directors’ and officers’ insurance, it could seek to cover future losses from lawsuits targeting Altman.

Aon analysed more than 300 AI-related legal cases and found insurers could also be exposed under crime, intellectual property, media liability, cyber security and technology errors and omissions policies. The cases lack precedent and have not been tested in court, and Hiscox chief executive Aki Hussain said it is too soon to know how US courts will treat AI-agent liability.

Quellen:ft.com

Forschung
Nächste Leseseite →