ReadingResearchRadarInvestment framework
Sign in / Sign up
Sign in / Sign up
ReadingResearchRadarInvestment framework
Reading archive →

Reading

2026-10-0740 posts

US Local Media Sue Microsoft and OpenAI for Copyright Infringement

Material
2026-10-07 18:19 GMT+8

Emmerich Newspapers and other US local media groups have filed a lawsuit against Microsoft and OpenAI, accusing the tech giants of unauthorized scraping of tens of thousands of copyrighted articles to train AI models.

The plaintiffs allege that the defendants bypassed paywall restrictions and generated billions in profits from the stolen content without compensating publishers. The suit cites violations of the Copyright Act and the Digital Millennium Copyright Act, seeking damages and an injunction to remove infringing training data from the models.

Reported by neowin and translated by IT Home. This case marks a shift as local newspaper chains join the copyright battle, indicating that AI litigation is expanding beyond major national media outlets.

Sources:ithome.com

Research

AI Self-Planning Beats Fixed Algorithms

Topic · 智能体搜索效率优化Material
2026-10-06 12:00 GMT+8

The AgentDiscover framework demonstrates that letting large language models (LLMs) plan their own search paths is more efficient than using fixed algorithms designed by humans.

Traditional AI research assistants typically act only as "proposers," while the search steps are controlled by human-designed algorithms like Monte Carlo tree search. As model capabilities grow, these manual constraints become bottlenecks. AgentDiscover promotes the model to a "planner," using its context as working memory to run experiments and record every attempt in a database. This database serves as long-term memory, allowing the model to freely use, combine, or replace the selection rules of classical algorithms.

Author-reported tests show the framework achieves higher scores at lower cost. In simulations of seven past AtCoder heuristic programming contests, its generated programs would have placed first among human competitors. On eleven mathematical and systems optimization tasks, it matched or exceeded baselines using the same underlying model.

This is an arXiv preprint (submitted October 4, 2026). Results are self-evaluated by the author team and have not yet been independently reproduced.

Sources:arxiv.org

Research

ForkPilot Cuts Tokens by 59%

Topic · 智能体搜索效率优化Material
2026-10-06 12:00 GMT+8

The ForkPilot framework achieves up to 59.2% token reduction in long-horizon agent tasks while maintaining state-of-the-art performance.

Traditional agents often struggle with retrospective search due to delayed outcomes causing attribution complexity and stale estimates from evolving execution evidence. ForkPilot introduces Search Value Dynamics (SVD) and uses a two-stage self-evolving policy: learning offline from completed trajectories, then making real-time decisions and updating itself.

Evaluated across 6 benchmarks and 7 LLM backbones (including GPT-5.6 Sol and Opus 4.8) against 9 baselines, the study reports significant efficiency gains. Note that these results come from a preprint and await independent reproduction.

Sources:arxiv.org

Research

MemTrace Boosts Long-Horizon Coding Agents

Topic · MemTrace记忆系统Material
2026-10-06 12:00 GMT+8

The MemTrace system improved the pass@1 metric on the DeepSWE benchmark by 21.2 percentage points.

Long-horizon coding agents often face context overflow from extended execution trajectories or invalidation of earlier evidence due to repository changes. MemTrace uses immutable memory traces and a dependency graph to validate evidence against the current state before retrieval, fetching only what is needed for the next action.

Under the Codex CLI environment, the system also increased the SWE-EVO Resolved Rate by 4.4 points and the SWE-Milestone Score by 17.8 points. These are author-reported results, not yet independently reproduced.

Sources:arxiv.org

Research

Small Model Cuts Video Agent Costs

Quick take
2026-10-06 12:00 GMT+8

The VideoResearchAgent framework enables Qwen3.5-4B to achieve 40.48% accuracy on the Video-BrowseComp benchmark, comparable to Gemini-3-Flash-Preview.

This result comes from a preprint submitted on October 4. Traditional video agents rely on slow, unreliable live web interactions, while fixed local simulations often induce retrieval-specific shortcuts that fail to generalize to the open web.

The team introduced the RDR-GRPO algorithm, which randomizes candidate rankings, distractors, and metadata during training to reduce overfitting. Compared to the untrained model, cumulative API token consumption dropped by 74.9%.

These are author-reported results based on a specific simulator environment; independent reproduction on real open-web video has not yet been verified.

Sources:arxiv.org

Research

TaSQ Boosts 1-Bit Cache Throughput

Quick take
2026-10-05 12:00 GMT+8

The TaSQ algorithm achieves 14x larger batch sizes and 1.87x higher peak throughput on an RTX 6000 Ada GPU.

The study addresses memory bottlenecks in long-context LLM inference by tailoring vector quantization target spaces. It maintains reasoning stability under extreme 1-bit compression using query-guided channel weighting.

Authors report significant performance gains over BF16 baselines via SGLang implementation. Results are self-reported preprint data without independent reproduction.

Sources:arxiv.org

Research

EV Firm Plans 100k Nvidia GPU Edge Network

Quick take
2026-10-06 20:45 GMT+8

EV charging company Xeal announced plans to deploy 100,000 Nvidia GPUs across its US network to create the "world's first edge inference compute network using idle EV charging capacity."

The firm claims its 1,600+ sites have 200MW of permitted power but typically operate at less than 10% capacity. Its Latient Pods contain up to 48 Hopper or Blackwell Ultra GPUs each, valued at roughly $2 million, requiring no water cooling and operating quietly.

CEO Nikhil Bharadwaj states this bypasses grid interconnect delays for new data centers. However, Tom's Hardware highlights significant theft risks for high-value equipment at public sites, citing prior cable thefts. The first pod is scheduled to go online by the end of the year.

Sources:tomshardware.com

Research

61% UK Workers Lack AI Threat Training

Quick take
2026-10-07 15:00 GMT+8

A new survey by UK training provider QA found that 61% of workers received no training on spotting AI-driven cyber threats in the last 12 months.

As AI enables convincing phishing emails and deepfake audio, traditional defenses like proofreading have become ineffective. The study also noted that 22% of organizations lack policies against such attacks, and younger employees (18-24) report lower confidence in identifying AI scams than older peers.

The findings, based on a sample of 1,000 people, underscore the need for foundational AI literacy to close security skill gaps.

Sources:techradar.com

Research

Ex-Ramp Engineers Raise $25M for Melius

Quick take
2026-10-07 06:34 GMT+8

AI creative generation platform Melius announced a total of $25 million in funding, comprising a $20 million Series A led by CRV and a $5 million seed round led by General Catalyst.

The founding team consists of former engineers from Ramp. They initially built a tool to help marketers optimize ad spend but abandoned the project after six months, deeming it unpromising. The team scrapped the entire codebase to pivot toward building an "agents lab" that generates creative assets and campaigns directly.

The company claims to have hit more than $1 million in annualized revenue within two months of coming out of stealth in July. This figure is self-reported and has not been independently audited. Melius will compete with rivals like Higgsfield and Krea in the AI marketing content space.

Sources:techcrunch.com

Research

Surface Laptop Ultra specs leak

Quick take
2026-10-07 11:30 GMT+8

Tech outlet Windows Central leaked detailed specifications for the Microsoft Surface Laptop Ultra on October 6, revealing a top configuration with 128GB of unified memory.

The device features the NVIDIA RTX Spark Superchip, integrating a Grace CPU and Blackwell RTX GPU. The 20-core version delivers 1 PetaFLOP of AI compute, claimed to support running frontier models with over 120 billion parameters locally.

It includes a 15-inch PixelSense Ultra display with 2000 nits HDR peak brightness. A new system feature allows users to adjust GPU memory allocation within the unified memory pool to optimize AI workflows.

Pricing remains unknown, though Microsoft showcased the product prototype in June this year.

Sources:ithome.com

Research

RAM-Net Optimizes Linear Attention with Sparse Addressing

Topic · RAM-Net架构Material
2026-10-07 12:00 GMT+8

RAM-Net proposes a sparse address-based access mechanism to address the degradation of long-range fine-grained recall caused by shared states in linear attention.

Traditional linear attention superimposes information from distinct tokens within a fixed-size recurrent state, creating inter-token interference. RAM-Net organizes the state as an array of independent slots and uses an Address Decoder to map each key or query to a sparse address, selecting only a small subset of slots to write to or read from at each step.

Authors report that this design achieves the lowest perplexity with competitive commonsense reasoning and outperforms strong baselines on fine-grained long-range retrieval. A key efficiency metric is accessing fewer state elements per step than all baselines, e.g., 8x fewer than Mamba2.

These are first-party benchmark results from a paper accepted at NeurIPS 2026; independent reproduction has not yet been reported.

Sources:arxiv.org

Research

Nash Decoding Beats Scale

Topic · 纳什解码Material
2026-10-06 12:00 GMT+8

Nash decoding allows masked language models to outperform autoregressive models up to 18 times larger on question-answering benchmarks.

The preprint models text revision as a multi-player game: each token position is a player, vocabulary items are actions, and the goal is a Nash equilibrium that maximizes joint probability. On CLAPNQ, PubMedQA and CoQA, the authors self-report that masked models (which predict all positions in parallel) using Nash decoding achieve higher F1 and ROUGE scores than autoregressive models (which generate token by token) with up to 18x more parameters, without fine-tuning. The cost is significantly increased test-time computation, not quantified in the abstract.

Results are from the authors' own evaluation; no independent reproduction has been reported.

Sources:arxiv.org

Research
Next reading page →