LecturaInvestigaciónRadarMarco de inversión
Iniciar sesión / Registrarse
Iniciar sesión / Registrarse
LecturaInvestigaciónRadarMarco de inversión
Archivo de lecturas →

Lectura

2026-10-084 publicaciones

Microsoft Open-Sources Lightweight Agentic RL Framework

Material
2026-10-08 00:00 GMT+8

Microsoft Research Asia open-sourced Agent Lightning v1.0 on October 7, a lightweight framework of approximately 3,500 lines of code designed to allow the same agent harness used in deployment to participate directly in reinforcement learning.

Traditional agentic RL often requires reimplementing agent logic within the training framework, leading to discrepancies between training and production environments and high costs. Agent Lightning inserts an LLM proxy between the agent and the model, allowing existing harness code to remain unchanged while connecting to RL training.

The framework natively supports running agents as Kubernetes jobs, eliminating dependency on paid commercial sandbox services. In coding agent experiments, using about 6,000 training samples, it improved Qwen3.5-9B's Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points. These results are self-reported by Microsoft and await independent reproduction.

Fuentes:microsoft.com

Investigación

Healthleap Raises $38M for AI Screening

Tema · Healthleap融资Resumen rápido
2026-10-07 23:07 GMT+8

Healthleap announced it has raised $38 million across two rounds: an $8 million seed co-led by Sequoia Capital and First Round Capital, and a $30 million Series A led by Hummingbird Ventures.

Founded in 2022, the startup uses large language models to analyze unstructured text in electronic health records, such as clinician notes, to identify patients at risk for undiagnosed conditions like malnutrition and delirium. The platform is currently deployed in over 50 hospitals, including Penn Medicine and Cedars-Sinai.

The company did not disclose its valuation. CEO Josiah Meyer stated that revenue has grown more than 10x over the past three years and that all customers have achieved at least 5x hard ROI, though these figures are self-reported and not independently audited. The new capital will support engineering, sales, and expansion into outpatient and home care settings.

Fuentes:techcrunch.com

Investigación

Microsoft Sandboxes AI Agents on Windows

Material
2026-10-08 01:24 GMT+8

Microsoft introduced an SDK that sandboxes AI agent code execution on Windows.

The MXC (Microsoft Execution Containers) SDK was unveiled at the Windows and Surface event, positioned as a security isolation layer that lets agents execute code and call tools locally while minimizing access to sensitive user data. Per The Verge's report, Microsoft described it as a key component of its Windows AI strategy.

Microsoft has not disclosed full technical specifications, developer availability, or a launch timeline.

Fuentes:ithome.com

Investigación

Fewer parameters, faster training, but don't treat it as a universal solution when switching tasks

Material
Verificado 2026-10-08 00:03 GMT+8

Lectura a largo plazo · 《ALBERT: A Lite BERT for Self-supervised Learning of Language Representations》(2019)

ALBERT is an approach to saving parameters in text models: it reduces the size of word representations and shares parameters across layers, enabling faster learning of sentence meanings with fewer parameters. When you see it achieving high scores on text tasks with fewer parameters, first ask where those saved parameters come from.

The mechanism it provides is this: parameters are saved by reducing the size of word representations and sharing them across layers, while using sentence-order prediction to retain multi-sentence task capabilities. Therefore, when reading about low-parameter text models, ask whether these two aspects still fit your specific task.

If dealing with non-English languages, extremely small datasets, or if the width per layer is too large, do not use it to judge general optimality. In some single-sentence tasks, removing the sentence-order objective has limited impact.

"ALBERT" (2019) | Next review date: 2027-09-20

Fuentes:arxiv.org

Investigación

Estás al día en esta vista