ReadingResearchRadarInvestment framework
Sign in / Sign up
Sign in / Sign up
ReadingResearchRadarInvestment framework
Reading archive →

Reading

2026-09-1854 posts

Codec metadata speeds VLM inference up to 3.3x

Quick take
Verified 2026-09-18 02:52 GMT+8

Video codec metadata can serve directly as a runtime signal for streaming VLM inference: pruning image patches before visual encoding and selectively refreshing the KV cache across windows by frame type, with no model-specific training or offline profiling, lets concurrent streams reach up to 3.3x the baseline and cuts executed FLOPs by up to 93%.

Previously, speeding up streaming VLM inference usually required model-specific training or offline profiling, which was costly and hard to transfer to new models and workloads.

Yulin Zou and eight other authors self-report: across three VLMs and four video workloads, their vLLM-based implementation reaches up to 3.3x the baseline in concurrent streams, up to 5.3x faster average time-to-first-token, up to 93% fewer executed FLOPs, and at most a 4.64 percentage point drop in task quality. These are the authors' self-reported benchmark results.

The results do not cover other models or workloads, and no third party has reproduced them; the preprint was first submitted on April 7 and updated to v4 on September 15, arXiv:2604.06036.

Sources:arxiv.org

Research

PICKT self-reported: difficulty features most informative for very hard items; text and knowledge-graph features estimate unseen items

Quick take
Verified 2026-09-18 02:47 GMT+8

New items in intelligent tutoring systems lack response histories, degrading diagnostic reliability; after PICKT integrates multiple feature types, experiments show difficulty features are most informative for very hard items with extremely low accuracy, and fusing text and knowledge-graph features can estimate representations of unseen items via semantically or structurally similar items seen in training; the authors suggest prioritizing feature annotation according to educational-service needs. Results are self-reported by Wonbeen Lee and three co-authors.

Boundary: results are author self-reported with no third-party replication; the work first appeared in December 2025 and was updated as v2 on September 15, 2026 (arXiv:2512.07179).

Source: arXiv:2512.07179 abstract page ↗

Sources:arxiv.org

Research

Skeleton-only emotion recognition hits 37.23% on hidden test, near human 39%

Quick take
Verified 2026-09-18 02:26 GMT+8

Skeleton motion alone can now recognize 12 categories of performed emotion unseen by the performers: an 11-model ensemble by Naoto Nishida and Yoshio Ishiguro of the University of Tokyo scored 37.23% Macro-F1 on the MMAC Challenge 2026 hidden test set, close to the human 39% cited in the text, and the authors say it won the Best Performance Award.

The prior approach was 10-fold leave-performer-out cross-validation under the same protocol, yielding 36.80%, with a reproduced baseline of only 25.73%.

The 37.23% is author-reported, measured on the MMAC Challenge 2026 hidden test set; masking and counterfactual audits show the models rely on body-region motion evidence. The emotions are performed, and the results have not been independently reproduced; the data is the September 15 v2 preprint (arXiv:2609.02510), code at nawta/diema-challenge, author page nawta.github.io/mmac2026.

Sources:arxiv.org

Research

HumanEgo self-reported: 30 minutes of human video per task trains manipulation policies at 92.5% success

Quick take
Verified 2026-09-18 01:46 GMT+8

With just 30 minutes of egocentric human video per task, you can train a manipulation policy that needs no robot data at all, achieving 92.5% average success on four real tasks — self-reported by Zhi Wang and six co-authors in HumanEgo. Previously, such policies relied on robot-collected data, which is costly and hard to transfer zero-shot to new robots, cameras, and environments.

The authors' self-reported measurements show: 30 minutes of human video per task, 92.5% average success across four real tasks, 41% higher than teleoperation of equal duration; the method converts egocentric human video into entity-level hand-object representations and trains a flow-matching policy; code and dataset are public.

Boundary: results await independent replication; the preprint was first posted May 24 and updated to v3 on September 14 (arXiv:2605.24934).

Sources: arXiv preprint ↗ | Code repository ↗ | Project page ↗

Sources:arxiv.org

Research

RL-trained LoRA rewriting policy aligns SFT data distribution, non-downstream degradation drops across three backbones

Quick take
Verified 2026-09-18 01:36 GMT+8

A lightweight LoRA rewriting policy trained with reinforcement learning can rewrite SFT data, optimizing QA-style distribution alignment and semantic diversity under a task-consistency hard gate; the authors self-report that across three instruction-tuned backbones, upstream-downstream gains are roughly on par with standard SFT, and non-downstream benchmark degradation is reduced in all evaluation settings.

Previously, SFT data rewriting lacked this kind of formalization, making it hard to control rewriting quality and downstream cost at the same time.

The result is self-reported by the authors: on three instruction-tuned backbones, upstream-downstream gains are roughly on par with standard SFT, and non-downstream benchmark degradation is reduced in all evaluation settings; cross-domain reuse of the rewriting policy is only preliminary evidence.

Boundary: the results have not been independently reproduced; the work is the arXiv preprint 2602.11220, first submitted February 11 and updated to v2 on September 16. Source: arXiv abstract page ↗

Sources:arxiv.org

Research

Shared selective persistent memory lifts agent task completion to 96%, while storing full history drops it to 71%

Quick take
Verified 2026-09-18 01:23 GMT+8

An agent LLM system using shared selective persistent memory reached 96% task completion across three enterprise deployment scenarios, compared with 79% with no memory and 71% when persisting the full history — results published by Apple's machine learning research team on September 16.

The prior approaches were persisting the full history or using no memory at all; the former actually hurt task completion, and session-level reasoning traces consume substantial context.

The method keeps only four reusable context types — task specifications, data schemas, tool configurations, and output constraints — discards session-level reasoning traces, and supports cross-user sharing. The team self-reports that summary-driven data representation cuts token cost by roughly 97x versus injecting raw data. All figures are vendor self-reported.

Boundary: results are limited to three enterprise deployment scenarios and no third-party replication yet; the preprint was submitted to arXiv on July 10 (arXiv:2607.09493) and updated to v2 on September 15.

Sources:machinelearning.apple.com

Research

One steady-state snapshot suffices to recover particle interaction kernels, no trajectories, self-reported

Quick take
Verified 2026-09-18 01:08 GMT+8

A single steady-state snapshot of collective behavior can now identify an interacting particle system, with no trajectory observations at all.

Identifying such systems previously relied on trajectory data; Baoli Hao, Mauro Maggioni and Ming Zhong instead regularize this ill-posed inverse problem using empirical distributions of configurations under different unobserved initial conditions.

The authors self-report stable and accurate recovery of the interaction kernels across multiple representative steady-state and quasi-steady-state models.

Boundary: the results are self-reported by the authors and not yet independently reproduced; the work is an arXiv preprint (arXiv:2609.12004), submitted September 10 and updated to v2 on September 16.

Sources:arxiv.org

Research

DiffAdapterVLA injects trajectory tokens into late VLM layers, self-reported low-latency closed-loop planning on NAVSIM

Quick take
2026-09-23 12:00 GMT+8

DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, letting trajectory states co-evolve with driving conditions depth-by-depth inside the backbone, refining trajectories recursively with lightweight layer-wise adapters and dropping the separate planner. Previously, trajectory generation in driving VLMs was decoupled from driving-condition evolution and often relied on a separate planner, adding latency and complexity.

The authors self-report high-quality, low-latency closed-loop planning with a small number of trainable parameters on the NAVSIM benchmark.

Boundary: results are self-reported with no third-party replication; the work is by a team of eight including Changxin Lu, submitted to arXiv as a preprint on September 14 (v2 updated on the 16th, arXiv:2609.15322).

Sources:arxiv.org

Research

Surgical video QA benchmark self-reports 14,256 pairs and 14.61% gain

Quick take
Verified 2026-09-18 00:59 GMT+8

Surgical video AI now has a question-answering benchmark covering five surgical task types, SurgCoTBench, which the authors self-report contains 14,256 QA pairs, with retrieval-augmented multi-agent reasoning accuracy exceeding supervised models by 14.61%.

Previously the field lacked a unified surgical video question-answering benchmark, making it hard to compare methods across five surgical tasks.

Chang Han Low et al. self-reported the above benchmark and the 14.61% accuracy gain, updating arXiv v3 on September 16, with the entry marked as citing IEEE RA-L 2026; code is open on GitHub and the dataset has been released.

No independent verification yet.

Sources:

  • arXiv abstract page (v3, revised 2026-09-16) ↗
  • SurgRAW code repository (GitHub) ↗

Sources:arxiv.org

Research

2026-09-1717 posts

Prism-SQA offers interpretable EMG quality assessment, with authors reporting parity or better versus black-box methods

Quick take
Verified 2026-09-17 23:58 GMT+8

You can now inspect the individual impact of five contamination components in surface EMG signals and customise quality criteria without retraining: Prism-SQA uses a U-Net plus bidirectional LSTM to decompose the signal into clean and five contamination components, then checks physiological plausibility via fingerprint verification.

Previous black-box quality assessment methods returned only a single score, showing nothing about which contamination damaged the signal, and changing quality criteria required retraining.

The authors self-report that on Ninapro synthetic-noise data and clinical dysphagia data, performance matches or exceeds black-box methods; the arXiv page notes acceptance to JBHI.

Boundary: results are author self-reported, with no third-party replication yet; the preprint by Kuan-Chen Wang and four others was submitted 11 September and revised 15 September (arXiv:2609.12724).

Sources:arxiv.org

Research

Geospatial metadata lifts cross-disciplinary dataset connection rate to 63.2%

Quick take
Verified 2026-09-17 23:36 GMT+8

After small language models were fine-tuned for geospatial metadata enrichment, the self-reported cross-disciplinary metadata connection rate rose from 58.5% to 63.2%; in a metadata knowledge graph, datasets are about twice as likely to be linked across scientific disciplines through shared geospatial metadata as through keyword paths.

Previously, in an analysis of Harvard Dataverse, the authors report that only 0.3% of research datasets contained geospatial bounding boxes, so cross-disciplinary connections relied mainly on keyword paths.

Ebanks and Jain measured this change: geospatial metadata enrichment raised the connection rate from 58.5% to 63.2%, with linking power about twice that of keyword paths.

The results have not been peer-reviewed or independently verified; the arXiv preprint was submitted on September 15 and revised on September 16.

Source: arXiv abstract page: Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines ↗

Sources:arxiv.org

Research

Expected free energy acquisition function self-reports competitive regret and MSE

Quick take
Verified 2026-09-17 23:28 GMT+8

Bayesian optimization now has a curvature-aware expected free energy acquisition function whose authors self-report competitive performance on both regret and mean squared error on a two-dimensional oscillator benchmark, where typical acquisition functions usually excel at only one.

Previously, Meera and Kouw proposed this acquisition function for the joint problem of optimizing and learning a function in parallel, claiming that under specific assumptions the objective reduces to UCB, LCB, or expected information gain, and proving an unbiased convergence guarantee for concave functions, which yields a curvature-aware update rule.

The empirical evidence is author-reported: a proof of concept using Van der Pol oscillator system identification, with self-reported competitive regret and mean squared error on the two-dimensional oscillator benchmark.

Boundary: the preprint (arXiv 2603.26339, first submitted March 27, revised to v2 on September 16) reports author-run experiments only, with no third-party replication yet.

Source: arXiv preprint 2603.26339, "Curvature-aware Expected Free Energy as an Acquisition Function for Bayesian Optimization", abstract page ↗

Sources:arxiv.org

Research
Next reading page →