ReadingResearchRadarInvestment framework
Sign in / Sign up中文
Sign in / Sign up中文
ReadingResearchRadarInvestment framework
Reading archive →

Reading

2026-09-1717 posts

Prism-SQA offers interpretable EMG quality assessment, with authors reporting parity or better versus black-box methods

Quick take

You can now inspect the individual impact of five contamination components in surface EMG signals and customise quality criteria without retraining: Prism-SQA uses a U-Net plus bidirectional LSTM to decompose the signal into clean and five contamination components, then checks physiological plausibility via fingerprint verification.

Previous black-box quality assessment methods returned only a single score, showing nothing about which contamination damaged the signal, and changing quality criteria required retraining.

The authors self-report that on Ninapro synthetic-noise data and clinical dysphagia data, performance matches or exceeds black-box methods; the arXiv page notes acceptance to JBHI.

Boundary: results are author self-reported, with no third-party replication yet; the preprint by Kuan-Chen Wang and four others was submitted 11 September and revised 15 September (arXiv:2609.12724).

Sources:arxiv.org

Research

Geospatial metadata lifts cross-disciplinary dataset connection rate to 63.2%

Quick take

After small language models were fine-tuned for geospatial metadata enrichment, the self-reported cross-disciplinary metadata connection rate rose from 58.5% to 63.2%; in a metadata knowledge graph, datasets are about twice as likely to be linked across scientific disciplines through shared geospatial metadata as through keyword paths.

Previously, in an analysis of Harvard Dataverse, the authors report that only 0.3% of research datasets contained geospatial bounding boxes, so cross-disciplinary connections relied mainly on keyword paths.

Ebanks and Jain measured this change: geospatial metadata enrichment raised the connection rate from 58.5% to 63.2%, with linking power about twice that of keyword paths.

The results have not been peer-reviewed or independently verified; the arXiv preprint was submitted on September 15 and revised on September 16.

Source: arXiv abstract page: Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines ↗

Sources:arxiv.org

Research

Expected free energy acquisition function self-reports competitive regret and MSE

Quick take

Bayesian optimization now has a curvature-aware expected free energy acquisition function whose authors self-report competitive performance on both regret and mean squared error on a two-dimensional oscillator benchmark, where typical acquisition functions usually excel at only one.

Previously, Meera and Kouw proposed this acquisition function for the joint problem of optimizing and learning a function in parallel, claiming that under specific assumptions the objective reduces to UCB, LCB, or expected information gain, and proving an unbiased convergence guarantee for concave functions, which yields a curvature-aware update rule.

The empirical evidence is author-reported: a proof of concept using Van der Pol oscillator system identification, with self-reported competitive regret and mean squared error on the two-dimensional oscillator benchmark.

Boundary: the preprint (arXiv 2603.26339, first submitted March 27, revised to v2 on September 16) reports author-run experiments only, with no third-party replication yet.

Source: arXiv preprint 2603.26339, "Curvature-aware Expected Free Energy as an Acquisition Function for Bayesian Optimization", abstract page ↗

Sources:arxiv.org

Research

Cross-domain association boosts human creative originality, but seven LLMs overall do not benefit

Quick take

Random cross-domain sources reliably boost the creative originality of products designed by humans, but seven large language models overall did not benefit from the same cross-domain association intervention; only when the source and target were semantically distant enough did the highest-rated LLM benefit.

It was previously unclear whether both humans and machines could gain originality from cross-domain association. A five-person team including Liu, Dubova and Griffiths had human subjects and seven LLMs design products using random cross-domain sources, finding that the LLMs' ideas were overall more original than humans' yet did not benefit from the intervention. Humans tend to transfer surface features, while LLMs tend to transfer structural and functional properties. These are the authors' self-reported experimental findings.

The results are self-reported by the authors, appear in arXiv preprint 2603.19087 (first submitted March 19, revised to v3 on September 15), and have not yet been independently reproduced.

Source: arXiv:2603.19087 abstract page ↗

Sources:arxiv.org

Research

Apple's Glyph lifts column-description retrieval NDCG@10 to 0.92

Quick take

Column description generation and sensitivity tagging for enterprise data catalogs can now be done automatically by collaborative LLM agents: Apple self-reports that after fine-tuning a 6-layer MiniLM, same-tag retrieval NDCG@10 on an internal held-out set rose from 0.55 to 0.92.

Previously this work relied on manual effort, costly and hard to scale; Glyph's description agent retrieves the pipeline source code that generates the column on demand, and the tagging agent runs three strategies in parallel over a 275-leaf classification ontology, fused with RRF.

The measurement comes from an Apple Machine Learning Research blog post published September 16, with the paper submitted to arXiv on September 9.

What it does not cover and who has not yet reproduced it are not stated in the original.

Sources:machinelearning.apple.com

Research

Visual cues boost video-planned robot navigation, CueNav reports 70% narrow-passage success

Quick take

Giving a video model visual cues — a bird's-eye view and partial body view — can substantially raise robot navigation success: the authors self-report that maze navigation success nearly doubles versus a cue-free planner, narrow-passage success reaches 70%, and they demonstrate zero-shot semantic-conditioned navigation and deployment of the same planner across robot platforms. Previously, video models planned navigation without cues before an inverse dynamics model converted predicted video into actions, with notably lower success.

The method, called CueNav, comes from a TUM/MIRMI and MIT CSAIL team; all figures above are author self-reported, the abstract does not state whether tests were simulation or real hardware, results are independently unverified, and the project page shows code not yet released; the preprint was submitted to arXiv on September 15 and revised September 16 (arXiv:2609.16737).

Sources:arxiv.org

Research

DM³-Nav self-reports: decentralized multi-robot semantic navigation matches or beats centralized baselines

Quick take

Multiple robots can perform semantic navigation without a central coordinator or shared global state, relying only on pairwise ad hoc communication to exchange local maps and navigation intents, and the team deployed it on two robots with onboard perception in an office — a result self-reported by a Northeastern University team in DM³-Nav.

Multi-robot semantic navigation previously typically relied on a central coordinator and shared global state; the costs of centralized approaches in communication and scalability are what this work tries to avoid.

The authors self-report that on the HM3Dv0.2 and GOAT-Bench benchmarks, the method matches or exceeds centralized and shared-map baselines; the measurement was done first-party by the Northeastern University team.

The results have no independent verification; the preprint was first submitted on April 23, revised to v3 on September 16, and arXiv lists it as accepted to IROS 2026 (arXiv:2604.22014).

Sources:arxiv.org

Research

Causal Mamba hits 3.32 PESQ in real time

Quick take

Real-time speech enhancement can reach 3.32 PESQ under a 25 ms algorithmic latency constraint — that is the self-reported result of RT-SEMamba by Rong Chao et al. on Voicebank-DEMAND, built on causal time-frequency Mamba blocks.

Previously real-time denoising forced a trade-off between latency and quality, and an 8-layer teacher model was good but compute-heavy. The authors use progressive knowledge distillation to compress the 8-layer teacher into a 1-layer student, raising PESQ from a naive 1-layer baseline of 3.06 to 3.18, with steady-state RTF unchanged and a 2.64x speedup over the teacher. The authors say the fixed-size recurrent state makes long-duration inference cheaper in memory and bandwidth.

All of the above are the authors' self-reported benchmark results, not yet reproduced by a third party. The preprint was first submitted on August 12, revised to v2 on September 16, and the arXiv page notes acceptance at INTERSPEECH 2026 (arXiv:2608.12099).

Sources:arxiv.org

Research

ACT encoder ablation's 35%→2% drop fails to reproduce

Quick take

A re-run of the CVAE encoder ablation for the robot imitation learning method ACT did not reproduce the reported 35%→2% success-rate drop: as relayed by Bo Kang, the original paper claimed that removing the encoder dropped average success rate on two simulation tasks from 35% to 2%, but re-running on the original code did not reproduce the drop.

Previously readers had only the original paper's ablation conclusion, with no way to know that training duration and checkpoint selection alone can flip which policy wins, for unknown reasons.

The re-run was performed by Bo Kang on the original code (arXiv preprint, submitted September 15, revised September 16, id 2609.16745); the author also reports that the latent variable is zeroed at inference, that skipping the encoder improves training throughput, and has released code and evaluation tools.

The result is not peer-reviewed and has not been reproduced by others; it covers two simulation tasks only, not real-robot experiments.

arXiv abstract page (primary source) ↗、HTML full text v2 (primary source) ↗

Sources:arxiv.org

Research

ADORE unifies global and local explanations with first/second derivatives; authors claim it beats LIME and SHAP

Quick take

Readers can now obtain global feature importance and local sample contributions within a single framework: ADORE uses first- and second-order derivatives to characterize nonlinear feature interactions, combined with randomized SVD and dynamic sparsity detection, covering tabular, text, and image data, with a Python package open-sourced.

Previously, getting both global and local explanations required separate tools such as LIME and SHAP; the authors self-report that ADORE outperforms both in interaction modeling and computational efficiency, though this comparison is the authors' own experimental conclusion.

The measurement was made by the authors themselves (first-party self-report): the benchmarks for their interaction-modeling and computational-efficiency comparison were LIME and SHAP, with ADORE reported as superior.

Boundary: the comparison has not been reproduced by third parties, and coverage is limited to the tabular, text, and image data the authors tested; Lemen Chao and two co-authors submitted it to arXiv on September 15 (v2 revised September 16, id 2609.17171).

Sources:arxiv.org

Research

Lightweight road segmentation hits 97.23% MaxF, 68.73 FPS on Jetson

Quick take

A lightweight vision-LiDAR fusion network, LiteViLNet, self-reports 97.23±0.15% MaxF on the KITTI Road benchmark with only 14.04M parameters in the full model, and runs at 68.73 FPS under TensorRT FP16 inference on a Jetson Orin NX.

Previous lightweight road segmentation approaches often traded off between accuracy and embedded real-time performance, struggling to achieve both.

The network consists of MobileNetV3 plus a 0.12M-parameter geometric encoder; the authors self-report 97.23±0.15% MaxF on the KITTI Road benchmark, 22.18 FPS for the model alone under PyTorch FP16 inference on a Jetson Orin NX, and separately measured 68.73 FPS under TensorRT FP16. These are the authors' self-reported public benchmark results, with no independent verification yet.

The preprint LiteViLNet by Daojie Peng et al., v1 submitted on May 20, v3 revised on September 16, arXiv:2605.21007.

Sources:arxiv.org

Research

EdiTikZ self-reported: 9B beats GPT-5.6-Sol in human eval

Quick take

EdiTikZ, 4B/9B models built on Qwen3.5, can edit scientific TikZ figures; the authors self-report that in a human evaluation with 9 raters and 4,320 ratings, the 9B model scored above GPT-5.6-Sol and on par with Gemini-3.1-Pro.

No dedicated model for scientific TikZ figure editing existed before; this work mined 391,000 pairs of TikZ revisions from arXiv, GitHub and TeX SE and inferred 781,000 editing instructions for training.

The measurement is a self-reported human evaluation: 9 raters, 4,320 ratings in total, with the 9B model above GPT-5.6-Sol and comparable to Gemini-3.1-Pro; the results are not peer-reviewed.

Boundary: covers only scientific TikZ figure editing, with no third-party replication; the preprint was first posted September 1 and updated September 14 (arXiv:2609.01409), models and data are released, and the full code is said to be coming soon.

Sources:arxiv.org

Research
Next reading page →