読書リサーチレーダー投資フレームワーク
ログイン / 新規登録
ログイン / 新規登録
読書リサーチレーダー投資フレームワーク
アーカイブ →

読書

2026-09-1854 投稿

4-bit quantization largely preserves LLM personality structure

クイックテイク
検証済み 2026-09-18 08:47 GMT+8

Model builders can now know that 4-bit quantization largely preserves the coarse-grained personality structure of large models, while 2-bit quantization breaks fine-grained prompt consistency and cross-precision consistency, and that personality decisions mainly form in the upper layers.

Previously, when choosing quantization for open-source large models, there was no layered MBTI personality analysis basis for whether personality structure changes with precision.

Yao Fu and five other authors performed a layered MBTI personality analysis on open-source large models, covering GPTQ, AWQ 4-bit and AQLM 2-bit quantization; the results are all self-reported by the authors.

The results have not been independently verified and serve only as a reference for quantization selection in personality-sensitive chat applications; preprint v2 was updated on September 15, first submitted on August 26, the page self-labels "in progress", arXiv:2608.25977.

ソース:arxiv.org

リサーチ

Meta self-reports Bumblebee beating HSTU and DLRM on industrial c-NE

クイックテイク
検証済み 2026-09-18 08:35 GMT+8

Bumblebee, a recommendation architecture from Meta's team, achieves an author-reported average c-NE of 0.7895 on industrial data with 155.2 million unique users per 1-hour timestep and 81 billion samples in total, below HSTU and DLRM.

Previously, industrial recommenders used HSTU (0.7964) and DLRM (0.8034) as baselines, with sequence modeling and feature crossing arranged sequentially, leaving the gains of interleaving untapped.

Author-reported: Bumblebee interleaves sequence modeling and feature crossing into stackable blocks, averaging c-NE 0.7895 versus HSTU's 0.7964 and DLRM's 0.8034; interleaving the same components improves NE by 0.20% over sequential arrangement, at the cost of roughly 7-10% slower training throughput.

The results have not been independently reproduced; the data comes from an arXiv preprint (2607.24804, updated to v3 on September 15).

Sources:

  • arXiv abstract page: Bumblebee (2607.24804) ↗
  • Full-text HTML v3 (with experiments and ablation data) ↗

ソース:arxiv.org

リサーチ

Agent skill attack trajectory library opens 3,014 cases

クイックテイク
2026-09-23 12:00 GMT+8

You can now search 3,014 agent skill attack cases, 6,589 trajectories, 233 affected skills and 8 risk categories, drawn from the vetted, desensitised public attack trajectory library SkillAtlas.

Previously, agent skill security reports were mostly private, and existing static, dynamic and benchmark-style evaluations rarely preserved publicly verifiable evidence.

The authors self-report that 42.5% of successful attacks succeeded only after a failed first attempt, and that trajectory-based labels raise pre-execution defence accuracy to 0.770; all these figures are self-reported by the authors.

The preprint was submitted to arXiv on 11 September by Yuxin Tian, Liang Pang and three other authors (arXiv:2609.13353), has been accepted to the EMNLP 2026 REALM workshop as a non-archival paper, and has not yet been reproduced by others.

ソース:arxiv.org

リサーチ

Training-free navigation framework tops 76% success

クイックテイク
検証済み 2026-09-18 08:15 GMT+8

The training-free embodied navigation framework HarnessVLN reaches success rates of 60.8%, 53.9%, 76.0% and 59.3% on the R2R, RxR, HM3D-v2 and HM3D-OVON benchmarks, surpassing prior training-free methods — a new reference point for what training-free approaches can achieve.

Prior training-free methods required separately building perception, memory and execution pipelines, which were hard to coordinate. HarnessVLN uses a unified Agent Harness to coordinate perception, memory and execution tools, and the authors say it has been deployed on a humanoid robot.

These figures are self-reported by Yang Chen and co-authors and have not been independently reviewed.

The preprint was submitted to arXiv on September 14 and updated to v2 on the 16th; the results have not yet been independently reproduced.

ソース:arxiv.org

リサーチ

AXON's dual-path gated attention beats dense baselines on six EEG tasks

クイックテイク
検証済み 2026-09-18 07:40 GMT+8

Readers can now know that AXON exceeds dense attention baselines on average balanced accuracy across six EEG downstream tasks, under both linear probing and full fine-tuning.

Prior EEG modeling relied on dense attention, without separating a temporal path along electrodes from a spatial path across electrodes. AXON proposes adaptive anisotropic attention: attention is split into a temporal path along electrodes and a spatial path across electrodes, with a gate predicting a two-way convex combination per token.

The result is self-reported by Mahir Jain and six other authors; audio-spectrum experiments suggest transferability beyond EEG.

No independent verification yet; the preprint was submitted September 8 and updated to v3 on the 16th, arXiv:2609.08788.

ソース:arxiv.org

リサーチ

DiffAdapterVLA embeds trajectory tokens into driving VLM late layers, low latency and few parameters on NAVSIM

クイックテイク
2026-09-23 12:00 GMT+8

DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, letting trajectory states co-evolve with driving conditions during the backbone's forward computation while training only lightweight adapter modules; the authors self-report high closed-loop planning quality, low latency and few trainable parameters on the NAVSIM benchmark.

Previous approaches did not let trajectory states co-evolve with driving conditions in the backbone's forward computation, making it hard to balance planning quality, latency and parameter count.

The measurement was done first-party by Changxin Lu and seven co-authors: on the NAVSIM benchmark they self-report high closed-loop planning quality, low latency and few trainable parameters.

Boundary: results are author self-reported with no third-party reproduction yet; the preprint was submitted to arXiv on September 14 (v2 updated on the 16th, arXiv:2609.15322).

ソース:arxiv.org

リサーチ

CUA-Net self-reports 91.52% accuracy classifying uterine malformations on external 3D ultrasound set

クイックテイク
検証済み 2026-09-18 07:24 GMT+8

CUA-Net can automatically classify congenital uterine malformations in 3D ultrasound without coronal plane reconstruction, self-reporting 93.88% accuracy on the internal test set and 91.52% on the external test set.

Previously this classification relied on manual reading by sonographers, requiring coronal plane reconstruction, with junior physicians performing below the model.

Self-tested by Yueyue Xu, Yuhao Huang and 12 others: built on 3D ResNet-18, self-reported as outperforming junior sonographers on all metrics and matching senior sonographers on most.

The results are not peer-reviewed or independently verified; it is an arXiv preprint (arXiv:2609.15225), submitted September 14 and revised as v2 on September 15.

ソース:arxiv.org

リサーチ

Human-written narrative descriptions boost zero-shot hidden-narrative detection; few-shot examples often hurt

クイックテイク
検証済み 2026-09-18 07:20 GMT+8

Providing large language models with human-written narrative descriptions significantly improves their zero-shot accuracy at detecting hidden narratives in social messaging; combined with majority-vote ensembling, the authors claim performance comparable to supervised systems.

The prior approach relied on auto-generated descriptions or few-shot examples, but these often reduced accuracy due to subtle framing shifts.

The measurement was run on the Dipromats and SemEval datasets, and the results are self-reported by the three authors, not independently verified.

Boundary: no third-party replication yet; the preprint was submitted September 15 and revised as v2 on September 16 (arXiv:2609.17310).

ソース:arxiv.org

リサーチ

Two automated pipelines screening South African packaged foods against sodium limits agree on only 69.5% of final verdicts

クイックテイク
検証済み 2026-09-18 06:41 GMT+8

When 442 South African packaged foods were screened by image against R214 sodium limits, two automated pipelines agreed on only 69.5% of final verdicts (307/442), with 93.9% agreement on category classification; insufficient-evidence cases were routed to REVIEW rather than forced to a verdict.

Previously such screening relied on manual checking of package labels, with no quantified cross-check of agreement between two independent automated pipelines.

The result was self-reported by Mayimunah Nagayi and four others: one pipeline used YOLO26s region detection plus OCR, the other was an independent Qwen2.5-VL 7B pipeline, and the agreement rates above were measured by comparing the two.

The result is author self-reported, not peer-reviewed, and not yet reproduced by any third party; it comes from an arXiv preprint (arXiv:2609.15427) submitted September 14 and revised as v2 on September 15.

ソース:arxiv.org

リサーチ

GP-BTS removes the batch-size Q factor from regret bounds without initial uncertainty sampling

クイックテイク
検証済み 2026-09-18 06:26 GMT+8

In parallel Gaussian process Bayesian optimization, GP-BTS achieves a regret upper bound without the multiplicative batch-size Q factor, without the initial uncertainty sampling stage required by existing analyses — and the noiseless bound is tighter. This is a self-reported result by Shion Takeno and Shogo Iwazaki.

Previous analyses required an initial uncertainty sampling stage first, which the authors say is often ineffective in practice.

The conclusion is self-reported by the authors: GP-BTS's regret upper bound carries no multiplicative factor of the batch size Q, and the noiseless bound is tighter.

Boundary: the results are not peer-reviewed, and v2 revised Lemma 4.2 of v1; the preprint was submitted August 17 and revised as v2 on September 16 (arXiv:2608.16492).

ソース:arxiv.org

リサーチ

Atomic user model retrieves 8 fields at 23% context to match full-profile personalization

クイックテイク
検証済み 2026-09-18 06:17 GMT+8

Build user profiles into an "Atomic User Model" (AUM) with a stable identity core and retrieve a few fields per task into the prompt: four authors self-report that 8 retrieved fields match the personalized writing quality of the full profile at only 23% of the context (211 vs 915 tokens).

The prior approach injected the entire user profile into the prompt, at high context cost; four preregistered controls show the gains come from profile structure, not the retrieval method.

Experiments used 16 simulated users, 6 style-sensitive task types, and 3 random seeds, self-tested by the authors, not peer-reviewed or independently replicated; the preprint was submitted September 10 and updated to v2 on September 16 (arXiv:2609.12086).

Source: arXiv:2609.12086 abstract page ↗

ソース:arxiv.org

リサーチ

Volumetric harmonic field navigation validated on a physical Crazyflie quadrotor for the first time, beating Dijkstra guidance on clearance and jerk

クイックテイク
検証済み 2026-09-18 06:08 GMT+8

Volumetric harmonic field navigation has been validated on a physical Crazyflie quadrotor for the first time: three authors self-report combining a precomputed volumetric harmonic field with a constrained predictive planner, querying field values at predicted positions rather than extracting a global reference path, and state that to their knowledge this is the first physical quadrotor demonstration of the method.

The prior approach was to extract a global reference path; in structured 3D tests, harmonic guidance achieved larger minimum clearance and lower RMS jerk than matched Dijkstra guidance, at the cost of longer paths.

Physical execution was verified in Crazyflie hardware trials; the measurements are self-reported by the authors, with no independent verification or peer review yet.

Boundary: the results have not been independently reproduced; the preprint was submitted September 14 and updated to v2 on September 16, see arXiv:2609.15680.

ソース:arxiv.org

リサーチ
次の読書ページ →