When Does Non-Uniform Replay Matter in Reinforcement Learning?
arXiv:2605.10236v3 Announce Type: replace-cross Abstract: Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains…
arXiv:2605.10236v3 Announce Type: replace-cross Abstract: Modern off-policy reinforcement learning algorithms often rely on simple uniform replay sampling and it remains…
arXiv:2605.15652v2 Announce Type: replace-cross Abstract: Vector-HaSH and the Tolman-Eichenbaum Machine propose the hippocampal-entorhinal circuit factorizes content from a grid-cell scaffold,…
arXiv:2605.12770v3 Announce Type: replace-cross Abstract: We introduce WriteSAE, the first sparse autoencoder that decomposes and edits the matrix cache write…
arXiv:2605.16391v1 Announce Type: cross Abstract: Inertial measurement units (IMUs) are fundamental sensing components in multi-source integrated navigation systems, and their…
arXiv:2605.17862v1 Announce Type: cross Abstract: Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution…
arXiv:2605.16393v1 Announce Type: cross Abstract: Semantic segmentation is essential for analysing anatomical features in biomedical research, yet a performance gap…
arXiv:2605.15202v1 Announce Type: new Abstract: Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize the…
arXiv:2605.15053v2 Announce Type: replace-cross Abstract: Continually pre-training a large language model on heterogeneous text domains, without replay or task labels,…
arXiv:2605.15120v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated…
arXiv:2605.16048v1 Announce Type: cross Abstract: State Space Models (SSMs) are inherently recurrent along the sequence dimension, yet depth-recurrence – reusing…