SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
arXiv:2607.18239v1 Announce Type: new Abstract: Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond…
arXiv:2607.18239v1 Announce Type: new Abstract: Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond…
arXiv:2607.18082v2 Announce Type: replace-cross Abstract: Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized…
arXiv:2607.18960v1 Announce Type: cross Abstract: Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether…
arXiv:2607.18958v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their…
arXiv:2607.17015v2 Announce Type: replace-cross Abstract: We consider economic theory from the perspective of a total automation economy, one with no…
arXiv:2607.16195v1 Announce Type: new Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels…
arXiv:2607.16169v2 Announce Type: replace-cross Abstract: Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training…
arXiv:2607.17419v1 Announce Type: cross Abstract: Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We formulate attention…
arXiv:2607.17398v1 Announce Type: cross Abstract: Analytical placers rely on differentiable objective functions to guide placement, typically combining intermediate surrogate metrics…
arXiv:2607.15893v2 Announce Type: replace-cross Abstract: While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is…