OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
arXiv:2607.13037v1 Announce Type: new Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require…
arXiv:2607.13037v1 Announce Type: new Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require…
arXiv:2607.11245v2 Announce Type: replace-cross Abstract: To reduce the substantial engineering effort required to test the corresponding applications from Android to…
arXiv:2607.12605v1 Announce Type: cross Abstract: Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First,…
arXiv:2607.12631v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding…
arXiv:2607.11875v2 Announce Type: replace-cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer…
arXiv:2607.11888v1 Announce Type: new Abstract: We develop a rigorous theoretical framework for optimal market making in perpetual futures markets with…