CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
arXiv:2511.18889v1 Announce Type: cross Abstract: Data contamination poses a significant challenge to the fairness of LLM evaluations in natural language…
