Multi-LLM Thematic Analysis with Dual Reliability Metrics: Combining Cohen’s Kappa and Semantic Similarity for Qualitative Research Validation
arXiv:2512.20352v1 Announce Type: cross Abstract: Qualitative research faces a critical reliability challenge: traditional inter-rater agreement methods require multiple human coders,…
Clust-PSI-PFL: A Population Stability Index Approach for Clustered Non-IID Personalized Federated Learning
arXiv:2512.20363v1 Announce Type: cross Abstract: Federated learning (FL) supports privacy-preserving, decentralized machine learning (ML) model training by keeping data on…
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
arXiv:2512.19025v2 Announce Type: replace-cross Abstract: Machine unlearning aims to remove specific data influences from trained models, a capability essential for…
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
arXiv:2511.16193v3 Announce Type: replace-cross Abstract: Rollout dominates the training time in large language model (LLM) post-training, where the trained model…
Bidirectional human-AI collaboration in brain tumour assessments improves both expert human and AI agent performance
arXiv:2512.19707v1 Announce Type: cross Abstract: The benefits of artificial intelligence (AI) human partnerships-evaluating how AI agents enhance expert human performance-are…
The 6th International Verification of Neural Networks Competition (VNN-COMP 2025): Summary and Results
arXiv:2512.19007v1 Announce Type: cross Abstract: This report summarizes the 6th International Verification of Neural Networks Competition (VNN-COMP 2025), held as…
Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
arXiv:2512.19011v1 Announce Type: cross Abstract: Prompt injection and jailbreaking attacks pose persistent security challenges to large language model (LLM)-based systems.…
TakeAD: Preference-based Post-optimization for End-to-end Autonomous Driving with Expert Takeover Data
arXiv:2512.17370v2 Announce Type: replace-cross Abstract: Existing end-to-end autonomous driving methods typically rely on imitation learning (IL) but face a key…
Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
arXiv:2512.06951v2 Announce Type: replace-cross Abstract: We present a vision-action policy that won 1st place in the 2025 BEHAVIOR Challenge –…
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
arXiv:2507.12898v4 Announce Type: replace-cross Abstract: Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous…
