Large Language Models

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
arXiv
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
Agents' Last Exam
arXiv
Agents' Last Exam
Asymmetric Capacity Allocation in Self-Refinement Pipelines
arXiv
Asymmetric Capacity Allocation in Self-Refinement Pipelines
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
ACL 2026 Findings
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
EMNLP 2026
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Mitigating Hallucinations in Large Language Models via Causal Reasoning
AAAI 2026
Mitigating Hallucinations in Large Language Models via Causal Reasoning
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
ACL 2025 Findings
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
A Large-Scale Simulation on Large Language Models for Decision-Making in Political Science
arXiv
A Large-Scale Simulation on Large Language Models for Decision-Making in Political Science
FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
ACM MM 2024
FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning