Publications

The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning
arXiv
The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
NeurIPS 2026 E&D
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
The Orchestration Gap: Why Process Automation Stalls in Operationally Complex Industries
arXiv
The Orchestration Gap: Why Process Automation Stalls in Operationally Complex Industries
Agents' Last Exam
NeurIPS 2026 E&D
Agents' Last Exam
Artificial Intelligence-Aided Digital Twin Design: A Systematic Review and Future Directions
Preprints 2026
Artificial Intelligence-Aided Digital Twin Design: A Systematic Review and Future Directions
Asymmetric Capacity Allocation in Self-Refinement Pipelines
arXiv
Asymmetric Capacity Allocation in Self-Refinement Pipelines
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
ACL 2026 Findings
LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines
Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability
arXiv
Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
EMNLP 2026
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Mitigating Hallucinations in Large Language Models via Causal Reasoning
AAAI 2026
Mitigating Hallucinations in Large Language Models via Causal Reasoning
H-FedSN: Personalized Sparse Networks for Efficient and Accurate Hierarchical Federated Learning for IoT Applications
Nature AI’26
H-FedSN: Personalized Sparse Networks for Efficient and Accurate Hierarchical Federated Learning for IoT Applications
NLP-ADBench: NLP Anomaly Detection Benchmark
EMNLP 2025 Findings
NLP-ADBench: NLP Anomaly Detection Benchmark
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
ACL 2025 Findings
AD-LLM: Benchmarking Large Language Models for Anomaly Detection
A Large-Scale Simulation on Large Language Models for Decision-Making in Political Science
arXiv
A Large-Scale Simulation on Large Language Models for Decision-Making in Political Science
FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning
ACM MM 2024
FedBCGD: Communication-Efficient Accelerated Block Coordinate Gradient Descent for Federated Learning