Human-AI-powered hypothesis testing: cost-aware selective AI scoring and sequential human escalation
Read the original at arxiv.org→arXiv:2609.28859v1 Announce Type: new Abstract: Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet...
Original headline: "Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation"
Coverage timeline
- Sep 25, 04:00 UTC arXiv cs.AI lead source Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation