Auditing pairwise equivalence judgments: self-critique effects and diversity measurement in multi-agent hypothesis generation
Read the original at arxiv.org→arXiv:2610.04133v1 Announce Type: new Abstract: Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and...
Original headline: "Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation"
Coverage timeline
- Oct 6, 04:00 UTC arXiv cs.AI lead source Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation