StochBench: a Lean 4 benchmark of 450 graduate stochastic-processes problems with natural-language sources
Read the original at arxiv.org→arXiv:2609.09264v1 Announce Type: new Abstract: Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly...
Original headline: "StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean"
Coverage timeline
- Sep 10, 04:00 UTC arXiv cs.CL lead source StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean