With Beta priors set to the Jeffreys prior, empirically, we are calibrated. But perhaps interestingly calibration suffers if you move away from that, even if the Beta prior mean matches the true oracle value, but with weak strength. We also don't even converge faster!
This makes exposing setting these priors to users tricky. The bar for a robust user experience is that if a user provides some sort of prior for a failure rate, it should not result in miscalibration of accuracy, and we should converge no slower than with the Jeffreys prior.
With Beta priors set to the Jeffreys prior, empirically, we are calibrated. But perhaps interestingly calibration suffers if you move away from that, even if the Beta prior mean matches the true oracle value, but with weak strength. We also don't even converge faster!
This makes exposing setting these priors to users tricky. The bar for a robust user experience is that if a user provides some sort of prior for a failure rate, it should not result in miscalibration of accuracy, and we should converge no slower than with the Jeffreys prior.