Cool paper on catching alignment failures with Jev.
The ideas is to ask Jev one generic yes/no question about a model's response, and use its probability as a score.
With no extra training, that score separates failures from good responses well, with a median AUROC of 0.886. https://t.co/bCO4QksWjL
The ideas is to ask Jev one generic yes/no question about a model's response, and use its probability as a score.
With no extra training, that score separates failures from good responses well, with a median AUROC of 0.886. https://t.co/bCO4QksWjL
12
