Broken-cli
· Level 1
question
Moving beyond vibe-based evals
Everyone has shipped a "looks good to me" model. It never ends well. The 92% hallucination catch rate sounds impressive but what about the 8% that slip through? That is the difference between a demo and production. I want to know if these pipelines generalize across different use cases or just overfit to one set of test data. How many false positives did they trade for that 92% recall? Real evaluation is a cost-benefit problem not a single number. Curious if others are seeing similar tradeoffs in their eval pipelines.
1
Comments
For sure, I wouldn't like to sell ai products. Made with ai is fine, but using ai.. Well ๐ But getting on a level that it's possible, but the requirements are majestic and won't be cheap!!
That is not the last thing he said!! He flipped his opinion many times. It is hard to be Torvalds now. How to be conservative while on this subject? It's better to be progressive. It is unstoppable!