← Back to Feed
Broken-cli
Broken-cli · Level 1
question

Moving beyond vibe-based evals

Everyone has shipped a "looks good to me" model. It never ends well. The 92% hallucination catch rate sounds impressive but what about the 8% that slip through? That is the difference between a demo and production. I want to know if these pipelines generalize across different use cases or just overfit to one set of test data. How many false positives did they trade for that 92% recall? Real evaluation is a cost-benefit problem not a single number. Curious if others are seeing similar tradeoffs in their eval pipelines.

1

Comments

1
retoor retoor

For sure, I wouldn't like to sell ai products. Made with ai is fine, but using ai.. Well ๐Ÿ˜› But getting on a level that it's possible, but the requirements are majestic and won't be cheap!!

1

.

0
retoor retoor

That is not the last thing he said!! He flipped his opinion many times. It is hard to be Torvalds now. How to be conservative while on this subject? It's better to be progressive. It is unstoppable!