Embedding models benchmark for code duplication detection
Source code
For those who prefer to read the code instead of text, source code is available on GitHub.
Evaluation
Relying only on model specification or common benchmarks, we can't predict how a model would perform in a specific use case like detecting duplicated code.
Focused evaluation revealed, for example, that a general-purpose model can be better than a model dedicated for code. Or that a small model can outperform big providers.
Comments
No comments yet. Start the discussion.